#method to get shape of tensor flow element
Saturday, March 11, 2023
Tensorflow general methods
Wednesday, March 8, 2023
Synchronously shuffle X,Y
import numpy as np
np.random.seed(seed)
m = X.shape[1] # number of training examples
permutation = list(np.random.permutation(m))
shuffled_X = X[:, permutation]
shuffled_Y = Y[:, permutation].reshape((1, m))
Saturday, March 4, 2023
Dropout
- Dropout is a regularization technique.
- You only use dropout during training. Don't use dropout (randomly eliminate nodes) during test time.
- Apply dropout both during forward and backward propagation.
- During training time, divide each dropout layer by keep_prob to keep the same expected value for the activations. For example, if keep_prob is 0.5, then we will on average shut down half the nodes, so the output will be scaled by 0.5 since only the remaining half are contributing to the solution. Dividing by 0.5 is equivalent to multiplying by 2. Hence, the output now has the same expected value. You can check that this works even when keep_prob is other values than 0.5.
Friday, March 3, 2023
python - Initialization of weights
The main difference between Gaussian variable (numpy.random.randn()) and uniform random variable is the distribution of the generated random numbers:
- numpy.random.rand() produces numbers in a uniform distribution.
- and numpy.random.randn() produces numbers in a normal distribution.
When used for weight initialization, randn() helps most the weights to Avoid being close to the extremes, allocating most of them in the center of the range.
An intuitive way to see it is, for example, if you take the sigmoid() activation function.
You’ll remember that the slope near 0 or near 1 is extremely small, so the weights near those extremes will converge much more slowly to the solution, and having most of them near the center will speed the convergence.
Initialization of weights
- The weights
𝑊[𝑙] should be initialized randomly to break symmetry. - However, it's okay to initialize the biases
𝑏[𝑙] to zeros. Symmetry is still broken so long as𝑊[𝑙] is initialized randomly. - Initializing weights to very large random values doesn't work well.
- Initializing with small random values should do better.
Wednesday, March 1, 2023
python code to plot cost
import matplotlib.pyplot as plt
%matplotlib inline
def plot_costs(costs, learning_rate=0.0075):
plt.plot(np.squeeze(costs))
plt.ylabel('cost')
plt.xlabel('iterations (per hundreds)')
plt.title("Learning rate =" + str(learning_rate))
plt.show()
#Assuming "costs" is a list of costs obtained during training iterations per hundred
#calling the method with some learning rate
plot_costs(costs, learning_rate)
output:
Deep Learning methodology using gradient descent
Usual Deep Learning methodology to build the model:
- Initialize parameters / Define hyperparameters
- Loop for num_iterations:
Sunday, December 18, 2022
split data set to train, cross validation and test sets
print(f"the shape of the original set (input) is: {x.shape}")
print(f"the shape of the original set (target) is: {y.shape}\n")
Tuesday, December 13, 2022
Epochs and batches
We provide epoch value while fitting/training the model as below.
Example: model.fit(X,Y,epoch=100)
Epochs and batches
In the fit statement above, the number of epochs was set to 100. This specifies that the entire data set
should be applied during training 100 times. During training, you see output describing the progress of
training that looks like this:
Epoch 1/100
157/157 [==============================] - 0s 1ms/step - loss: 2.2770The first line, Epoch 1/100, describes which epoch the model is currently running. For efficiency,
the training data set is broken into 'batches'. The default size of a batch in Tensorflow is 32.
if given an model has are 5000 examples(X_train) it will set or roughly to 157 batches.
The notation on the 2nd line 157/157 [==== is describing which batch has been executed.
Monday, December 12, 2022
Derivative using python
Libraries for derivative
from sympy import symbols, diff
Let's try this out. Let's look at the derivative of the function
Sunday, December 11, 2022
SparseCategorialCrossentropy or CategoricalCrossEntropy
Tensorflow has two potential formats for target values and the selection of the loss defines which is expected.
- SparseCategorialCrossentropy: expects the target to be an integer corresponding to the index. For example, if there are 10 potential target values, y would be between 0 and 9.
- CategoricalCrossEntropy: Expects the target value of an example to be one-hot encoded where the value at the target index is 1 while the other N-1 entries are zero. An example with 10 potential target values, where the target is 2 would be [0,0,1,0,0,0,0,0,0,0].
Friday, December 9, 2022
Get the output of each layer in Neural network
Lets Consider following simple neural network
import keras.backend as K
from keras.models import Model
from keras.layers import Input, Dense
input_layer = Input((10,))
layer_1 = Dense(10)(input_layer)
layer_2 = Dense(20)(layer_1)
layer_3 = Dense(5)(layer_2)
output_layer = Dense(1)(layer_3)
model = Model(inputs=input_layer, outputs=output_layer)# some random input
import numpy as np
features = np.random.rand(100,10)and consider this model is trained
# With a Keras function get the ouputs of all the layers
get_all_layer_outputs = K.function([model.layers[0].input],
[l.output for l in model.layers[0:]])
layer_output = get_all_layer_outputs([features]) # return the same thing
#layer_output is a list of all layers outputs
#if the model is trained you will get the output for input with trained weights other wise
it will give outout with initial weights