Neural Network Fundamentals - Part 2: Training Neural Networks
How Neural Networks Learn by Adjusting Their Weights from Training Data.
This part is sequel to: Neural Network Fundamentals - Part 1: Understanding How Neural Networks Work
1. Initializing the Weights
Before a neural network can start training, it needs values for its weights and biases.
These values are called the model's initial parameters.
Usually, the weights are initialized with small random values, while biases are often initialized to zero.
For example:
These initial values are not meaningful predictions. They simply give the network a starting point.
The network will then use training data to gradually adjust these values.
Initial weights
↓
Make predictions
↓
Compare with expected output
↓
Adjust weights
↓
Make better predictions
A neural network therefore does not start with the correct weights. Training is the process of finding weights and biases that produce useful predictions.
2. What Is Training?
Training is the process of adjusting a neural network's weights and biases so that its predictions become closer to the expected outputs.
The basic idea is:
Training Data
↓
Neural Network
↓
Prediction
↓
Compare with Expected Output
↓
Adjust Weights
↓
Repeat
For example, suppose we are training a network to classify cats and dogs.
The network starts with its initial weights and makes predictions.
Those predictions will usually be wrong or imperfect at first.
The network then uses the difference between its prediction and the expected output to determine how its parameters should be adjusted.
This process is repeated over many training examples, gradually improving the model's predictions.
The key idea is:
Training means learning the weights and biases from examples.
3. Training Data
A neural network learns from training data. Training data consists of examples containing:
- Inputs: the features given to the network.
- Expected outputs: the correct answers for those inputs.
For example, for a cat-versus-dog classifier:
| Weight | Height | Ear Length | Expected Output |
|---|---|---|---|
| 4 kg | 25 cm | 6 cm | Cat |
| 15 kg | 45 cm | 12 cm | Dog |
| 5 kg | 28 cm | 7 cm | Cat |
| 20 kg | 50 cm | 14 cm | Dog |
The network takes the input features and produces a prediction.
It then compares that prediction with the expected output.
A training dataset normally contains many examples, not just the few examples shown above.
The variety of examples is important because the network needs to learn a general relationship between the inputs and outputs rather than simply memorize individual examples.
In the next step, we need a way to measure how different the prediction is from the expected output.
This is the role of the loss function.
4. Loss Function
After making a prediction, the neural network needs to know how wrong that prediction is.
A loss function measures the difference between the network's prediction and the expected output.
For example:
The prediction is close to the expected output, so the loss should be relatively small.
But if:
the prediction is much further away, so the loss should be larger.
Conceptually:
The loss function converts the prediction error into a number.
Small loss → prediction is closer to expected output
Large loss → prediction is further from expected output
During training, the goal is to reduce the loss by changing the network's weights and biases.
However, the loss value alone does not tell us exactly which weights should change or by how much.
That is where backpropagation comes in.
5. Backpropagation
Backpropagation is the process used to determine how the neural network's weights and biases should be adjusted to reduce the loss.
The network first makes a prediction and calculates the loss:
Backpropagation then works backward from the loss, calculating how the loss changes with respect to each weight and bias.
For example:
A gradient tells us the direction and magnitude of change needed for a parameter to reduce the loss.
The weights are then updated using these gradients.
So, at a high level:
Backpropagation determines how the weights and biases should change to reduce the error.

The actual mathematical update of the weights is performed using an optimization algorithm, such as gradient descent.
6. Gradient Descent
Once backpropagation has calculated the gradients, the neural network needs a way to use those gradients to adjust its weights and biases.
Gradient descent is the most widely used optimization algorithm for minimizing the loss in neural networks.
The basic idea is simple:
The weights are adjusted in the direction that reduces the loss.
Other commonly used optimization algorithms include:
- Stochastic Gradient Descent (SGD)
- Momentum
- RMSProp
- Adam
- Adagrad
Different optimizers use different strategies for updating the weights, but their goal is the same: reduce the loss and improve the network's predictions.
The mathematical details of gradient descent and other optimization algorithms will be covered separately.
Different loss functions are also used for different types of neural network tasks. Examples include:
- Mean Squared Error (MSE)
- Mean Absolute Error (MAE)
- Cross-Entropy Loss
- Binary Cross-Entropy
- Categorical Cross-Entropy
Various loss functions and when to use them will be covered in other articles in this AI in Context series.
7. Putting Training Together
We can now put the complete training process together.
For each training example, the neural network:
1. Takes the input x
↓
2. Performs forward propagation
↓
3. Produces a prediction
↓
4. Calculates the loss
↓
5. Backpropagation calculates the gradients
↓
6. Gradient descent adjusts the weights and biases
↓
7. Repeat with more training examples
This process is repeated many times over the training data.
As the weights and biases are adjusted, the network generally becomes better at producing predictions that match the expected outputs.
In simple terms:
Training a neural network means repeatedly making predictions, measuring the error, calculating how the weights should change, and updating them to reduce the error.
Relevant Link(s)
Neural Network Fundamentals - Part 1: Understanding How Neural Networks Work