# Part 22: Gradient Descent: How Neural Networks Find Better Weights

Imagine you are standing on top of a mountain, but it's completely dark. 🌙

Your goal is to reach the **lowest point in the valley**.

You can't see the entire path, so what would you do?

You might take a small step, check which direction goes downward, and continue taking steps in that direction until you reach the bottom.

This is the basic idea behind **Gradient Descent**.

In machine learning, Gradient Descent helps a model find better parameter values by gradually moving in a direction that **reduces the error (loss)**.

## 🤔 **What Is Gradient Descent?**

**Gradient Descent is an optimization algorithm that helps a machine learning model reduce its error by gradually adjusting its parameters in the direction that lowers the loss.**

Remember our previous blog on **Backpropagation**?

Backpropagation calculates the **gradient for each weight**, showing how a change in that weight affects the loss.

Gradient Descent then uses these gradients to **update the weights in a direction that reduces the loss**.

In simple terms:

**"Backpropagation calculates the gradients. Gradient Descent uses those gradients to update the weights."**

Together, they help neural networks learn from their mistakes.

![](https://cdn.hashnode.com/uploads/covers/6942cafec9d5320a12aa01b1/e7c3fe65-1532-4a9d-a94d-464bba5c13b1.png align="center")

## ⛰ **Think of It Like Walking Down a Mountain**

Imagine the loss of a model as a mountain landscape.

The **height represents the amount of error**.

*   High point → High error ❌
    
*   Low point → Low error ✅
    

Our goal is to reach the lowest point, also called the **minimum**.

But the model doesn't know where the lowest point is.

So, it uses the gradient to determine which direction leads toward increasing loss and then moves in the **opposite direction**, toward lower loss.

This repeated process of moving toward lower loss is called **Gradient Descent**.

## 🧭 **What Does “Gradient” Mean?**

A **gradient** tells us how the loss changes when a model parameter, such as a weight, changes.

Think of standing on a hill.

The gradient points toward the direction where the loss increases most rapidly.

Since we want to **reduce** the loss, Gradient Descent moves in the opposite direction of the gradient.Here:

The weight-update formula is:

![](https://cdn.hashnode.com/uploads/covers/6942cafec9d5320a12aa01b1/02fc2305-7523-4927-a448-f81e639ca89d.png align="left")

Here:

![](https://cdn.hashnode.com/uploads/covers/6942cafec9d5320a12aa01b1/2e6ba6af-53dc-492b-b0f4-2f7de8111534.png align="left")

![](https://cdn.hashnode.com/uploads/covers/6942cafec9d5320a12aa01b1/24ffa827-b4a7-49f5-9aa6-ea0827e7ec95.png align="left")

![](https://cdn.hashnode.com/uploads/covers/6942cafec9d5320a12aa01b1/330cde03-62de-4ad6-b222-5f29c7974572.png align="left")

![](https://cdn.hashnode.com/uploads/covers/6942cafec9d5320a12aa01b1/a5e9691a-3eb8-45dc-9a56-57d531544ce9.png align="left")

The formula basically says:

***"Take the current weight and adjust it in the direction that reduces the loss."***

## ⚖️ **Weights vs. Gradients**

It is easy to confuse these two, but they have different roles.

<table style="width: 673px;"><colgroup><col style="width: 108px;"><col style="width: 282px;"><col style="width: 283px;"></colgroup><tbody><tr><td colspan="1" rowspan="1" colwidth="108"><p>&nbsp;</p></td><td colspan="1" rowspan="1" colwidth="282"><p><strong>Weights</strong></p></td><td colspan="1" rowspan="1" colwidth="283"><p><strong>Gradients</strong></p></td></tr><tr><td colspan="1" rowspan="1" colwidth="108"><p><strong>What are they?</strong></p></td><td colspan="1" rowspan="1" colwidth="282"><p>Numbers that control how strongly inputs influence the model's output.</p></td><td colspan="1" rowspan="1" colwidth="283"><p>Numbers that tell us how changing a weight affects the loss.</p></td></tr><tr><td colspan="1" rowspan="1" colwidth="108"><p><strong>Purpose</strong></p></td><td colspan="1" rowspan="1" colwidth="282"><p>Help the model make predictions.</p></td><td colspan="1" rowspan="1" colwidth="283"><p>Tell the model how the weights should change.</p></td></tr><tr><td colspan="1" rowspan="1" colwidth="108"><p><strong>Example</strong></p></td><td colspan="1" rowspan="1" colwidth="282"><p>A weight might be <strong>0.5</strong>.</p></td><td colspan="1" rowspan="1" colwidth="283"><p>Its gradient might be <strong>-0.2</strong>.</p></td></tr><tr><td colspan="1" rowspan="1" colwidth="108"><p><strong>During training</strong></p></td><td colspan="1" rowspan="1" colwidth="282"><p>They are updated.</p></td><td colspan="1" rowspan="1" colwidth="283"><p>They are calculated from the loss.</p></td></tr></tbody></table>

So remember:

**Weights → Make predictions**

**Gradients → Guide corrections**

**Gradient Descent → Updates weights**

## 👣 **What Is the Learning Rate?**

The **learning rate is a value that controls how big a step the model takes when updating its weights.**

Think of it as the size of your steps while walking down a mountain.

### **Large learning rate** 🚶

The model takes bigger steps.

It may move toward the minimum faster, but it could **overshoot the lowest point** or make the training unstable.

### **Small learning rate** 🐢

The model takes smaller steps.

It can make more careful adjustments, but it may take much longer to reach a good solution.

Therefore, the learning rate needs to be chosen carefully so the model can reduce its loss efficiently.

## 🔁 **How Does Gradient Descent Work?**

The process can be simplified into a few steps:

### 1\. Start with initial weights:

The model begins with some initial values.

### 2\. Make a prediction

The model uses those weights to produce an output.

### 3\. Calculate the loss

We measure how far the prediction is from the correct answer.

### 4\. Calculate the gradients

Backpropagation calculates how each weight affects the loss.

### 5\. Update the weights

Gradient Descent uses the gradients to adjust the weights in a direction that reduces the loss.

### 6\. Repeat

The model repeats these steps for many iterations, gradually reducing the loss.

## 🍳 **A Simple Real-Life Example**

Imagine you're cooking a dish for the first time.

You taste it and realize:

“It's too salty.”

So, next time, you add less salt.

You taste it again:

“Now it's too bland.”

So you make another small adjustment.

After several attempts, you get the taste closer to what you want.

That's similar to Gradient Descent.

The **difference in taste** is like the error, while changing the amount of an ingredient is like **changing a model's weights**.

The information about how much an ingredient should be changed is similar to a **gradient**, while the size of your adjustment is influenced by the **learning rate**.

The goal isn't to get everything perfect in one attempt, but to **gradually improve through repeated corrections**.

## 🚀 **Final Thoughts**

Gradient Descent may sound complicated, but its core idea is simple:

**Measure the error, calculate the gradients, adjust the weights, and repeat.**

By making these adjustments repeatedly, a model can gradually reduce its loss and improve its predictions.

This combination of **loss, gradients, and weight updates** is one of the fundamental ideas behind how machine learning models learn.

## 🔮 Coming Up Next

In the next blog, we’ll explore the **Types of Gradient Descent** and understand how Batch, Stochastic, and Mini-Batch Gradient Descent help neural networks learn and update their weights efficiently.
