Working and Architecture of GPT

Generative Pre-trained Transformers also knows as GPT. It is a Transformer that is pretained on the input tokens and generates the output in the form of next prediction token.
Example: Why the color of sky is blue?
Now how our GPT/LLM will process this input and generate the next prediction token as output?
Let's undersand the actual workign and architecture of GPT.
Step-1: Tokenization: It is basically the bifircation of natural language input text into tokens.
Input: Why the color of sky is blue?
Our GPT will first visualize the input and breaks it into the tokens and assigns each token a token ID as below.
Now, after this it will produce one next predicted token based on the input tokens bifercations. So, behind the computations scenes will be like Why the color of sky is blue? + because (this is the generetd next predition token) then this combitnation will go again the model as the new input and then it will generated the next predicted token.
Next will be : Why the color of sky is blue? + because + air molecules and so on...
Step-2: Input Embeddings: It will create a 3D visualization of each token in the vector space as shown below. Basically, Input embeddings are just a way of sorting words by how they feel and what they mean.
Step-3: Positional Encodings: Positional encoding is a secret number tag given to each word so the computer knows the exact order they stand in a line.
Words change their whole meaning if they move around. Lets's say "dog bites boy," it is sad. If you say "boy bites dog," it is very strange and funny! The words are the same, but the order makes all the difference. Why the Model Gets Confused:
Mixing up words: When a model looks at input embeddings, it sees a big pile of words sorted by what they mean, but it forgets who came first.
Finding the spot: Without a little extra help, the model reads all the words at the exact same time as a big soup.
How Positional Encoding Fixes It?
Number tags: The model glues a tiny extra number tag to every word.
Standing in line: The first word gets tag number one, the second word gets tag number two, and so on.
Keeping track: Now, even if the word is mixed up in the big box, the model looks at the number tag and says, "Ah, this word goes at the very front of the sentence."
Step-4: Self Attention Mechanism: Self-attention is a superpower that lets every single word in a sentence look at all the other words at the same time to figure out who belongs with whom.
Step-5: Feed-Forward Network: Also know as the Thinking Cap every word goes through a mini brainstorming machine that processes everything it just learned to decide what word should come next!
Architecture:


