NeuroText MLP

Building a character-level Multi-Layer Perceptron (MLP) language model from scratch in PyTorch. This project extends the Bigram model by learning character embeddings and predicting the next character using multiple previous characters as context.
Overview
NeuroText MLP is the next step in my journey of building language models from scratch.
After implementing statistical and neural Bigram models, this project introduces a Multi-Layer Perceptron (MLP) language model that learns richer representations of text through character embeddings and hidden layers.
Instead of predicting the next character using only the previous one, the model learns from multiple preceding characters, enabling it to capture more meaningful patterns in language.
Motivation
Bigram models are excellent for understanding probability and language modeling fundamentals, but they have a significant limitation—they only consider one previous character when making predictions.
The goal of this project is to overcome that limitation by introducing:
- Character embeddings
- Context windows
- Hidden layers
- Non-linear activations
- Gradient-based optimization
These ideas form the foundation of more advanced neural language models.
Architecture
The model follows a simple neural language modeling pipeline:
Previous Characters
│
▼
Character Embeddings
│
▼
Embedding Concatenation
│
▼
Hidden Layer (MLP)
│
▼
Output Layer
│
▼
Probability Distribution
│
▼
Next Character Prediction