Fake News Classifier
A supervised text classifier that reads a news article and scores it as real or fake. Paste any article URL below — the page is scraped, flattened into the same shape the model was trained on, and run through the pipeline.
Try it
Confidence is the model’s own probability for the class it picked, not a measure of whether it is right. Treat results on unfamiliar sites as a rough signal rather than a verdict.
How it works
- Scrape. The URL is fetched and stripped of navigation, scripts, and boilerplate. The headline comes from
og:titlewith anh1fallback; the body is the first container that yields at least three real paragraphs, so a stray sidebar does not get mistaken for the story. - Normalize. Punctuation is stripped, text is tokenized and lowercased, English stopwords are dropped, and the remainder is stemmed with a Lancaster stemmer — the same transform used to fit the vectorizer.
- Vectorize and score. The cleaned tokens are TF-IDF weighted and passed to a multi-layer perceptron, which returns a class and a probability.
The label leak worth knowing about
The training corpus carries a subject column, and it separates the classes perfectly: politicsNews and worldnews are real without exception, the other six subjects are fake without exception. Training on it produces gorgeous accuracy and a model that has learned the column, not the writing.
A scraped page carries no such label, so this demo deliberately omits it. Supply one and the leak does the work — prefixing politicsNews is enough to make the model call an Onion piece real, while the same article with no prefix is correctly called fake. Leaving it off keeps the prediction resting on the headline and body, which is the only part that generalizes.
How the demo is served
The pipeline is a ~430 MB pickled scikit-learn object, which cannot run inside a Next.js route. Instead the Node process spawns a small Python service as a child process at server start, and the site talks to it over loopback through /api/classify. The model is unpickled once on a background thread while Next boots, so the load cost is paid before the first click rather than during it — and the sidecar binds 127.0.0.1 only, so the single public entry point stays the rate-limited API route.