Building TheBulletin AI: An Autonomous, Bias-Aware News Aggregator

In the modern digital landscape, staying informed is harder than ever. It鈥檚 not because of a lack of information, but an abundance of it鈥攕pecifically, the polarization, sensationalism, and sheer volume of duplicate reporting. Traditional aggregators simply fetch headlines and list them side-by-side. To solve this, I built TheBulletin AI (historically codenamed Veritas Aggregator): a fully autonomous, bias-aware news aggregator. Instead of just organizing links, it acts as an automated editorial room that ingests articles, semantically clusters them into single events, writes objective summaries, and evaluates the political bias and credibility of the coverage. ...

June 7, 2026 路 6 min 路 Pranav Buradkar

Building TheBulletin AI: An Autonomous, Bias-Aware News Aggregator

馃敆 GitHub Repository: pranavburadkar/veritas_aggregator In the modern digital landscape, staying informed is harder than ever. It鈥檚 not because of a lack of information, but an abundance of it鈥攕pecifically, the polarization, sensationalism, and sheer volume of duplicate reporting. Traditional aggregators simply fetch headlines and list them side-by-side. To solve this, I built TheBulletin AI (historically codenamed Veritas Aggregator): a fully autonomous, bias-aware news aggregator. Instead of just organizing links, it acts as a automated editorial room that ingests articles, semantically clusters them into single events, writes objective summaries, and evaluates the political bias and credibility of the coverage. ...

June 7, 2026 路 6 min 路 Pranav Buradkar

A Simple Guide to LLM Serving: Quantization, KV Caching, and Continuous Batching

When you type a prompt into ChatGPT or Claude, the model generates a response word-by-word (or token-by-token) in real-time. Behind this smooth user interface lies a massive engineering challenge: LLM Inference is incredibly expensive and slow. If you run a base LLM without optimizations, it will devour your GPU memory, process requests one-by-one, and make your users wait. To solve this, developers and researchers use three core serving techniques that work in tandem to speed up generation by up to 10x while drastically cutting hardware costs. ...

May 10, 2026 路 5 min 路 Pranav Buradkar

A Simple Guide to LoRA (Low-Rank Adaptation)

Have you ever tried to download or fine-tune a modern Large Language Model (LLM) like Llama 3 or Mistral? If so, you probably ran into a massive wall: GPU memory. Fine-tuning models with billions of parameters requires specialized, high-end hardware, costing thousands of dollars. But what if you could achieve the exact same performance by training less than 1% of the model鈥檚 parameters? That is the magic of LoRA (Low-Rank Adaptation). It is currently the most popular technique for making fine-tuning fast, cheap, and accessible to everyone. In this guide, we鈥檒l break down how it works, why it is so effective, and how you can use it. ...

March 20, 2026 路 4 min 路 Pranav Buradkar