Future Ventures: Scaling with Clarity

Nikola Borisov — Why the Real AI Battle Isn’t Training—It’s Deployment | Future Ventures Podcast Ep. 023

• Maxim Atanassov

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 57:34

Send us Fan Mail

Nikola Borisov is CEO and co-founder of DeepInfra, offering open-source models like DeepSeek, Llama, Kimi, GLM, and GPT-OSS via APIs. He previously scaled IMO Messenger to over 200 million users, handling up to a million new users daily with cheaper self-built infrastructure. A Northwestern CS grad and Bulgarian programming veteran, he learned that distributed systems succeed at the margins. This conversation is crucial because AI inference—execution—is the silent battleground of the AI economy, yet few discuss the compute layer, which now limits access outside big labs. Nikola applies his hyperscale infrastructure knowledge to compute-intensive workloads. For those in AI, understanding this layer is vital. 

Topics Covered 

  • From Sofia to scale — How Nikola progressed from Bulgarian programming contests to scaling IMO Messenger's backend to 200M+ monthly users, teaching him about doing more with less. 
  • Why the bet is on open source — The strategic and structural case for hosting open-source models, why a startup like DeepInfra had no realistic alternative, and why the US has an open-source gap worth worrying about 
  • The supply crunch nobody is pricing in — Demand for AI compute has roughly 4x'd in the last four months, driven mostly by coding agents finally crossing the line from "useful sometimes" to "useful always." Nikola explains why this is a GPU shortage now, not a demand problem. 
  • How Nikola actually thinks about optimization — Cache token pricing is the key cost lever most teams overlook, especially on agentic workloads. We discuss why inference hardware is diverging from training hardware and Nikola's habit of the build-then-measure loop since IMO days. 

Key Insights 

  1. AI's bottleneck is supply, not demand. Coding agents alone — Claude Code, Cursor, the rest — have driven token usage up roughly 4x in the last four months. There aren't enough GPUs to go around, and the big labs are eating most of the available capacity. Smaller buyers are getting priced out of the market entirely. 
  2. Cache token pricing is the cost lever most teams overlook. Agentic workflows reuse massive amounts of context across tool-calling loops. Pricing — and architecting around — cached tokens separately is what separates teams that can afford to scale agents from those that can't. 
  3. Open source vs. closed source won't end 100-to-zero. Nikola thinks we land at roughly 80/20 — the question is just which side gets the 80. His reasoning: information leaks, engineers move between labs, and the Anthropic GitHub commit that exposed model details earlier this year is exactly the kind of thing that's going to keep happening. You can't really isolate a model from the outside world for long.  

Links 

  • DeepInfra: https://deepinfra.com/ 
  • Nikola Borisov on LinkedIn: https://www.linkedin.com/in/nikola-borisov 
  • Future Ventures Corp: https://ca.linkedin.com/company/future-ventures-corp 
  • Maxim Atanassov on LinkedIn: https://ca.linkedin.com/in/maxim-atanassov 

About the Guest 

Nikola Borisov is the CEO and co-founder of DeepInfra, an inference-only AI cloud that hosts open-source models like DeepSeek, Llama, and GPT-OSS for developers and enterprises. Before DeepInfra, he led backend infrastructure at IMO Messenger, scaling the platform past 200 million monthly users on infrastructure his team built and ran themselves — at a fraction of what comparable cloud-native competitors were spending. He's a computer science graduate of Northwestern University with over a decade of experience building distributed systems at scale.