<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Llm on CloudWizz</title><link>/tags/llm/</link><description>Recent content in Llm on CloudWizz</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 14 May 2026 02:12:00 +0400</lastBuildDate><atom:link href="/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>llm-d</title><link>/blog/llm-d/</link><pubDate>Thu, 14 May 2026 02:12:00 +0400</pubDate><guid>/blog/llm-d/</guid><description>&lt;p&gt;I&amp;rsquo;ve been watching the LLM inference space settle since KubeCon EU in March. Most discussions center on model quality. For platform teams the bottleneck is operational - not algorithmic. Getting low latency at scale means solving scheduling, caching and capacity problems that Kubernetes wasn&amp;rsquo;t designed for.&lt;/p&gt;</description></item></channel></rss>