<?xml version="1.0" encoding="UTF-8" ?>
<rss version="2.0">
    <channel>
      <title>krsna.space</title>
      <link>https://www.krsna.space</link>
      <description>Last 10 notes on krsna.space</description>
      <generator>Quartz -- quartz.jzhao.xyz</generator>
      <item>
    <title>projects</title>
    <link>https://www.krsna.space/projects/</link>
    <guid>https://www.krsna.space/projects/</guid>
    <description><![CDATA[ things i’ve built. write-ups and notes for each are listed below. ]]></description>
    <pubDate>Mon, 28 Sep 2026 13:23:11 GMT</pubDate>
  </item><item>
    <title>AI-ML</title>
    <link>https://www.krsna.space/Notes/AI-ML/AI-ML</link>
    <guid>https://www.krsna.space/Notes/AI-ML/AI-ML</guid>
    <description><![CDATA[ byte latent transformers differential transformers large concept models Residual networks layer normalization the approximation of gelu nanogpt explaination MOE scaling laws titans learning to memorize at test time pytorch adam optimizer DINO v3 object centric learning with slot attention incontext ... ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item><item>
    <title>GradMem</title>
    <link>https://www.krsna.space/Notes/AI-ML/GradMem</link>
    <guid>https://www.krsna.space/Notes/AI-ML/GradMem</guid>
    <description><![CDATA[ arxiv.org/abs/2603.13875 titans learning to memorize at test time flashbacks what this paper introduces/does they want to use grad steps to use and update memory priors over an LLM to make it efficcient than KV cache over long horizion. ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item><item>
    <title>RL&#039;s razor</title>
    <link>https://www.krsna.space/Notes/AI-ML/RL's-razor</link>
    <guid>https://www.krsna.space/Notes/AI-ML/RL's-razor</guid>
    <description><![CDATA[ paper explaining why RL works over SFT or how in general works. ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item><item>
    <title>SPD via capability selective subspace projection</title>
    <link>https://www.krsna.space/Notes/AI-ML/SPD-via-capability-selective-subspace-projection</link>
    <guid>https://www.krsna.space/Notes/AI-ML/SPD-via-capability-selective-subspace-projection</guid>
    <description><![CDATA[ arxiv.org/abs/2605.22675 ok so the problem this paper attacks : self distillation where the student is the same model as the teacher should be free improvement right? except it isn’t. ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item><item>
    <title>deep-seek v3.2 sparse attention</title>
    <link>https://www.krsna.space/Notes/AI-ML/deep-seek-v3.2-sparse-attention</link>
    <guid>https://www.krsna.space/Notes/AI-ML/deep-seek-v3.2-sparse-attention</guid>
    <description><![CDATA[ github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/main/DeepSeek_V3_2.pdf instead of doing attention for all tokens in the sequence they do it for selected tokens. ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item><item>
    <title>genie</title>
    <link>https://www.krsna.space/Notes/AI-ML/genie</link>
    <guid>https://www.krsna.space/Notes/AI-ML/genie</guid>
    <description><![CDATA[ arxiv.org/abs/2402.15391v1 world envs… i love google papers they follow no conventions whatsoever lmao lmao the build up alright how do they do it how does it work it’s trained on video only data. ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item><item>
    <title>locating and editing factual associations in GPT</title>
    <link>https://www.krsna.space/Notes/AI-ML/locating-and-editing-factual-associations-in-GPT</link>
    <guid>https://www.krsna.space/Notes/AI-ML/locating-and-editing-factual-associations-in-GPT</guid>
    <description><![CDATA[ arxiv.org/abs/2202.05262 wow nice everything is opensource. which makes sense given it’s an interpretability adjacent paper. ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item><item>
    <title>memory R1</title>
    <link>https://www.krsna.space/Notes/AI-ML/memory-R1</link>
    <guid>https://www.krsna.space/Notes/AI-ML/memory-R1</guid>
    <description><![CDATA[ arxiv.org/abs/2508.19828 train models to index and organize data and retrieve to do query etc. ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item><item>
    <title>weight-sparse transformers have interpretable circuits</title>
    <link>https://www.krsna.space/Notes/AI-ML/weight-sparse-transformers-have-interpretable-circuits</link>
    <guid>https://www.krsna.space/Notes/AI-ML/weight-sparse-transformers-have-interpretable-circuits</guid>
    <description><![CDATA[ openai.com/index/understanding-neural-networks-through-sparse-circuits/ this paper is along for interpretibility but instead of the anthropic approach they are trying, to instead start with sparse models and prune the weights until they can nail specific functions over activations. ]]></description>
    <pubDate>Mon, 28 Sep 2026 12:57:28 GMT</pubDate>
  </item>
    </channel>
  </rss>