<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>蒸馏 on Jiaqi</title>
    <link>https://blog.jiaqiguo.xyz/tags/%E8%92%B8%E9%A6%8F/</link>
    <description>Recent content in 蒸馏 on Jiaqi</description>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Fri, 04 Sep 2026 00:50:00 +0800</lastBuildDate>
    <atom:link href="https://blog.jiaqiguo.xyz/tags/%E8%92%B8%E9%A6%8F/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>看过答案的自己教没看过答案的自己：深入浅出 OPSD 在线自蒸馏</title>
      <link>https://blog.jiaqiguo.xyz/posts/opsd-self-distillation/</link>
      <pubDate>Fri, 04 Sep 2026 00:50:00 +0800</pubDate>
      <guid>https://blog.jiaqiguo.xyz/posts/opsd-self-distillation/</guid>
      <description>做错题后翻答案复盘，是最常见的学习方式；OPSD 把它搬进 LLM：同一个模型，看过参考解的当老师、没看过的当学生，老师在学生自己写的每个 token 上逐字批注。不需要外部大模型，不需要奖励函数，1024 token 的短采样追平 8 × 16k 的 GRPO。这篇拆开它的目标函数、两个稳定器和策略梯度视角，带数字算一遍裁剪到底裁掉了什么，也聊聊后来者的质疑。</description>
    </item>
  </channel>
</rss>
