<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>评测 on Jiaqi</title>
    <link>https://blog.jiaqiguo.xyz/tags/%E8%AF%84%E6%B5%8B/</link>
    <description>Recent content in 评测 on Jiaqi</description>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Thu, 24 Sep 2026 17:00:00 +0800</lastBuildDate>
    <atom:link href="https://blog.jiaqiguo.xyz/tags/%E8%AF%84%E6%B5%8B/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>量语音助手的三十多把尺子：Qwen-Audio-3.1-Realtime 报告里的每个测试集，测什么、题长什么样、分怎么算</title>
      <link>https://blog.jiaqiguo.xyz/posts/qwen-audio-3-1-realtime-benchmarks/</link>
      <pubDate>Thu, 24 Sep 2026 17:00:00 +0800</pubDate>
      <guid>https://blog.jiaqiguo.xyz/posts/qwen-audio-3-1-realtime-benchmarks/</guid>
      <description>一份技术报告的分数表，真正读得懂的人不多：OpenAudioBench 的 Overall 是怎么平均出来的、Full-Duplex-Bench 的接管率在什么窗口里数、τ-Voice 的任务成功为什么要看数据库、VoiceChat 的 Spoken 和 Rubric 为什么一升一降。这篇把 Qwen-Audio-3.1-Realtime 报告用到的三十多个测试集逐个拆开：每个是什么、一道题长什么样、分数怎么算、报告里的数字该怎么读。</description>
    </item>
  </channel>
</rss>
