<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>强化学习 on Yiwen Cai</title><link>https://yiwen-cai.github.io/tags/%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0/</link><description>Recent content in 强化学习 on Yiwen Cai</description><generator>Hugo -- gohugo.io</generator><language>zh-cn</language><managingEditor>caiyiwen.cs@foxmail.com (Yiwen Cai)</managingEditor><webMaster>caiyiwen.cs@foxmail.com (Yiwen Cai)</webMaster><copyright>© 2026 Yiwen Cai</copyright><lastBuildDate>Tue, 04 Aug 2026 15:52:37 +0800</lastBuildDate><atom:link href="https://yiwen-cai.github.io/tags/%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0/index.xml" rel="self" type="application/rss+xml"/><item><title>RL Meets LLMs：大语言模型全生命周期强化学习综述</title><link>https://yiwen-cai.github.io/notes/papers/rl-meets-llms-survey/</link><pubDate>Tue, 04 Aug 2026 00:00:00 +0000</pubDate><author>caiyiwen.cs@foxmail.com (Yiwen Cai)</author><guid>https://yiwen-cai.github.io/notes/papers/rl-meets-llms-survey/</guid><description>梳理 RL 如何贯穿 LLM 全生命周期——预训练/Mid-training、RLHF 对齐微调、RLVR 强化推理三条主线；从策略梯度、PPO 到 GRPO 的算法演进，以及 RLVR 是否真正扩展推理能力、熵坍缩与性能上限等争议。</description></item></channel></rss>