Self-Questioning Language Models – arXiv最新論文の紹介

Self-Questioning Language Models [51.8]
本稿では,提案者がトピックを与えられ,解答者に対する質問を生成する非対称なセルフプレイフレームワークを提案する。提案者と解答者はともに強化学習を通じて訓練される。 3桁の乗算、OMEGAベンチマークの代数問題、Codeforcesのプログラミング問題である。
論文参考訳（メタデータ） (Tue, 05 Aug 2025 17:51:33 GMT)
「Our method leverages the intrinsic capabilities of large language models by casting them in dual roles of proposer and solver within an asymmetric self-play setup. By rewarding the generation of problems that are neither too easy nor too difficult, and by reinforcing answers via internal agreement or external verification, we demonstrate that models can meaningfully improve their reasoning skills through interaction with self-generated content alone.」というフレームワークの提案。R-Zero: Self-Evolving Reasoning LLM from Zero Data – arXiv最新論文の紹介にも近いなーと思う。
プロジェクトサイトはSelf-Questioning Language Models

コメントを残す

コメントを残す コメントをキャンセル