staka – ページ 4 – arXiv最新論文の紹介

Diffusion Language Models are Super Data Learners

Diffusion Language Models are Super Data Learners [61.7]
ユニークなデータが限られている場合、拡散言語モデル(DLM)は、よりエポックなトレーニングによって、常に自己回帰モデル(AR)を上回ります。本研究の目的は,(1) 任意の次数モデリング,(2) 反復的双方向 denoising からの超高次計算,(3) モンテカルロ増分という3つの複合的要因に起因する。
論文参考訳（メタデータ） (Wed, 05 Nov 2025 08:17:42 GMT)
「The main empirical finding is a Crossover: when total training tokens are fixed but the number of unique tokens is limited, DLMs consistently surpass equally sized AR counterparts. This crossover is not an isolated artifact—it systematically shifts with core factors.　With more unique data, it shifts later; with higher data quality, it shifts later; with larger models, the crossover arrives earlier; and it persists across dense and sparse (MoE) architectures (Figures 2, 3, 4). Under compute-bound settings with abundant unique data, AR recovers its edge by fitting the data more rapidly; but in data-bound regimes, which is our focus and, increasingly, the practical reality, DLM is the final winner.」との主張。Diffusion Beats Autoregressive in Data-Constrained Settings – arXiv最新論文の紹介の主張とも整合的であるように思う。
プロジェクトサイトはDiffusion Language Models are Super Data Learners、リポジトリはGitHub – JinjieNi/dlms-are-super-data-learners: The official github repo for “Diffusion Language Models are Super Data Learners”.

同著者の下記論文も興味深い。

Training Optimal Large Diffusion Language Models [61.7]
拡散言語モデル(DLM)の最初の体系的スケーリング法則であるQuokkaを紹介する。この結果が、DLMのトレーニングにおける短期的な実践的なガイダンスと、AIコミュニティ全体の長期的なインスピレーションをもたらすことを期待しています。
論文参考訳（メタデータ） (Wed, 05 Nov 2025 08:32:08 GMT)
リポジトリはGitHub – JinjieNi/Quokka: The official github repo for “Training Optimal Large Diffusion Language Models”, the first-ever large-scale diffusion language models scaling law..

RoboOmni: Proactive Robot Manipulation in Omni-modal Context

RoboOmni: Proactive Robot Manipulation in Omni-modal Context [165.1]
我々は,音声対話や環境音,視覚的手がかりから意図を導出する,クロスモーダルな文脈指示を導入する。目的認識,インタラクション確認,アクション実行を統一する,エンドツーエンドのOmni-Modal LLMに基づくフレームワークであるRoboOmniを提案する。シミュレーションと実世界の設定の実験では、Robo OmniはテキストベースとASRベースのベースラインを越え、成功率、推論速度、意図認識、積極的に支援している。
論文参考訳（メタデータ） (Mon, 27 Oct 2025 18:49:03 GMT)
「There arises a key research question: Can a robot integrate cross-modal context, including speech, environmental audio, and visual observations, to proactively infer and verify user intent?」という疑問に対してのマルチモーダルモデル「we propose RoboOmni, an end-to-end omni-modal framework for manipulation that closes the loop of intent recognition, interaction confirmation, and action execution. Unlike prior approaches, RoboOmni supports direct speech interaction without ASR, infers latent commands by fusing human speech, environmental audio, and vision through spatiotemporal modeling, and verifies intent via interaction.」
プロジェクトサイトはRoboOmni: Proactive Robot Manipulation in Omni-modal Context

Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts

Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts [113.1]
オフ・ポリティクス強化学習(RL)における重要サンプリング重み付けを最適化する新しいルータ認識手法を提案する。具体的には、ルータロジットによって誘導される再スケーリング戦略を設計し、勾配のばらつきを効果的に低減し、トレーニングのばらつきを軽減する。実験により, 本手法は収束安定性とMoEモデルの最終的な性能の両方を著しく改善することが示された。
論文参考訳（メタデータ） (Mon, 27 Oct 2025 05:47:48 GMT)
MoEに対する強化学習のための「Router-Shift Policy Optimization (RSPO), an RL algorithm specifically designed for MoE architectures to achieve stable and efficient training.」を提案。

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks [33.7]
下流認識タスクを強化するための新しい合成データ生成フレームワークであるDream4Driveを紹介する。 Dream4Driveは入力ビデオを複数の3D対応誘導マップに分解し、これらの誘導マップに3Dアセットをレンダリングする。駆動世界モデルは、下流の知覚モデルをトレーニングするために使用できる編集されたマルチビュービデオを作成するために微調整される。
論文参考訳（メタデータ） (Fri, 24 Oct 2025 10:10:43 GMT)
「We propose Dream4Drive, a 3D-aware synthetic data generation framework that edits the video with dense guidance maps, producing synthetic data with diverse appearances and geometric consistency.」とデータ合成フレームワークの提案。
プロジェクトサイトはRethinking Driving World Model as Synthetic Data Generator for Perception Tasks

Social Simulations with Large Language Model Risk Utopian Illusion

Social Simulations with Large Language Model Risk Utopian Illusion [61.4]
社会シミュレーションにおける大規模言語モデルの行動分析のための体系的枠組みを提案する。本手法は,チャットルーム型会話を通してマルチエージェントインタラクションをシミュレートし,5つの言語的側面にわたって解析する。以上の結果から,LSMは真の人間の行動を忠実に再現するのではなく,過度に理想化されたバージョンを反映していることが明らかとなった。
論文参考訳（メタデータ） (Fri, 24 Oct 2025 06:08:41 GMT)
様々なところで試されているLLMを用いた社会シミュレーションに関する報告、「Our findings reveal that LLMs do not faithfully reproduce genuine human behavior but instead reflect overly idealized versions of it, shaped by the social desirabil- ity bias. In particular, LLMs show social role bias, primacy effect, and positivity bias, resulting in “Utopian” societies that lack the complexity and variability of real human interactions.」と否定的見解。

Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark

Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark [124.0]
我々は、ビデオモデルがゼロショット推論器として機能する準備が整っているかどうかを実証研究する。私たちは、人気の高いVeo-3に注力しています。我々は,空間的,幾何学的,物理的,時間的,具体的論理を含む12次元にわたる推論行動を評価する。
論文参考訳（メタデータ） (Thu, 30 Oct 2025 17:59:55 GMT)
「Video models are zero-shot learners and reasoners – arXiv最新論文の紹介」という主張もあるが、異なるチームによる論文。「Our findings reveal that while current video models demonstrate promising reasoning patterns on short-horizon spatial coherence, fine-grained grounding, and locally consistent dynamics, they remain limited in long-horizon causal reasoning, strict geometric constraints, and abstract logic. Overall, they are not yet reliable as standalone zero-shot reasoners, but exhibit encouraging signs as complementary visual engines alongside dedicated reasoning models.」とのことで可能性を感じる結果ではある。
プロジェクトサイトはAre Video Models Ready as Zero-Shot Reasoners?

DeepAgent: A General Reasoning Agent with Scalable Toolsets

DeepAgent: A General Reasoning Agent with Scalable Toolsets [111.6]
DeepAgentは、自律的な思考、ツール発見、アクション実行を実行するエンドツーエンドのディープ推論エージェントである。長期にわたる相互作用の課題に対処するために,過去の相互作用を構造化エピソード,動作,ツール記憶に圧縮する自律的メモリ折り畳み機構を導入する。 LLMシミュレートされたAPIを活用し、ツール呼び出しトークンにきめ細かいクレジットを割り当てるツールコールアドバンテージ属性を適用した、エンドツーエンドの強化学習戦略であるToolPOを開発した。
論文参考訳（メタデータ） (Fri, 24 Oct 2025 16:24:01 GMT)
ツール利用等も可能になるエージェントフレームワークの紹介。QwQ-32Bをバックボーンとして有効性を検証している。
リポジトリはGitHub – RUC-NLPIR/DeepAgent: 🛠️ DeepAgent: A General Reasoning Agent with Scalable Toolsets

ImpossibleBench: Measuring LLMs’ Propensity of Exploiting Test Cases

ImpossibleBench: Measuring LLMs’ Propensity of Exploiting Test Cases [58.4]
タスク完了のための「ショートカット」は、大規模言語モデルの信頼性評価と展開に重大なリスクをもたらす。我々は,LLMエージェントがテストケースを利用するための正当性を測定するベンチマークフレームワークであるImpossibleBenchを紹介する。実践的なフレームワークとして、ImpossibleBenchは単なる評価ではなく、汎用的なツールである。
論文参考訳（メタデータ） (Thu, 23 Oct 2025 06:58:32 GMT)
「we introduce ImpossibleBench, a benchmark framework that systematically measures LLM agents’ propensity to exploit test cases.」と不正行為を測るためのベンチマーク。「frontier models frequently cheat when faced with these impossible tasks, and stronger models generally exhibit higher cheating rates.」という指摘が興味深いし感覚にも合う・・・
リポジトリはGitHub – safety-research/impossiblebench

ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows

ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows [109.3]
CS-54k(CS-54k)は、コンピュータ科学におけるQ&Aペアの高品質なコーパスである。 CS-4kは、科学研究を支援するAIの能力を評価するためのベンチマークである。 CS-50kは大規模なトレーニングデータセットである。
論文参考訳（メタデータ） (Thu, 23 Oct 2025 07:07:35 GMT)
「We introduce CS-4k, the first benchmark that systematically evaluates the end-to-end research workflow in computer science through open-ended scientific question answering, offering a rigorous yardstick to assess LLMs’ ability to assist scientific research.」というベンチマーク。また、これらデータを用いたポストトレーニングの有効性を主張。
リポジトリはGitHub – wph6/ResearchGPT: Official repo for ReseachGPT

Human-AI Interactions: Cognitive, Behavioral, and Emotional Impacts

Human-AI Interactions: Cognitive, Behavioral, and Emotional Impacts [0.0]
過度な信頼感、認知的オフロード、社会的および感情的な操作、および人間の代理店の曖昧な劣化と判断の潜在的なリスクが強調される。観察によると、AIは記憶、創造性、エンゲージメントを大幅に向上させることができるが、批判的思考の減少、スキルの侵食、不安の増加といったリスクももたらしている。本稿は、人間中心の新たなリスクと利益のバランスをとるための、縦断的研究と評価フレームワークのギャップを浮き彫りにして、責任とコンテキストを意識したAI設計の必要性を明らかにすることを目的としている。
論文参考訳（メタデータ） (Mon, 20 Oct 2025 17:06:46 GMT)
人間とAIのかかわりに関してのサーベイ。リスク面で注意すべきかもしれない事例が多く紹介されている。

月	火	水	木	金	土	日
					1	2
3	4	5	6	7	8	9
10	11	12	13	14	15	16
17	18	19	20	21	22	23
24	25	26	27	28	29	30