arXiv最新論文の紹介

Frontier LLMs Still Struggle with Simple Reasoning Tasks

Frontier LLMs Still Struggle with Simple Reasoning Tasks [53.5]
この研究は、フロンティア言語モデルの性能を、幅広い「容易」推論問題に対して研究する。計算,一階述語論理,証明木,旅行計画など,手続き的に生成された単純な推論タスクのスイートを作成します。最先端の思考モデルでさえ、このような問題や同様の理由で一貫して失敗することを示します。
論文参考訳（メタデータ） (Wed, 09 Jul 2025 22:22:49 GMT)
「By extending previous work in the literature, we create a suite of procedurally generated simple reasoning tasks, including counting, first-order logic, proof trees, and travel planning, with changeable parameters (such as document length. or the number of variables in a math problem) that can arbitrarily increase the amount of computation required to produce the answer while preserving the fundamental difficulty. While previous work showed that traditional, non-thinking models can be made to fail on such problems, we demonstrate that even state-of-the-art thinking models consistently fail on such problems and for similar reasons (e g , statistical shortcuts, errors in intermediate steps, and difficulties in processing long contexts).」と簡単だがLLM/LRMによって解きにくいタスクを作成。
「Similarly to other recent works, our results suggest that LLMs mimic training data rather than performing true reasoning, making it relatively easy to find out-of-distribution problems where the models fail, and this problem is also present at the newest thinking models. This suggests that users remain careful when relying on the output of LLMs.」と指摘している。下記のCatAttackの時も感じたがLLM/LRMは人間の能力とはかなり異なっていることは意識したほうが良いと思う。
リポジトリはhttps://github.com/google-deepmind/unpuzzles_and_simple_reasoning/とのこと

Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models [25.1]
本稿では,問合せに依存しない逆引き金を導入することで,段階ごとの問題解決を訓練した推論モデルのロバスト性について検討する。より弱く安価なプロキシモデル上でトリガを生成する自動反復攻撃パイプラインであるCatAttackを提案する。我々の研究結果は、推論モデルにおける重大な脆弱性を浮き彫りにして、最先端モデルでさえ、微妙な敵の入力に影響を受けやすいことを明らかにした。
論文参考訳（メタデータ） (Mon, 03 Mar 2025 18:10:54 GMT)
「For example, appending, Interesting fact: cats sleep most of their lives, to any math problem leads to more than doubling the chances of a model getting the answer wrong. Our findings highlight critical vulnerabilities in reasoning models, revealing that even state-of- the-art models remain susceptible to subtle adversarial inputs, raising security and reliability concerns.」という面白い攻撃。一方で、ノイズ（無関係）な事例がRAGの改善に有効という話もあり動作は本当に謎。
リポジトリはcollinear-ai/cat-attack-adversarial-triggers · Datasets at Hugging Face

The Power of Noise: Redefining Retrieval for RAG Systems [19.4]
Retrieval-Augmented Generation (RAG) は、大規模言語モデルの事前学習知識を超えて拡張する方法として登場した。我々は、RAGソリューションが取得すべきパスIRシステムの種類に焦点を当てる。
論文参考訳（メタデータ） (Wed, 1 May 2024 08:15:07 GMT)
「Finally, and even more surprisingly, random, noisy documents are actually helpful in increasing the accuracy of these systems when correctly positioned within a prompt.」と無関係な事例が有効なのは興味深い

MemOS: A Memory OS for AI System, MIRIX: Multi-Agent Memory System for LLM-Based Agents

RAGでは厳しい問題を扱うためのMemory関連の研究がとても盛ん。

MemOS: A Memory OS for AI System [115.3]
大規模言語モデル(LLM)は、人工知能(AGI)にとって不可欠な基盤となっている。既存のモデルは、主に静的パラメータと短命なコンテキスト状態に依存しており、ユーザの好みを追跡したり、長い期間にわたって知識を更新する能力を制限する。 MemOSはメモリを管理可能なシステムリソースとして扱うメモリオペレーティングシステムである。
論文参考訳（メタデータ） (Fri, 04 Jul 2025 17:21:46 GMT)
MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models – arXiv最新論文の紹介からのアップデート、AgenticなアプローチのLLM用メモリ。時系列性など通常のRAGでは簡単ではない部分の性能向上が大きい。（が、「To ensure architectural parity, all methods are implemented over the same LLM backbone (GPT-4o-mini)」とベースモデルがGPT-4o miniで良いのかは若干謎ではある）
リポジトリはGitHub – MemTensor/MemOS: MemOS (Preview) | Intelligence Begins with Memory

MIRIX: Multi-Agent Memory System for LLM-Based Agents [7.1]
MIRIXは言語モデルのためのモジュール型マルチエージェントメモリシステムである。 MIRIXは、リッチな視覚的およびマルチモーダル体験を受け入れるためにテキストを超越する。 MIRIXはメモリ拡張LDMエージェントの新たなパフォーマンス標準を設定している。
論文参考訳（メタデータ） (Thu, 10 Jul 2025 17:40:11 GMT)
こちらもAgenticなアプローチのメモリ管理フレームワーク。ベースモデルが異なるためMemOSと直接比較が困難だが、他システムと比べ高い性能を主張。
リポジトリはGitHub – Mirix-AI/MIRIX: Mirix is a multi-agent personal assistant designed to track on-screen activities and answer user questions intelligently. By capturing real-time visual data and consolidating it into structured memories, Mirix transforms raw inputs into a rich knowledge base that adapts to your digital experiences.

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions [19.5]
メモリ機構を持つエージェントをメモリエージェントと呼ぶ。本稿では,メモリエージェントに不可欠な4つのコア能力,すなわち,正確な検索,テスト時間学習,長距離理解,コンフリクト解決の4つを同定する。既存のデータセットは、限られたコンテキスト長に依存するか、書籍ベースのQAのような静的で長いコンテキスト設定用に調整されている。既存のベンチマークでは4つの能力をすべてカバーしていないため、メモリエージェント用に特別に設計された新しいベンチマークであるMemoryAgentBenchを紹介します。
論文参考訳（メタデータ） (Mon, 07 Jul 2025 17:59:54 GMT)
こちらはMemoryを持つエージェントのためのベンチマークの提案
「we identify four core competencies essential for memory agents: accurate retrieval, test-time learning, long-range understanding, and conflict resolution.」とのこと。
結果にある「While Mem0 has demonstrated relatively strong performance on conversational tasks such as LOCOMO—where information density is comparatively low—it tends to perform poorly on benchmarks containing dense informational content, including RULER and ∞-Bench. For tasks emphasizing Time-to-Live (TTL) and Least Recently Used (LRU) retrieval, these limitations are often even more pronounced.」という指摘は興味深く、ドメインを選ばない汎用的な構造を作るのは大変そうという印象。
リポジトリはai-hyz/MemoryAgentBench · Datasets at Hugging Face、GitHub – HUST-AI-HYZ/MemoryAgentBench: Open source code for Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

FlexOlmo: Open Language Models for Flexible Data Use

FlexOlmo: Open Language Models for Flexible Data Use [184.9]
我々は、データ共有なしで分散トレーニングをサポートする新しい言語モデル(LM)であるFlexOlmoを紹介します。 FlexOlmoはエキスパートの混成アーキテクチャを採用しており、各専門家はクローズドデータセットで独立して訓練される。我々は、公開データで訓練された一般専門家と、他のデータ所有者から独立した訓練を受けた専門家とを効果的に組み合わせることができることを示す。
論文参考訳（メタデータ） (Wed, 09 Jul 2025 16:54:21 GMT)
「Standard MoEs train all experts and the router jointly on all data. In contrast, FLEXOLMO trains experts independently by teaching them to coordinate (§3.3.1) and merges them at inference using a domain-informed router (§3.3.2).」と連合学習やMoEと聞いて思い浮かべるが現実的には難しいそれぞれの場所で構築されたAIが統合的に動作するフレームワークの提案と効果検証。
「Organizations in regulated industries require LMs that can leverage their closed datasets while maintaining strict data privacy and access controls. Healthcare institutions, financial firms, and other entities possess valuable domain-specific data but cannot share it externally due to HIPAA, GDPR [14, 15], data sovereignty laws [16], and intellectual property (IP) protections. 　These organizations need training paradigms that enable AI improvement on their sensitive data while ensuring such sensitive data never leaves certain environments and can be removed from the model after training, e g , when data usage rights expire. In such settings, modular training approaches, where individual experts are trained independently and asynchronously on locally maintained data, are essential.」はまさにその通りで非常に有用な技術に思える。
プロジェクトサイトはIntroducing FlexOlmo: a new paradigm for language model training and data collaboration | Ai2、リポジトリはGitHub – allenai/FlexOlmo: Code and training scripts for FlexOlmo

Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers

Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers [31.5]
LimitGenは、初期のフィードバックをサポートし、人間のピアレビューを補完するLLMの能力を評価するための最初のベンチマークである。提案手法は, LLMシステムによる研究論文の限界を生じさせる能力を高め, より具体的で建設的なフィードバックを提供する。
論文参考訳（メタデータ） (Thu, 03 Jul 2025 15:04:38 GMT)
「We propose LIMITGEN, a comprehensive bench- mark specifically designed to assess the ability of models to identify and address limitations in scientific research, with a reliable and systematic evaluation framework.」というベンチマークの提案と検証。「Even the best-performing LLM, GPT-4o, can only identify about half of the limitations that humans consider very obvious. Although MARG lever- ages multi-agent collaboration and generates more comments, successfully identifying more limita- tions, the feedback it provides still lacks specificity, which is reflected in the fine-grained scores.」とのこと。MARGはマルチエージェントフレームワーク。
リポジトリはGitHub – yale-nlp/LimitGen: Data and Code for ACL 2025 Paper “Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers”

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact [31.6]
本稿では,人工知能,認知神経科学,心理学,生成モデル,エージェントベースシステムの学際的合成について述べる。我々は汎用知能のアーキテクチャと認知の基礎を分析し、モジュラー推論、永続記憶、マルチエージェント協調の役割を強調した。我々は、人工知能への道の鍵となる科学的、技術的、倫理的課題を特定します。
論文参考訳（メタデータ） (Tue, 01 Jul 2025 16:52:25 GMT)
AGIを目指すうえでの整理「Several challenges remains, such as the need for grounded world models, dynamic memory, causal reasoning, robust handling of aleatory and epistemic uncertainty, developing perception of emotional and social contexts and collective agent architectures. Significant advancements have been made, such as Large Concept Models, Large Reasoning Models and Mixture of Experts, which improve LLM performance beyond next-token prediction by incorporating biologically inspired behaviors into output generation.」と指摘。
MoEなど技術的なとらえ方に違和感がなくはないが興味深い整理

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools [15.9]
ファンデーションモデル(FM)は、科学的発見のためにスケーラブルで汎用的でマルチモーダルなAIシステムを実現する。この調査は、この成長分野をサポートする基盤モデル、エージェントシステム、データセット、計算ツールの包括的概要を提供する。
論文参考訳（メタデータ） (Wed, 25 Jun 2025 18:10:30 GMT)

The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure

The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure [25.0]
生成のための暗黙的なタスク解決–>翻訳パイプラインの存在を実証する。 108言語対にわたる単語翻訳タスクに対して,この仮説を検証した。全体的な失敗のかなりの部分は、翻訳失敗に起因していることが分かりました。
論文参考訳（メタデータ） (Sat, 28 Jun 2025 02:09:21 GMT)
「We find that a significant portion of overall failures indeed stems from translation failure, or the model’s inability to translate correctly solved intermediate concepts into the target language. This is especially true for low-resource target languages.」という指摘
動作自体はBeyond English-Centric LLMs: What Language Do Multilingual Language Models Think in? – arXiv最新論文の紹介からもそうなんだろうと思いつつ、中間言語は学習の中心になった言語に影響されているんだろうなと思うとそれでよいのかという気がしなくはない。

FineWeb2: One Pipeline to Scale Them All — Adapting Pre-Training Data Processing to Every Language

FineWeb2: One Pipeline to Scale Them All — Adapting Pre-Training Data Processing to Every Language [48.8]
我々は、FineWebをベースにした、新しいトレーニング済みデータセットキュレーションパイプラインを導入する。我々のパイプラインは、以前のデータセットよりもパフォーマンスの高いモデルを生成する非英語コーパスを作成するために使用できることを示す。パイプラインを約100のCommon Crawlスナップショットを使用して1000以上の言語に拡張し、新たに20テラバイト(50億ドキュメント)のマルチリンガルデータセットであるFinWeb2を生成しました。
論文参考訳（メタデータ） (Thu, 26 Jun 2025 01:01:47 GMT)
大規模、マルチリンガル、高品質なデータセットの提案。重複データへの対応やフィルタリングによって他のデータセットよりも効率的な学習が可能とのこと
リポジトリはGitHub – huggingface/fineweb-2、データセットはHuggingFaceFW/fineweb-2 · Datasets at Hugging Face

Embodied AI Agents: Modeling the World

Embodied AI Agents: Modeling the World [165.0]
本稿では,視覚的,仮想的,物理的形態を具現化したAIエージェントの研究について述べる。我々は,世界モデルの開発が,具体的AIエージェントの推論と計画の中心であることを提案する。また,より優れた人間とエージェントのコラボレーションを実現するために,ユーザのメンタルワールドモデルを学ぶことを提案する。
論文参考訳（メタデータ） (Fri, 27 Jun 2025 16:05:34 GMT)
「We propose that the development of world models is central to reasoning and planning of embodied AI agents, allowing these agents to understand and predict their environment, to understand user intentions and social contexts, thereby enhancing their ability to perform complex tasks autonomously. World modeling encompasses the integration of multimodal perception, planning through reasoning for action and control, and memory to create a comprehensive understanding of the physical world.」という整理

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements [87.6]
科学的進歩への重要な能力は、既存の作品を再現する能力である。アクティブな研究領域においてAIエージェントが結果を再現する能力を評価するために,自動LLM高速化ベンチマークを導入する。最近のLSMとSoTAの足場を組み合わせると、ベンチマークですでに知られているイノベーションを再実装するのに苦労していることが分かりました。
論文参考訳（メタデータ） (Fri, 27 Jun 2025 17:44:32 GMT)
「We find that recent reasoning LLMs combined with SoTA scaffolds struggle to reimplement already-known innovations in our benchmark, even when given detailed hints.」というやや意外な結果。
リポジトリはGitHub – facebookresearch/llm-speedrunner: The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in language modeling.

2025年8月
月	火	水	木	金	土	日
				1	2	3
4	5	6	7	8	9	10
11	12	13	14	15	16	17
18	19	20	21	22	23	24
25	26	27	28	29	30	31