رادار AI والتقنية

أحدث اتجاهات AI والتقنية

عناوين وم excerpts وروابط مصادر محدّثة يومياً من feeds موثوقة.

Browse by Category

Aggregated legally from RSS and public APIs. Full articles remain at the original publisher.

60 articles

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
نماذج الذكاء الاصطناعي arXiv cs.AI

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents

arXiv:2606.13544v1 Announce Type: cross Abstract: Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations. We propose ModeratorLM, a role-playing voice agent that conditions turn-taking behavior…

ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages
نماذج الذكاء الاصطناعي arXiv cs.AI

ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages

arXiv:2606.13572v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios. This gap is critical in re…

One Polluted Page Is Enough: Evaluating Web Content Pollution in Generative Recommenders
نماذج الذكاء الاصطناعي arXiv cs.AI

One Polluted Page Is Enough: Evaluating Web Content Pollution in Generative Recommenders

arXiv:2606.13610v1 Announce Type: cross Abstract: Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This creates a new risk: generative recommenders may consume polluted web content, such as fake reviews and promotional pages crafted to mislead recommendation…

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
نماذج الذكاء الاصطناعي arXiv cs.AI

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

arXiv:2511.02627v4 Announce Type: replace Abstract: We introduce DecompSR, decomposed spatial reasoning, a large benchmark dataset (over 5m datapoints) and generation framework designed to analyse compositional spatial reasoning ability. The generation of DecompSR allows users to independently vary several aspects of…

DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems
نماذج الذكاء الاصطناعي arXiv cs.AI

DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems

arXiv:2601.13591v2 Announce Type: replace Abstract: Recent LLM-based data agents aim to automate data science tasks ranging from data analysis to deep learning. However, the open-ended nature of real-world data science problems, which often span multiple taxonomies and lack standard answers, poses a significant chall…

Epistemic Constitutionalism Or: how to avoid coherence bias
نماذج الذكاء الاصطناعي arXiv cs.AI

Epistemic Constitutionalism Or: how to avoid coherence bias

arXiv:2601.14295v4 Announce Type: replace Abstract: Large language models increasingly function as artificial reasoners: they evaluate arguments, assign credibility, and express confidence. Yet their belief-forming behavior is governed by implicit, uninspected epistemic policies. This paper argues for an epistemic co…

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration
نماذج الذكاء الاصطناعي arXiv cs.AI

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

arXiv:2603.21563v5 Announce Type: replace Abstract: Collaborative multi-agent large language models (LLMs) can solve complex reasoning tasks by decomposing roles, but reinforcement learning for such systems is limited by credit assignment: shared terminal rewards obscure individual contributions and can encourage fre…

LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning
نماذج الذكاء الاصطناعي arXiv cs.AI

LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning

arXiv:2604.27960v2 Announce Type: replace Abstract: Recent large language models (LLMs) have achieved impressive reasoning milestones but continue to struggle with high computational costs, logical inconsistencies, and sharp performance degradation on high-complexity problems. While neuro-symbolic methods attempt to…

Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture
نماذج الذكاء الاصطناعي arXiv cs.AI

Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture

arXiv:2606.09500v3 Announce Type: replace Abstract: As autonomous research agents and AI co-scientist systems push large language models (LLMs) from drafting toward end-to-end manuscript production, the bottleneck shifts from generation to verification. Fluent LLM output can hide fabricated citations, numbers that dr…

A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design
نماذج الذكاء الاصطناعي arXiv cs.AI

A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design

arXiv:2606.12040v2 Announce Type: replace Abstract: The design of reinforced concrete highway barriers is a safety-critical process that requires strict compliance with regulatory provisions such as the AASHTO-LRFD bridge design guidelines. Current engineering practice relies heavily on manual, iterative, and heurist…

ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs
نماذج الذكاء الاصطناعي arXiv cs.AI

ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs

arXiv:2606.12451v1 Announce Type: new Abstract: Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck. As embedding-based retrieval approaches rely on compact encoders that may under-capture specialized tool semantics, parametric tool retrieval addresses this by e…

Emergence of Hierarchical Emotion Organization in Large Language Models
نماذج الذكاء الاصطناعي arXiv cs.AI

Emergence of Hierarchical Emotion Organization in Large Language Models

arXiv:2507.10599v2 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly power conversational agents, understanding how they model users' emotional states is critical for ethical deployment. Inspired by emotion wheels, i.e., a psychological framework that argues emotions organize hierarc…

Authorship Attribution in Multilingual Machine-Generated Texts
نماذج الذكاء الاصطناعي arXiv cs.AI

Authorship Attribution in Multilingual Machine-Generated Texts

arXiv:2508.01656v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult. While early efforts in MGT detection have focused on binary classification, th…

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
نماذج الذكاء الاصطناعي arXiv cs.AI

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models

arXiv:2508.04427v2 Announce Type: replace-cross Abstract: Multimodal learning has witnessed remarkable advancements in recent years, particularly with the integration of attention-based models, leading to significant performance gains across a variety of tasks. Parallel to this progress, the demand for explainable ar…

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System
نماذج الذكاء الاصطناعي arXiv cs.AI

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System

arXiv:2606.12702v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks tend to measure correctness rather than user acceptance, aggregate performance across quer…

Examining the Usage of Generative AI Models in Student Learning Activities for Software Programming
نماذج الذكاء الاصطناعي arXiv cs.AI

Examining the Usage of Generative AI Models in Student Learning Activities for Software Programming

arXiv:2511.13271v2 Announce Type: replace-cross Abstract: The rise of Generative AI (GenAI) tools like ChatGPT has created new opportunities and challenges for computing education. Existing research has primarily focused on GenAI's ability to complete educational tasks and its impact on student performance, often ove…

PhononBench:A Large-Scale Phonon-Based Benchmark for Dynamical Stability in Crystal Generation
نماذج الذكاء الاصطناعي arXiv cs.AI

PhononBench:A Large-Scale Phonon-Based Benchmark for Dynamical Stability in Crystal Generation

arXiv:2512.21227v3 Announce Type: replace-cross Abstract: In recent years, generative artificial intelligence has made significant advances in the design of crystalline materials, giving rise to approaches based on graph neural networks, diffusion models, and large language models. Existing evaluations commonly follo…

Prefill Awareness in Large Language Models
نماذج الذكاء الاصطناعي arXiv cs.AI

Prefill Awareness in Large Language Models

arXiv:2606.12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs. If AI models can recognize and act on the fact their prior assistant messages have been inserted or edited, the…

A Tutorial on World Models and Physical AI
نماذج الذكاء الاصطناعي arXiv cs.AI

A Tutorial on World Models and Physical AI

arXiv:2606.12783v1 Announce Type: new Abstract: World modeling is emerging as a central principle for building intelligent systems capable of prediction, reasoning, and decision making. A central distinction can be drawn between explicit world models, which learn structured dynamics for rollout-based reasoning and pl…

The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements
نماذج الذكاء الاصطناعي arXiv cs.AI

The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements

arXiv:2606.12797v1 Announce Type: new Abstract: Agentic large language model systems that autonomously invoke tools, maintain persistent memory, and execute multi-step plans are increasingly deployed in public-facing domains, including government services, healthcare triage, and financial advising. We ask whether the…

CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters
نماذج الذكاء الاصطناعي arXiv cs.AI

CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters

arXiv:2601.04885v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) serve a global audience, alignment must transition from enforcing universal consensus to respecting cultural pluralism. We demonstrate that dense models, when forced to fit conflicting value distributions, suffer from \textbf{Me…

MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs
نماذج الذكاء الاصطناعي arXiv cs.AI

MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs

arXiv:2606.12809v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are trained on massive multimodal data, making data unlearning increasingly important as data owners may request the removal of specific content. In practice, these requests often arrive sequentially over time, giving rise to the…

GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models
نماذج الذكاء الاصطناعي arXiv cs.AI

GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models

arXiv:2606.12821v1 Announce Type: new Abstract: Environmental scientists spend disproportionate effort on data wrangling rather than analysis, and AI agents that automate geospatial workflows remain unvalidated: no benchmark evaluates agents operating through structured tool calling against real APIs. We introduce th…

HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation
نماذج الذكاء الاصطناعي arXiv cs.AI

HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation

arXiv:2601.19072v3 Announce Type: replace-cross Abstract: Large Language models (LLMs) have shown strong capabilities in code review automation, such as review comment generation, yet they suffer from hallucinations -- where the generated review comments are ungrounded in the actual code -- poses a significant challe…

Topical Phase Transitions in Artificial Intelligence Research: Large-Scale Evidence and an Early-Warning Signature for Emerging Topics
نماذج الذكاء الاصطناعي arXiv cs.AI

Topical Phase Transitions in Artificial Intelligence Research: Large-Scale Evidence and an Early-Warning Signature for Emerging Topics

arXiv:2606.12828v1 Announce Type: new Abstract: Do research topics in artificial intelligence grow gradually, or do they advance through abrupt, detectable jumps? Analyzing 80,814 accepted main-track papers from five premier AI conferences (ACL, CVPR, ICLR, ICML, NeurIPS) spanning 2017 to 2025, we show major AI topic…

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering
نماذج الذكاء الاصطناعي arXiv cs.AI

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

arXiv:2601.19827v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-reasoning loops meaningfully outperform static RAG, particularly in scientific domains with multi-hop reasoning, s…

(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable
نماذج الذكاء الاصطناعي arXiv cs.AI

(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable

arXiv:2606.12848v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions. We argue that the reliability of AI-assisted research depends not only on model capabilit…

DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks
نماذج الذكاء الاصطناعي arXiv cs.AI

DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks

arXiv:2606.12871v1 Announce Type: new Abstract: Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information into comprehensive responses. For SAs evaluation, prior benchmarks mainly focus on special…

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
نماذج الذكاء الاصطناعي arXiv cs.AI

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs

arXiv:2602.00462v4 Announce Type: replace-cross Abstract: Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM. Intriguingly, this mapping can be as simple as a shallow MLP transformation. To…

HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness
نماذج الذكاء الاصطناعي arXiv cs.AI

HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness

arXiv:2606.12882v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents for long-horizon tasks, yet their performance is shaped not only by model capability and environment design, but also by the harness that mediates agent--environment interaction. Existing harnesses are largely ma…

Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models
نماذج الذكاء الاصطناعي arXiv cs.AI

Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models

arXiv:2602.07106v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely unexplored despite its importance for natural human-computer interaction. A key…

Zero-source LLM Hallucination Detection with Human-like Criteria Probing
نماذج الذكاء الاصطناعي arXiv cs.AI

Zero-source LLM Hallucination Detection with Human-like Criteria Probing

arXiv:2606.12900v1 Announce Type: new Abstract: Large language models (LLMs) often hallucinate by generating factually incorrect or unfaithful content, posing significant risks to their safe use. Detecting such hallucinations is particularly challenging under the zero-source constraint, where no model internals or ex…

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
نماذج الذكاء الاصطناعي arXiv cs.AI

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings

arXiv:2602.07294v4 Announce Type: replace-cross Abstract: With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures. However, existing benchmarks often focus on isolated details, failing to reflect the complexity of pro…

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce
نماذج الذكاء الاصطناعي arXiv cs.AI

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce

arXiv:2606.12924v1 Announce Type: new Abstract: We present a modular two-agent simulation framework for evaluating conversational shopping assistant architectures. An independent buyer agent, configured with personas, missions, and patience levels, is paired with an interchangeable responder that integrates with a re…

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
نماذج الذكاء الاصطناعي arXiv cs.AI

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem

arXiv:2602.14367v2 Announce Type: replace-cross Abstract: The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. The fundamental nature of scientific evaluation needs knowledgeable grounding, co…

PRISMR: Overcoming Parse Collapse in Multimodal Listwise Ranking via Parameterized Representation Internalization
نماذج الذكاء الاصطناعي arXiv cs.AI

PRISMR: Overcoming Parse Collapse in Multimodal Listwise Ranking via Parameterized Representation Internalization

arXiv:2606.12942v1 Announce Type: new Abstract: Generative listwise ranking with Large Multimodal Models (LMMs) aims to capture global list context in a single forward pass, but its effectiveness degrades in long-context multimodal scenarios. We identify a recurring failure mode, parse collapse, where the autoreg…

FENCE: A Financial and Multimodal Jailbreak Detection Dataset
نماذج الذكاء الاصطناعي arXiv cs.AI

FENCE: A Financial and Multimodal Jailbreak Detection Dataset

arXiv:2602.18154v2 Announce Type: replace-cross Abstract: Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces. However, available resource…

Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models
نماذج الذكاء الاصطناعي arXiv cs.AI

Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models

arXiv:2606.12969v1 Announce Type: new Abstract: The power distribution network is critical to reliable electricity delivery, yet traditional inspection methods face limitations in semantic understanding, generalization, and closed-loop automation. To address these challenges, this paper proposes a Multi-Modal Agent f…

Contextual Invertible World Models: A Neuro-Symbolic Agentic Framework for Colorectal Cancer Drug Response
نماذج الذكاء الاصطناعي arXiv cs.AI

Contextual Invertible World Models: A Neuro-Symbolic Agentic Framework for Colorectal Cancer Drug Response

arXiv:2603.02274v3 Announce Type: replace-cross Abstract: Precision oncology is currently limited by the small-N, large-P paradox, where high-dimensional genomic data is abundant but pharmacological response samples are sparse. While deep learning achieves predictive accuracy, it frequently fails to provide the mecha…

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
نماذج الذكاء الاصطناعي arXiv cs.AI

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment

arXiv:2603.06652v2 Announce Type: replace-cross Abstract: Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the rig…

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation
نماذج الذكاء الاصطناعي arXiv cs.AI

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation

arXiv:2606.12983v1 Announce Type: new Abstract: Automated testbench generation has become a critical bottleneck in large language model (LLM)-driven Register Transfer Level (RTL) workflows, where large numbers of candidate designs must be verified rapidly and reliably. Existing prompt-based approaches treat testbench…

Nous: An Attempt to Extract and Inject the Cognition Behind Prediction-Market Behavior
نماذج الذكاء الاصطناعي arXiv cs.AI

Nous: An Attempt to Extract and Inject the Cognition Behind Prediction-Market Behavior

arXiv:2606.13038v1 Announce Type: new Abstract: As LLM agents proliferate in prediction markets and collective decision-making, they risk a cognitive monoculture: agents built on shared foundation models produce correlated forecasts, and recent measurement finds frontier-model errors correlated at r ~ 0.77. We ask wh…

DCD: Domain-Oriented Design for Controlled Retrieval-Augmented Generation
نماذج الذكاء الاصطناعي arXiv cs.AI

DCD: Domain-Oriented Design for Controlled Retrieval-Augmented Generation

arXiv:2604.07590v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) is widely used to ground large language models in external knowledge sources. However, when applied to heterogeneous corpora and multi-step queries, Naive RAG pipelines often degrade in quality due to flat knowledge represe…

AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction
نماذج الذكاء الاصطناعي arXiv cs.AI

AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction

arXiv:2606.13051v1 Announce Type: new Abstract: Despite advances in information extraction driven by deep learning and large language models, performance gaps remain in highly specialized biomedical fields, where domainspecific complexity poses challenges for generalist models. In this work, we focus on the domain of…

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?
نماذج الذكاء الاصطناعي arXiv cs.AI

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

arXiv:2606.13148v1 Announce Type: new Abstract: Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate foundation models can forecast well, but do not reas…

The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
نماذج الذكاء الاصطناعي arXiv cs.AI

The Pragmatic Persona: Discovering LLM Persona through Bridging Inference

arXiv:2604.24079v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) reveal inherent and distinctive personas through dialogue. However, most existing persona discovery approaches rely on surface-level lexical or stylistic cues, treating dialogue as a flat sequence of tokens and failing to capture t…

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment
نماذج الذكاء الاصطناعي arXiv cs.AI

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment

arXiv:2606.13176v1 Announce Type: new Abstract: Mental health problems such as anxiety, depression, and suicide remain urgent global challenges, where timely and accurate assessment is critical for effective intervention. Recently, large language models have been explored for mental health assessment. However, existi…

Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach
نماذج الذكاء الاصطناعي arXiv cs.AI

Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach

arXiv:2606.13192v1 Announce Type: new Abstract: User experience (UX) centered on usability, perceived consistency, and functional clarity is fundamental to real-world user interfaces (UI). The application of multimodal large language models (MLLMs) in the field of user interfaces is evolving rapidly, such as visual…

BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning
نماذج الذكاء الاصطناعي arXiv cs.AI

BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning

arXiv:2604.27277v3 Announce Type: replace-cross Abstract: Brain MRI underpins a wide range of neuroscientific and clinical applications, yet most learning-based methods remain task-specific and require substantial labeled data. Here we show that a single self-supervised representation can generalize across heterogene…

Under What Conditions Can a Machine Become Genuinely Creative?
نماذج الذكاء الاصطناعي arXiv cs.AI

Under What Conditions Can a Machine Become Genuinely Creative?

arXiv:2606.13196v1 Announce Type: new Abstract: Recent AI systems can generate texts, software architectures, hypotheses, designs, and scientific workflows that appear creative. This paper asks under what conditions a machine can become genuinely creative, and how human agency can be preserved within shared cognitive…

ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning
نماذج الذكاء الاصطناعي arXiv cs.AI

ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning

arXiv:2606.13197v1 Announce Type: new Abstract: Multi-agent debate (MAD) can improve large language model reasoning, but fixed debate pipelines often waste computation and can amplify correlated errors among similar agents. We propose ARMOR-MAD, a training-free heterogeneous MAD framework that treats debate as condit…

Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints
نماذج الذكاء الاصطناعي arXiv cs.AI

Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints

arXiv:2606.13211v1 Announce Type: new Abstract: AI systems are being deployed across medical imaging faster than their failure modes are understood. At this point in time, the failure of greatest clinical concern is hallucination: clinically plausible but factually incorrect outputs, including fabricated anatomical s…

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis
نماذج الذكاء الاصطناعي arXiv cs.AI

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis

arXiv:2606.13220v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as interactive assistants for technical problem solving. However, when users provide incomplete descriptions or plausible but unverified explanations, LLMs may prematurely align with these assumptions and propose soluti…

Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability
نماذج الذكاء الاصطناعي arXiv cs.AI

Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability

arXiv:2605.25225v2 Announce Type: replace-cross Abstract: Mechanistic interpretability often studies Transformer behavior by intervening on internal activations through activation patching, causal tracing, path patching, and steering directions. This paper develops Transformer Field Theory: a response-theoretic frame…

From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification
نماذج الذكاء الاصطناعي arXiv cs.AI

From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification

arXiv:2606.13262v1 Announce Type: new Abstract: Recent approaches combining Large Language Models (LLMs) with retrieval-augmented reasoning have shown promise for automated fact verification. To process complex claims, these verification pipelines typically execute multi-stage workflows that coordinate tightly couple…

If LLMs Have Human-Like Attributes, Then So Does Age of Empires II
نماذج الذكاء الاصطناعي arXiv cs.AI

If LLMs Have Human-Like Attributes, Then So Does Age of Empires II

arXiv:2605.31514v3 Announce Type: replace-cross Abstract: Much research has been carried out on large language models (LLMs) and LLM-powered agentic workflows. However, many works within the field state emergence of, ascribe to, or assume, generalised anthropomorphic attributes to them (e.g., morality or understandin…

ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space
نماذج الذكاء الاصطناعي arXiv cs.AI

ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space

arXiv:2606.13282v1 Announce Type: new Abstract: As AI systems are deployed in high-stakes ethical contexts such as healthcare triage, autonomous vehicle control, and employment screening, formal methods for evaluating their robustness against adversarial manipulation of ethical reasoning remain underdeveloped. This p…

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
نماذج الذكاء الاصطناعي arXiv cs.AI

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning

arXiv:2606.13316v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods often encourage unnecessarily long reasoning rollouts, which can degrade reasoning coherence…

Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems
نماذج الذكاء الاصطناعي arXiv cs.AI

Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems

arXiv:2606.06525v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as powerful foundation models with strong reasoning capabilities across domains. Beyond reactive text generation, agentic LLMs enable autonomous workflow execution through modular task decomposition and coordinated too…