Perceive the world in its nature and explore the unknown
Perception, reasoning, and interaction: The world reveals itself not through pixels, but through the symbols and structures underneath. Vision speaks its own language, directing where to look and what to understand. To live in the world is to engage with the knowledge it offers—absorbing it and transforming experience into capability, not just caching isolated data.
3.7k curated diagrams and 18.3k human-validated questions across six scientific domains, evaluating MLLMs on diagram-to-code parsing, editing, and question answering — alongside 16 controllable settings covering both foundational and agentic (context, tool use, state, planning) capabilities.
A framework that enables structured visual reasoning in spatial and object-centric space, improving visual perception tasks through reinforcement learning with verifiable reasoning chains.
Current MLLMs perform poorly on basic diagram perceptual tasks, relying on textual shortcuts rather than visual understanding (math blind). Representing diagrams as graphs of primitives is crucial; our results show that strong low-level perception drives faithful high-level mathematical reasoning.
A self-supervised symbolic auto-encoder that encodes diagrams into structured primitives and their interrelationships, achieving 98.2% MSE reduction in geometric diagram reconstruction, improving by +13% on the diagram perception benchmark, and by +3% on MathVerse and GeoQA reasoning benchmarks.
An agentic learning framework that enables progressive improvement through multimodal semantic memory, integrating visual and logical memory to refine both perception and reasoning for lifelong and cross-domain agentic learning.
Pre-trained Artemis models for structured visual reasoning and perception policy learning across various visual tasks.
The model trained with GEOMETRIC for enhanced geometric diagram understanding and mathematical reasoning.
3.7k scientific diagrams paired with 18.3k human-validated questions, spanning diagram-to-code parsing, diagram editing, and diagram question answering across 16 controllable foundational and agentic evaluation settings.
A benchmark that isolates diagram perception from reasoning in MLLMs, featuring 1.2K diagrams and 1.6K curated questions across four tasks: shape classification, counting, relationship identification, and grounding.
A geometric logic-form dataset of 134K diagram-description pairs across three subsets (SymParser-100K, SymVAE-16K, SymHPR-9K), annotating each diagram with point, line, and shape instances, their geometric relations, and normalized coordinates for symbolic visual learning.
A structure-aware geometric diagram-description dataset encoding shapes, attributes, and interrelationships as graphs with fine-grained spatial annotations for model training.
Official implementation of the Artemis framework for structured visual reasoning and perception policy learning with reinforcement learning.
Official implementation of the ViLoMem framework, featuring multimodal semantic memory architecture and agentic learning algorithms.