ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
· Source: arXiv cs.AI
ZGCM‑1 is a 7‑billion‑parameter base model trained from scratch using a strategy that prioritizes data, resource, and algorithmic efficiency. The authors argue that while compact models cannot store all the information available on the web, they can offset their capacity limits by combining internal deliberative reasoning with active use of external tools. To handle contexts up to 256 k tokens, the architecture alternates sliding windows with full attention, and employs a low‑precision optimizer (FP8) called Muon. Training follows a progressive curriculum that gradually expands context size (16 K, 64 K, 256 K) and reformulates interaction sequences as Markov decision processes, easing agent‑behavior learning. An automated workflow lets swarms of agents manage cluster infrastructure, data selection, and rapid diagnostic evaluation. Benchmarks show that ZGCM‑1‑7B remains competitive within its family and, on mathematical reasoning and agent‑based search tasks, matches the performance of much larger models such as Qwen3‑235B‑A22B and GLM‑5.1. The project also reports roughly a 4.2‑fold speedup in training at 16 K tokens. All weights, code, and data are publicly available for the community to explore and extend.
This news is significant because it demonstrates how smaller models can approach the performance of large‑scale systems through architectural and training innovations, opening the door to more accessible and sustainable AI applications. The full openness of the project also encourages collaborative research and accelerates the development of new AI‑based tools.
Read the original article on arXiv cs.AI
This summary is an informational synthesis produced by dataqbs.com. All rights to the original content belong to its author and the cited media outlet. We act solely as curators of technology news and claim no authorship.