What was released, by whom, under which licence, in what format — and what it does not support yet. Every figure comes from the model card, the config or the vendor's own announcement.
12 August 2026 · Alibaba
Alibaba released Qwen3.8-2.4T-A95B: 2.4 trillion total parameters with 95B active, 512 experts, 262K context. The licence is no longer Apache 2.0. Specs and what is missing.
11 August 2026 · Meta
Meta released Muse Glimmer 30B under Apache 2.0 — a multimodal agent model with a 1.8B vision encoder, 131k context, GGUF and ExecuTorch builds. Full specs, licence and runtime support.
11 August 2026 · Motif Technologies
Motif Technologies released Motif 3, a 314B-parameter Mixture-of-Experts model with 13.2B active per token, 256K context and an MIT licence. Architecture, sizes and what is missing.
11 August 2026 · Kurakura AI
Kurakura AI released Luth-2, French-specialised small language models at 0.8B and 2B under Apache 2.0, built on Qwen3.5 with GGUF builds from 1.2 GB. Specs, training and limits.
11 August 2026 · Moonshot AI
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model with 104B active per token, native vision and a 1M-token context. Specs, licence and what it takes to run.
11 August 2026 · DeepSeek
DeepSeek V4 Flash 0731 ships under MIT with a 1,048,576-token context and FP8 weights at 166.9 GB. Architecture, expert configuration, licence and runtime support.
11 August 2026 · Z.ai
Z.ai's GLM-5.2 is an MIT-licensed Mixture-of-Experts flagship with a 1M-token context and an IndexShare sparse attention design that cuts per-token FLOPs by 2.9× at full length.
11 August 2026 · Alibaba
Qwen3.5 spans 0.8B to 397B-A17B, all Apache 2.0, with 262,144-token native context extensible to about one million. Every published size, with what each one is for.
11 August 2026 · NVIDIA
NVIDIA released Nemotron 3.5 Lightning, a 30B Mixture-of-Experts model with 3B active parameters, interleaved Mamba-2 layers, 1M context and day-one GGUF from the llama.cpp project itself.
11 August 2026 · inclusionAI
inclusionAI released Ling-3.0-tiny under MIT: 7.9B total parameters with 1.3B active, a 3:1 KDA/MLA hybrid architecture, and measured throughput on an M4 Pro MacBook. Specs and runtime status.
11 August 2026 · Liquid AI
Liquid AI released LFM2.5-2.6B with agentic post-training, a 128K context window and official GGUF, MLX, NVFP4 and MXFP4 builds. Specs, licence and runtime support.
11 August 2026 · Google DeepMind
Google DeepMind published quantisation-aware-trained mobile builds of Gemma 4 E2B and E4B under Apache 2.0. What QAT changes, sizes, context and multimodal support.
11 August 2026 · Prism ML
Bonsai 27B stores each weight as a single sign bit — trained that way, not compressed afterwards. 27B parameters in 3.8 GB, 262K context, vision tower, Apache 2.0.
11 August 2026 · Microsoft Research
Microsoft Research released Fara 1.5 at 4B, 9B and 27B under MIT — multimodal computer-use agents that drive a browser from screenshots. Built on Qwen3.5.
11 August 2026 · Mistral AI
Mistral's Ministral 3 8B Instruct 2512 is an Apache 2.0 model with vision capability and a 262,144-token context, published alongside a reasoning variant.
11 August 2026 · Meta
Meta's newest open weights ship under the Muse name, not Llama. What the meta-llama account still holds, what Llama 4 is, and what changed.