GGUF is the format to use for local inference today. It bundles model weights, metadata, and the tokenizer into one extensible file that loads fast and supports quantization, which is why most modern runtimes build around it. GGML, the older list-based format, still shows up in archives and old workflows, but it has no real place in a current setup.
TL;DR:
- GGUF bundles weights, tokenizer, and typed metadata in one aligned file, letting compatible runtimes map it from disk and skip unknown fields.
- Published head to head benchmarks are scarce; model speed and RAM depend more on size and quantization than format, while lower precision can reduce quality.
- Hugging Face Transformers dequantizes GGUF weights to fp32 in training contexts, so quantized files save memory during inference, not fine tuning.
- If conversion fails, check for a missing config.json, a Python version mismatch, or an incomplete download; most readers should choose a trusted, already converted build instead.
- Ollama, llama.cpp, and Hugging Face support GGUF directly, while remaining GGML support mainly serves backward compatibility and archival files rather than new model workflows.
Table of Contents
- At-a-glance comparison: GGUF vs GGML
- How the formats differ technically
- What these differences mean when you run models locally
- How to convert a Hugging Face model to GGUF and run it
- Real-world example: GGUF powering local, privacy-first apps
- Comparative performance benchmarks between GGUF and GGML models in terms of speed and memory usage
- Compatibility and support differences for different local AI applications beyond llama.cpp
- Format adoption and community momentum
- Security and data privacy implications of using GGUF vs GGML locally
- A short take on where the format is headed
- A straightforward way to run GGUF models locally
- FAQ
- Sources
At-a-glance comparison: GGUF vs GGML
The two formats solve the same basic problem (packaging a model for local use) in very different ways, and the gap explains why one replaced the other.
- Metadata model: GGUF stores metadata as typed key-value pairs; GGML used an untyped list that broke when new fields were added, according to IBM's comparison.
- Packaging: GGUF keeps tensors, metadata, and the tokenizer in a single file, so there is nothing else to track down before running a model.
- Extensibility: New architectures and tokenizer types can be added to GGUF without breaking older files, a design goal the GGUF specification documents directly.
- Quantization and loading: GGUF's layout is friendly to memory-mapped (mmap) loading and supports several quantization types, which keeps RAM use down on modest hardware.
GGML's list-based hyperparameters worked fine for early, simple models, but every new architecture risked breaking compatibility with existing tools.
How the formats differ technically
The practical gap between GGUF and GGML comes down to how each one stores information about the model, not just the raw weights.
GGML described a model's hyperparameters as a flat, untyped list. That worked when models were simple, but adding a new field, say, a different tokenizer vocabulary or an extra architecture parameter, meant changing the list's shape. Older files and older code often could not read newer ones, and newer code sometimes misread older files. IBM's technical explainer describes this as the core reason GGML caused breaking changes as the ecosystem grew.
GGUF replaced that list with typed key-value metadata. Each field has a name and a type, so adding a new key (for a new tokenizer, a new architecture flag, or anything else) does not disturb fields that already exist. Readers can skip keys they don't recognize instead of failing outright.
Beyond metadata, GGUF defines a consistent file header, a version number, and alignment settings for tensor data, details laid out in the gguf.h specification inside the llama.cpp repository. That alignment matters because it lets runtimes memory-map the file directly instead of copying it into RAM first, which speeds up loading.

GGUF also supports a range of block-wise quantization types, the low-precision number formats that shrink a model's footprint on disk and in memory. Because the format was built with these types in mind from the start, tools that produce and read GGUF files handle quantized weights predictably, which is a big part of why quantized models distributed as GGUF just work across different runtimes.
What these differences mean when you run models locally
On a real machine, these design choices turn into faster load times and fewer compatibility headaches.
Because GGUF packs everything into one aligned, mmap-friendly file, runtime projects can map it straight from disk instead of parsing scattered files first. That is a major reason llama.cpp, Ollama, and similar tools standardized on it. The format is built for inference, not training: when Hugging Face Transformers loads a GGUF file, it dequantizes the weights to fp32, according to the Transformers GGUF documentation, so the memory savings apply when you run the model, not when you fine-tune it.
A few practical points follow from this:
- Loading speed: A single aligned file with embedded metadata loads faster than a multi-file setup with separate configuration.
- Memory use: Quantized GGUF files take up less RAM at rest; training or fine-tuning in a framework like Transformers expands that usage back toward full precision.
- Tool support: llama.cpp, Ollama, and Hugging Face's own model hub all expect or support GGUF directly.
GGUF's quantization options let a model that would otherwise need far more memory run on a machine with limited RAM, which is the main reason quantized GGUF builds dominate the local model listings people use to pick a model for their hardware.
How to convert a Hugging Face model to GGUF and run it
Converting a model you found on Hugging Face into something you can run locally follows a fairly fixed path.
- Run the
convert-hf-to-gguf.pyscript from llama.cpp against the source model folder, picking an--outtypequantization flag (common choices trade off size against accuracy). - If the model uses an architecture llama.cpp does not already recognize, you need to define that architecture and add the matching GGML graph logic, a process the llama.cpp HOWTO walks through step by step.
- Watch for common snags: a missing
config.jsonin the source folder, a Python version mismatch with the conversion script's requirements, or an incomplete download, all of which community conversion discussions flag repeatedly. - When a pitfall shows up, it is usually faster to re-download a trusted, already-converted GGUF build than to debug a conversion by hand.
- Once you have a working GGUF file, the runtime path is simple: place it where your tool expects it, load it with llama.cpp or Ollama, and start running prompts.
Because adding a brand-new architecture to llama.cpp is genuinely technical work, most people are better off sourcing a preconverted GGUF file from a reputable model repository rather than building the pipeline themselves.
Real-world example: GGUF powering local, privacy-first apps
Desktop apps built for offline AI are a direct, practical example of what GGUF's design makes possible. A single quantized file that loads quickly and runs without a network connection is exactly what a downloadable, privacy-first assistant needs. We built GreenCube around that same idea: it runs entirely on a Windows or Mac computer, downloading one model once during setup rather than streaming anything from a server afterward.
Chats and files never leave the device, and the model keeps working with no internet connection once it's installed. That only works because the model ships as a single, self-contained, quantized file, the same packaging GGUF was designed to provide.
Comparative performance benchmarks between GGUF and GGML models in terms of speed and memory usage
Independent, apples-to-apples benchmark numbers comparing GGUF and GGML side by side are not widely published, so it's more useful to look at what the format design predicts and what practitioners consistently report.
GGUF's aligned, mmap-compatible layout lets a runtime map the file directly from disk rather than parsing it piece by piece, which tends to shorten load times compared to older, less structured formats. Because GGUF also supports a range of quantization types natively, a quantized GGUF build of a given model generally uses less RAM at rest than a quantized or loosely packaged equivalent.
GGML files, by contrast, predate most of this tooling and the ecosystem built around them has largely moved on, so current runtimes rarely optimize for GGML loading paths at all. In practice this means the realistic choice is not between two actively maintained, equally fast formats. It's between a modern format with active performance work behind it and a legacy one with little ongoing attention.
For a given model, the speed and memory profile you experience will depend far more on the quantization level you pick (lower precision generally means less RAM and faster loading, at some cost to output quality) than on any GGUF versus GGML comparison. That's why guides on choosing models for your hardware tier focus on quantization level and parameter count rather than file format alone.

Compatibility and support differences for different local AI applications beyond llama.cpp
GGUF's reach extends well past llama.cpp itself. Ollama, one of the most common ways people run local models through a simple command-line or app interface, reads GGUF natively. Hugging Face's model hub recognizes and displays GGUF files directly, and the Transformers library can load them too, with the caveat that it dequantizes weights to fp32 for training contexts.
Beyond these, a growing set of desktop chat apps, local API servers, and developer tools standardize on GGUF as their expected input format, largely because its single-file packaging removes the need to hunt down separate config files, tokenizer files, and weight files that other formats often require.
GGML support, by comparison, is effectively frozen. Tools that still mention it generally do so for backward compatibility or archival reading, not as a format they recommend for new models. If you come across a .ggml file today, the practical move is to look for a GGUF version of the same model rather than trying to find software that still loads the original.
This matters for anyone evaluating a local AI application beyond llama.cpp: checking whether a tool explicitly lists GGUF support is a fast way to judge whether it is actively maintained for modern local inference, or built on older assumptions.
Format adoption and community momentum
Format adoption tends to follow the tools people actually use, and GGUF has that momentum. llama.cpp standardized on it, Ollama built its model distribution around it, and Hugging Face added native GGUF recognition to its hub, which together cover most of the on-ramps people use to find and run local models.
That kind of layered adoption, where the inference engine, the distribution hub, and the user-facing tools all agree on one format, is what keeps a format alive and growing rather than frozen in place. New quantization schemes and architecture support tend to land in GGUF first, because that's where the active development and the active user base both are.
GGML, by contrast, stopped being the focus of new feature work once GGUF's extensibility advantages became clear. Maintainers and model publishers migrated their distribution pipelines, and the community's documentation, troubleshooting threads, and conversion tools shifted with them. A format without active contributors adding features or fixing edge cases tends to fall further behind with each new model architecture that appears, and that's the position GGML now occupies.
For anyone picking a format today, the practical signal is simple: look at where new quantization types, new architectures, and new tooling show up first. Right now, that's GGUF, consistently.
Security and data privacy implications of using GGUF vs GGML locally
Neither GGUF nor GGML changes the basic privacy picture of local inference: a model file, regardless of format, runs on your own machine without needing to send your prompts or documents anywhere. The format governs how weights and metadata are packaged, not whether your data leaves the device.
That said, format choice does affect how safely you can source a model. Because GGUF is now the default output of official conversion tools and the format most model hubs display prominently, it's easier to find GGUF builds from identifiable, traceable sources. Older .ggml files circulating without clear provenance are harder to verify, since the tooling and documentation around them has largely moved on, meaning fewer eyes are checking that a given file matches its claimed source.
The practical privacy and security advice is the same regardless of format: download model files from a source you trust, keep the application running the model up to date, and favor tools that make clear where a model came from and how it was converted. A model file packaged as a single, well-documented GGUF file with embedded metadata makes that kind of verification more straightforward than a loose collection of older files with no built-in record of their origin.
A short take on where the format is headed
GGUF is the practical choice for local inference, full stop. The sensible path forward is to prefer GGUF-distributed models, lean on official conversion tools like convert-hf-to-gguf.py instead of hand-editing files, and re-download a converted build rather than trying to rework a legacy .ggml file by hand. Favor distributions with clear versioning and complete metadata: that clarity is exactly what GGUF was built to provide.
— Hector Gras
A straightforward way to run GGUF models locally
We built our app around the same idea that makes GGUF useful in the first place: one file, one setup step, and nothing sent anywhere afterward. Install it on Windows or Mac, sign in once with Google or Microsoft to unlock the license, and pick a model to download during setup. From there, everything runs on your own computer.

- Runs fully offline once set up, with no subscription and no ongoing cloud dependency.
- Chats and documents stay on your device; nothing is uploaded unless you turn on a feature that explicitly goes online.
- Costs a one-time fee, with everything included.
- Works on both Windows and Mac, with tiers reflecting what different computers handle well (a slower machine will still answer and build things, just at a slower pace).
If local, private AI without a subscription sounds like what you need, get GreenCube and start from a single download.
FAQ
Are GGUF models better than GGML models?
For current local inference, yes: GGUF's key-value metadata, single-file packaging, and quantization support make it more reliable and better supported than GGML's older list-based format. GGML still works for archival purposes, but new tools rarely add features for it.
Can I still use old GGML files?
Some older tools can still read .ggml files, but active development has moved to GGUF, so support is shrinking. The more reliable path is to find a GGUF version of the same model rather than keeping the original .ggml file in your workflow.
Does converting to GGUF reduce model quality?
Conversion itself does not change quality; the quantization level you choose during conversion does. Picking a lower-precision --outtype flag in the conversion script trades some accuracy for a smaller, faster-loading file.
Do I need to convert models myself to use a tool like GreenCube?
No. GreenCube downloads its model automatically during setup, so there's no manual conversion, file hunting, or command-line work involved. You pick a model option once, and the app handles the rest.
Why do some apps need more RAM to run local models than others?
RAM needs come down to model size and quantization level: a larger or less-compressed model needs more memory to load and run. Guides on matching models to hardware or options like cloud-based GPU desktops are both ways to handle a model that's heavier than your current machine comfortably supports.
Sources
- GGUF versus GGML | IBM
- Transformers documentation: GGUF
- GGUF documentation (ggml-org/ggml)
- gguf.h (llama.cpp repository)
