Foundation Models has one awkward property that no amount of prompting or architecture will fix: the framework only works where Apple Intelligence is on. Older device, unsupported region, or the person turned the feature off — and your feature simply isn't there.
In iOS 27 the way around it became official. The LanguageModel protocol is open, and a model exported through Core AI can be dropped into the same LanguageModelSession, with the same prompts, tools, and structured output. This piece walks the whole chain, from listing models in a terminal to the first response in the app.
This is part three; Core AI itself and what's new in Foundation Models are separate. Requirements: macOS 27, iOS 27, and Xcode 27 or later.
When this is actually worth doing#
Three reasons Apple itself gives for bringing a model that isn't theirs:
- you need a specialized capability that model provides;
- you need to support devices that don't have Apple Intelligence;
- you need cross-platform parity — the same model on your server and in the app.
There's a fourth, unstated but obvious: the built-in model ships with the OS. It already changed in 26.4 and changed again in 27, and each time Apple wrote "test your prompts." A model in your bundle changes when you decide it does.
Getting a model: registry and export#
Export lives in the open-source coreai-models Swift package, which carries both the export recipes and the utilities. It all starts in a terminal: install the uv package manager, clone the repository, and change into the coreai-models directory.
Then look at what's supported:
uv run coreai.model.registry --list-models # models in the registry and their export presetsThe column you want in the output is HF_ID — the identifier used for export. Make your first model something around 0.6B parameters: it downloads quickly and runs comfortably on device. Fighting with quantization of a seven-billion-parameter model while you're still establishing "does this work at all" is one variable too many.
Models are specialized for the hardware they run on, so the platform is set at export time:
uv run coreai.llm.export HF_ID # export for macOS
uv run coreai.llm.export HF_ID --platform iOS # same model for iOSWhat comes out is a resource folder: the .aimodel plus the tokenizer and whatever else the model needs. Add the whole folder to the app.
Do check the models directory inside coreai-models — every model has its own README with the exact recipe and its own requirements. There's no universal "export anything" command here, which is more honest than a promise that breaks on the second model.
Wiring up the package#
CoreAILanguageModel lives in that same coreai-models. It's added like any dependency: File > Add Package Dependencies, search for coreai-models, add it. In the product table, CoreAILM will show None next to it — pick your app there, or the package attaches and the module never appears. That's the standard way to lose twenty minutes to "why won't it import."
The code itself#
The whole bridge is four lines:
import FoundationModels
import CoreAILanguageModels
// The resource folder you exported and bundled with your app.
guard let modelURL = Bundle.main.url(forResource: "The model name",
withExtension: nil) else {
// Handle the missing resource.
return
}
// Load the model and create a session that runs requests through it.
let model = try await CoreAILanguageModel(resourcesAt: modelURL)
let session = LanguageModelSession(model: model)withExtension: nil isn't a typo — you're pointing at a folder, not a file.
CoreAILanguageModel conforms to LanguageModel, so the session is created with exactly the same initializer as for the built-in model. From there, your feature code sees no difference at all:
let response = try await session.respond(
to: "Summarize the key points from this meeting transcript: \(meetingTranscript)."
)Streaming, tools, @Generable, GenerationOptions — all of it carries over without a single edit. Which is the entire point of the exercise: your feature layer doesn't know which model sits underneath.
Loading is async, and the user can tell#
The try await on CoreAILanguageModel's initializer is covering real work: the framework compiles the model and loads its tokenizer before the first request. Call it at the moment someone taps the button, and someone waits.
The workable pattern is loading early, when the request is at least a second or two out: a screen opened, typing started, a flow began. That's also the place to prewarm the session with prewarm(promptPrefix:), so instructions and tool definitions land in the KV cache before the first respond.
For a large model, async alone won't cover it — you want coreai-build ahead-of-time compilation and explicit control over the specialization cache. That's the first article's territory, and language models have a specific wrinkle there: expectFrequentReshapes in SpecializationOptions. An LLM's sequence length grows one token per step, and per-shape optimization eats more than it returns.
Reasoning models behave correctly out of the box#
Open models that emit a chain of thought are a nuisance to integrate by hand: their intermediate text has to be separated from the answer, and that usually ends in regexes over tags.
Core AI recognizes that output itself and routes it into the transcript as a reasoning segment. It doesn't reach response.content — the person sees only the answer. You can still read the reasoning when you're working out why an answer came out odd.
Whether a model reasons at all depends on what you exported, and it's checked explicitly:
if model.capabilities.contains(.reasoning) {
// The model supports reasoning.
}What to measure#
Core AI picks the engine for the device on its own — GPU, CPU, or Neural Engine, depending on how the model was exported. You don't drive that from feature code, but the result is worth verifying.
Look in Instruments, at the Foundation Models instrument: asset load times, token counts, and per-request durations. The first number I'd check is cached input tokens over total input tokens between turns. If it's low, the prefix is being recomputed every time, and the problem isn't the model — it's how the session is put together.
What it costs you#
No illusions about the trade. You get independence from Apple Intelligence and control over the model version. In exchange you take on three things the system used to handle: the model's weight in your delivery (or downloading and updating it), on-device specialization with its first-run latency, and quality — an open 0.6B model owes you nothing, and that has to be checked against your data rather than your impressions.
So the order I'd hold to: the built-in model first, with an honest evaluation of the feature on it. Core AI when you've hit the wall of Apple Intelligence availability or need a capability the system model lacks. The good news is that moving between those costs one line rather than a rewrite.



