For six years, Core ML was the only answer to "how do I run my model on an iPhone." iOS 27 added a second one — Core AI, a separate framework with its own format, its own compiler, and its own debugger. Core ML hasn't gone anywhere: decision trees and tabular feature engineering still belong there. Core AI is for neural networks, and for getting modern architectures onto CPU, GPU, and Neural Engine without hand-tuning for every chip.
Here's the stack in order: how a model gets into the project, what happens on the first load (and why it takes so long), how to control that, and where I'd slow down before shipping.
Versions: Core AI ships with iOS 27, iPadOS 27, macOS 27, tvOS 27, visionOS 27, and watchOS 27. The signatures below come from Apple's documentation as of late September 2026. The API is released, not beta — but different parts of the stack support different platform sets, and I'll come back to that.
What Core AI actually includes#
The framework is only the visible part. Apple put five things under the Core AI name, and mixing them up gets expensive:
- Core AI framework — the Swift API:
AIModel,InferenceFunction,NDArray,AIModelCache. .aimodel— the portable model format. Works across Apple devices, but doesn't execute on its own.- coreai-torch — PyTorch extensions: convert a model to
.aimodel, export several inference functions into one artifact, use built-in hardware-optimized ops for attention and normalization, plug in custom Metal 4 kernels. - coreai-optimization — quantization and palettization with per-layer control over the technique.
- coreai-models — a catalog of export-ready models plus a Swift package of helpers.
Separate from all of that is the Core AI Debugger, a macOS app that does something Core ML never could: trace tensor values back to your original Python source. Xcode also gained a Core AI debug gauge and a Core AI instrument.
The first thing that breaks the build#
Xcode can't build a project containing .aimodel out of the box. It needs the Metal Toolchain, which isn't installed by default:
% xcodebuild -downloadComponent MetalToolchainOr through Xcode > Settings > Components > Other Components > Metal Toolchain. Without it, the build fails on a missing Metal compiler — and from the error text alone, nothing suggests the fix is one checkbox in settings.
The file itself goes in by dragging it into the Project Navigator; after that it should show up in the target's Compile Sources phase. If it doesn't, the model never reaches the bundle.
Look at the model before you write code against it#
Select the .aimodel in the navigator and the viewer opens. The General tab carries the size in parameters and in bytes, the metadata (description, author, license, arbitrary key-value pairs — editable inline, Xcode saves automatically), and, more usefully, numeric precision split between compute and storage. It also shows the operation distribution across the graph, sorted by count.
The Functions tab holds the signatures: names and types of every input and output, with a question mark in an NDArray dimension wherever that dimension is resolved at runtime. Most models have a single function.
Open this tab before the first line of code. Half of all integration bugs are a shape or scalar-type mismatch, and the viewer shows them in ten seconds.
Loading: why await, and why it's slow#
import CoreAI
// Specialize the model for this device and load it.
let model = try await AIModel(contentsOf: urlOfModel)
// Load a function from the model.
guard let function = try model.loadFunction(named: "main") else {
// Handle case where expected function is not found.
}init(contentsOf:) isn't async out of politeness. Between a .aimodel file and a working model sits specialization — Core AI looks at the compute units this particular device offers and generates executable code for that hardware and OS version. On large models this takes real time, and the user will see it.
loadFunction(named:) isn't cheap either; it prepares the resources for one specific function. It throws on a load failure and returns nil when no function by that name exists — two distinct cases that collapse very easily into one try? and then cost you three hours of debugging. All names are available through functionNames.
The same inference function can safely be called from different tasks at once. The documentation says so explicitly, so there's no need to wrap it in a serializing actor of your own.
Inference: NDArray and two kinds of memory access#
Inputs and outputs are InferenceValue — either an NDArray or an image. Which one is visible in the Functions tab, or through the descriptor at runtime.
You need the runtime check when the model arrives from a server and can change between releases without an app update:
let function: InferenceFunction = ...
let functionDescriptor = function.descriptor
guard let valueDescriptor = functionDescriptor.inputDescriptor(of: "input"),
case .ndArray(let arrayDescriptor) = valueDescriptor else {
// Handle input not found, or an unexpected type.
}
guard arrayDescriptor.shape == [3, 4] else {
// Handle an unexpected shape.
}
guard arrayDescriptor.scalarType == .float32 else {
// Handle an unexpected scalar type.
}Then the array itself. Note how access is split: an NDArray is read-only by default, and writes go through mutableView(as:). Swift enforces that at compile time, so the code always shows what's happening to the memory.
// An array matching the expected shape and type.
var input = NDArray(shape: [3, 4], scalarType: .float32)
// A mutable view for writing.
var mutableView = input.mutableView(as: Float.self)
guard let elements = mutableView.contiguousElements else {
// Handle non-contiguous memory layout.
}
writeInputData(into: elements)
// Run it.
var outputs = try await function.run(inputs: ["input": input])
guard let predictionValue = outputs.remove("prediction") else {
// Handle output not found.
}
guard let prediction = predictionValue.ndArray else {
// Handle output of unexpected type.
}
processOutput(prediction.view())The keys in inputs are the names set at conversion time, not something you get to invent on your side. contiguousElements returns nil for a non-contiguous layout — a rare path, but not one to force-unwrap through.
For images, CVPixelBuffer replaces NDArray, and the descriptor reports the expected width, height, and pixel format; -1 in a dimension means it's dynamic. Inputs are assembled with InferenceFunction.Inputs() and insert(_:for:).
The specialization cache is the most useful part of the API#
By default AIModel specializes the model and caches the result: the first run pays full price, later ones load what's already there. Except the system is free to delete cached assets under storage pressure — at which point a "later" run quietly becomes a first run again.
So checking the cache before loading is worth doing every time — it answers whether you need to show progress:
func loadModel(from modelURL: URL) async throws -> AIModel {
let cache = AIModelCache.default
// A non-nil result means the model was previously specialized and cached.
if let model = try cache.model(for: modelURL, options: .default) {
return model
}
// No cached specialization. Tell the person, then specialize now.
Task { @MainActor in
informUser("Preparing AI features. This may take a while…")
}
return try await AIModel(contentsOf: modelURL, options: .default)
}cache.model(for:options:) specializes nothing — it only answers yes or no. If the model is downloaded, you can specialize it ahead of time at a convenient moment with AIModel.specialize(contentsOf:options:): it stores the assets in the cache and returns the ready model, and every later init with the same URL and options loads straight from cache.
Retention is governed by cachePolicy. The default lets the system reclaim assets under storage pressure. .persistent forbids that, and it exists for one specific case: you delete the source .aimodel so the device isn't holding two copies of the same weights. .persistent is unavailable on tvOS — local storage there has to stay purgeable.
If you have several apps or extensions sharing a model, set up an app group and create the cache with AIModelCache(appGroup:). One specialization for the group instead of a copy per target.
One more wrinkle: you can't delete the source file and keep calling AIModel(contentsOf:) — the source URL is the key the specialization is indexed under. For that case there's bookmarkData: save it after specializing, and on the next launch restore the model through AIModel(resolvingBookmark:), bypassing the source. A bookmark can be invalidated by an OS update, so the "not found" branch has to be a working path rather than a fatalError.
Specialization options#
SpecializationOptions.default lets the system pick the CPU/GPU/Neural Engine mix that minimizes latency. There's also .cpuOnly and init(preferredComputeUnitKind:).
I see exactly one solid reason to leave the default: a small model running in the background that shouldn't compete for the GPU with your UI. Then .cpuOnly earns its place. Everywhere else the default usually wins, and that's worth measuring rather than assuming. Available compute units differ by device — check ComputeUnitKind.
One more flag to know up front is expectFrequentReshapes. For dynamic-shape models, Core AI optimizes the function for each new input shape by default. For a language model, where sequence length grows one token per step, that optimization starts costing more than it saves. Setting the flag to true switches to the generic dynamic version of the function.
Ahead-of-time compilation: move the heavy part to your Mac#
Part of specialization can happen on the build machine. coreai-build compiles a .aimodel into a set of .aimodelc assets, one per device architecture:
% xcrun coreai-build compile MyModel.aimodel --platform iOS --min-deployment-version 27.0 --output compiled/The output is MyModel.<arch>.aimodelc, where <arch> matches what AIModel.deviceArchitectureName returns at runtime. Each compiled asset works on any OS version at or above the --min-deployment-version you passed.
Then there's a fork in the road. Bundling every architecture means shipping several copies of the model, of which a device uses one. Apple's recommendation is to host the .aimodelc files yourself and download the matching variant after checking the architecture at runtime:
let arch = AIModel.deviceArchitectureName
let assetName = "MyModel.\(arch).aimodelc"A .aimodelc loads through the same AIModel(contentsOf:) — the loading code doesn't change. Downloads and updates belong in Background Assets.
Now the honest part about this feature's edges. Ahead-of-time compilation covers only devices that support Apple Intelligence: iPhone and iPad with A17 Pro or later, Mac with M1 or later, Vision Pro with M2. On tvOS and watchOS it doesn't exist at all — even though Core AI itself runs there. And even with AOT, some specialization still happens on device; how much depends on the model and the compute units it uses. Apple's wording is careful here: less work, not no work.
Where I'd slow down#
Three things worth checking before Core AI goes into a release plan.
First, platform coverage is uneven. The framework is listed for all six OSes, but .persistent is missing on tvOS, AOT is missing on tvOS and watchOS, and AOT is absent entirely on devices without Apple Intelligence. The feature-by-platform matrix is denser than the framework page suggests, and you want it filled in before the architecture decisions, not after.
Second, the cost of the first run. Specialization happens on the user's device and it isn't a constant: it depends on the model, on the hardware, and on whether the previous specialization survived in the cache. Designing the UX around "the model is ready" doesn't work — a preparing state has to exist in the design.
Third, delivery size. A .aimodel in the bundle, .aimodelc files per architecture, and the specialization cache on disk are three separate copies of the same weights, each taking space. The "download → specialize → save a bookmark → delete the source" sequence exists precisely because the naive version inflates the app.
And if the model you're planning to carry through Core AI is a language model, you probably won't write NDArray inference by hand at all: you can hand it to a LanguageModelSession from Foundation Models and get the usual prompts, streaming, and structured output. How that fits together is a separate piece.



