The hard limit on Foundation Models in iOS 26 was never model quality — it was arithmetic: 4096 tokens covering everything, instructions, tool definitions, and the entire conversation transcript. In iOS 27 the same session gained a mode with 32K and noticeably stronger reasoning, and it switches on in one line. What it costs isn't money or an API key, and that part is less familiar.
I wrote the base guide to the framework back in August against iOS 26 — SystemLanguageModel, @Generable, tool calling. This piece covers only the delta: what landed by September 2026 and which product decisions it changes.
Versions: everything below is iOS 27, iPadOS 27, macOS 27, and visionOS 27 unless noted. One exception worth flagging: watchOS only got Foundation Models in 27 — on 26 the framework wasn't there.
The framework is no longer about one Apple model#
A year ago the documentation described this as access to the Apple Intelligence model. The first line of the overview now reads differently: access to any large language model — on-device, server, or one you bring yourself. There's a LanguageModel protocol, and all three are first-class citizens of the same API.
In practice: LanguageModelSession, Instructions, Tool, @Generable, streaming — you write those once, and what changes is the model you plug in. Three of those models follow.
Private Cloud Compute: same session, higher ceiling#
The switch looks like this:
// Create a session with the server-side model.
let session = LanguageModelSession(model: PrivateCloudComputeLanguageModel())Everything else carries over untouched: respond, streaming, tools, instructions. Both SystemLanguageModel and PrivateCloudComputeLanguageModel conform to LanguageModel, so the session initializer is the same one.
What you get: a 32K context and stronger reasoning, for long documents and extended multi-turn conversations. What you lose: offline. PCC needs a network, and when a request fails on connectivity Apple's own advice is to retry it on the on-device model — meaning you write the fallback regardless.
Availability is checked separately, for its own reasons:
let model = PrivateCloudComputeLanguageModel()
switch model.availability {
case .available:
// Show your intelligence UI.
case .unavailable(.deviceNotEligible):
// Show an alternative UI.
case .unavailable(.systemNotReady):
// PCC isn't ready to serve requests.
case .unavailable(let other):
// The model is unavailable for an unknown reason.
}Plus an #available check for iOS 27 / macOS 27 / watchOS 27 / visionOS 27, falling back to the on-device model on earlier versions.
And here's the unfamiliar payment. To develop against PCC you have to meet Apple's eligibility requirements and request the managed entitlement com.apple.developer.private-cloud-compute. This isn't a checkbox in Capabilities — it's an application. If you're planning a PCC feature, budget calendar time for access, or the demo will be beautiful code that doesn't run.
The quota belongs to the user, not to you#
The familiar server-LLM model: you pay for tokens, the user never learns tokens exist. PCC inverts it. There's no authentication and no keys at all — instead, every person gets a daily request limit, extended by upgrading their iCloud+ subscription.
Architecturally that's a reversal. You can't buy more headroom on a user's behalf, and you can't predict how much they have left when your screen opens. So the framework exposes quota state, and it belongs in the UI rather than behind a dismissible alert:
let model = PrivateCloudComputeLanguageModel()
if model.quotaUsage.isLimitReached {
Text("Usage limit exceeded")
.foregroundStyle(Color.red)
} else if case .belowLimit(let info) = model.quotaUsage.status {
if info.isApproachingLimit {
Text("Nearing usage limit")
.foregroundStyle(Color.orange)
}
}
if let suggestion = model.quotaUsage.limitIncreaseSuggestion {
Button("Show options") {
suggestion.show()
}
}limitIncreaseSuggestion.show() presents the system upgrade UI — you don't build your own pricing screen, and from Apple's phrasing, you're not meant to.
Exhausting the quota arrives as its own error, distinct from rate limiting: with rate limiting a person waits, with an exhausted quota they wait for a reset or upgrade. The framework exposes the reset date, though it can be empty.
You can test this without burning a real limit: Product > Scheme > Edit Scheme > Run > Options > Simulated Apple Foundation Models Availability, which offers "Approaching Quota Usage Limit" and "Quota Usage Limit Reached."
Reasoning became a dial you turn#
let response = try await session.respond(
to: "What are the tradeoffs in this architecture?",
contextOptions: ContextOptions(reasoningLevel: .deep)
)There are three levels; the ends are .light and .deep. Deeper reasoning costs latency, and it also spends more of the context window on the model's own reasoning text. The second effect bites harder in practice: you turned on .deep for quality and got exceededContextWindowSize on the third turn of a conversation.
The reasoning segments themselves stay out of the response content — they're visible in the transcript, which helps when you're working out why the model said something strange.
Apple's advice is to start at the lowest level and raise it based on evaluation. Dull advice, but it's backed now: the framework gained prompt evaluation tooling and a dedicated Foundation Models instrument showing latency, prompts, model output, tool calls, and token usage.
Images in the prompt#
Multimodal input arrived in ordinary respond:
func compareImages(imageOne: CGImage, imageTwo: CGImage) async throws -> String {
let session = LanguageModelSession()
let response = try await session.respond {
"Compare these two images by using three bullet points:"
Attachment(imageOne)
// When the image doesn't have a rotation applied — a frame from
// AVFoundation, say — pass the orientation and the framework
// performs the transform.
Attachment(imageTwo, orientation: .right)
}
return response.content
}Scaling and color conversion are handled for you; no preprocessing required. It accepts CGImage, raw data, and file URLs (the type is inferred from the UTType).
Pairing this with @Generable works better than free text. Classification in a single call:
@Generable
enum ImageLabel {
case cat
case dog
case frog
case bird
}
func classifyImage(_ image: CGImage) async throws -> ImageLabel {
let session = LanguageModelSession()
let response = try await session.respond(
generating: ImageLabel.self,
options: GenerationOptions(samplingMode: .greedy)
) {
"Choose the label that best represents the following image:"
Attachment(image)
}
return response.content
}.greedy earns its place here: without it the model may pick a label that's merely close, which in a classifier is a silent bug.
There are also ready-made image tools — barcode reading (BarcodeReaderTool) and text extraction. With several tools in play, label the image — Attachment(image).label("barcode-image") — so the model knows what each tool applies to.
A small prompting detail with an outsized effect: "List all food items in this photo" beats "What's in this image?" by a wide margin.
Dynamic profiles: instructions that rebuild before every request#
A session used to freeze its instructions at creation. For a multi-step flow — find a recipe, substitute ingredients, check the pantry, walk through cooking — that left you with either one bloated instruction covering every case, or recreating the session and losing the history.
Now the body of DynamicInstructions is re-evaluated before each model request:
struct PresentationInstructions: DynamicInstructions {
var isEditingImage = true
var isEditingAnimation = false
var body: some DynamicInstructions {
// The part that holds for any state.
Instructions {
"Help people improve their presentation."
}
ListPhotosTool()
AddPhotoTool()
// Whatever the app's current state actually needs.
if isEditingImage {
ImageEditingInstructions()
}
if isEditingAnimation {
AnimationEditingInstructions()
}
}
}
let session = LanguageModelSession(
dynamicInstructions: PresentationInstructions()
)A level up sits Profile, binding instructions to session settings, and DynamicProfile, which switches between profiles. The switch is compiler-checked: exactly one profile may be active, so the branches are written as if / else if / else rather than parallel if blocks.
Profile {
// Custom instructions and tools for a creative task.
}
.model(pccModel)
.temperature(likesPoetry ? 0.8 : 0.1)
.reasoningLevel(likesAstronomy ? .deep : .light)Modifiers resolve across three tiers: options passed directly to respond(to:options:) beat everything; a subprofile's modifier beats the dynamic profile's; the dynamic profile sets the defaults.
One rule, without which profiles make things worse#
Declaration order inside body affects performance, and it affects it a lot.
A session lays out as a token sequence: instructions first, then tool definitions, then the transcript. The provider's KV cache stays valid up to the first changed token — everything after it gets recomputed. So static instructions and tools go at the top of the body, conditional blocks at the bottom. Put a conditional first and every toggle of that flag invalidates the cache for the whole conversation.
The consequences are easy to violate with good intentions:
- Fix your tool set when you create the session. Adding a tool mid-conversation breaks the cache and works poorly: the model has already learned the pattern of earlier turns.
- When you remove a tool, strip its calls from the transcript too — otherwise the model sees references to something that no longer exists in the definitions.
- Trim the transcript rarely and in bulk. Trimming after every turn means invalidating after every turn; one consolidation near the context limit is cheaper.
- Switching profiles rewrites the whole prefix, which resets the cache entirely. Do it at natural boundaries in the flow, not on every turn.
When you know a request is at least a second or two away, prewarm(promptPrefix:) computes the prefix into the cache before it arrives.
All of this is measurable in Instruments: divide cached input tokens by total input tokens. A low ratio between turns means the cache is being thrown away and the model is chewing through the prefix every time.
What to do with all this#
The order I'd keep today: start on the on-device model and evaluate the feature quantitatively. If it hits the 4096-token wall or the quality of reasoning, add PCC — having requested the entitlement early and designed the quota state. Bring in profiles when the flow genuinely has several modes, not because the API is new.
There's a third option that isn't an Apple model at all. The LanguageModel protocol is open, and a model exported through Core AI drops into the same LanguageModelSession: that cuts the dependency on Apple Intelligence and runs on devices where it isn't available.



