Scope: Recent public materials from Apple, Microsoft, NVIDIA, and AMD on edge / personal / cloud AI, NPU and GPU compute, unified memory, model routing, privacy, and local AI.

Research date: 2026-08-09. Vendor performance figures depend on model, quantization, software, power, and configuration; they should not be treated as universal benchmarks.

1. Executive Summary

The view that the future will combine edge and cloud computing is directionally correct. However, “simple tasks run locally and complex tasks run in the cloud” is too rigid for the long term.

A more accurate model is:

AI applications will use a policy-driven, tiered execution architecture. Device, personal-compute, and cloud models will form one AI runtime, with execution determined by capability, latency, privacy, energy, memory, connectivity, cost, context location, freshness, and reliability.

                 AI APPLICATION / AGENT
                           │
                    Task / Policy Router
                           │
        ┌──────────────────┼──────────────────┐
        ▼                  ▼                  ▼
   DEVICE / EDGE      PERSONAL COMPUTE       CLOUD
   Phone / Watch      Phone / AI PC / Hub      Frontier model
   Always-on          Larger local model     Long context
   Sensor-native      Private context        Tools / search
   Low latency        Offline capability     Shared knowledge
   Privacy-first      High memory capacity   Elastic compute

Key conclusions:

  1. NPU is becoming a client-platform capability, not merely a marketing feature. Copilot+ PCs made 40+ TOPS NPUs a defining product requirement, while Apple, Microsoft, AMD, NVIDIA, and Qualcomm are building software layers around local inference.
  2. Unified memory addresses model capacity and data movement, not just raw compute. Apple Silicon, NVIDIA DGX Spark, and AMD Ryzen AI Max/Halo all point toward tighter integration of CPU, GPU/NPU, and large shared memory.
  3. Local AI will usually be provided as a shared OS/runtime capability, not as a separate large model bundled inside every application. Apple Foundation Models and Microsoft Foundry Local/Windows ML are examples.
  4. Cloud AI will remain essential, but increasingly as an escalation layer. Frontier reasoning, very long context, web search, tools, cross-user knowledge, centralized training, and elastic capacity remain naturally cloud-oriented.
  5. Apple’s distinctive contribution is a system-level loop: personal-context selection → local model → Private Cloud Compute escalation → verifiable privacy → developer APIs.

2. A Better Model for Choosing the Execution Location

The execution target should be understood as a function of several variables:

Execution Target = f(
  capability,
  latency,
  privacy,
  energy,
  memory,
  connectivity,
  cost,
  context locality,
  freshness,
  reliability,
  tool requirements
)
VariableFavors local executionFavors cloud execution
LatencyReal-time, interactive, always-onLatency is acceptable
PrivacyPersonal health, photos, messages, sensorsData can be safely minimized or anonymized
CapabilityStable, narrow, structured tasksOpen-ended reasoning and multi-tool workflows
ContextData is already on the deviceSearch, cross-device, or shared knowledge is needed
Energy / costHigh-frequency repeated callsLow-frequency, high-value calls
MemoryModel fits available local memoryModel or context exceeds local capacity
ReliabilityMust work offlineNeeds the newest model or centralized governance

The resulting application abstraction is increasingly likely to be:

App capability
      │
      ▼
AI runtime / policy router
      │
      ├── deterministic code / DSP
      ├── device NPU / local model
      ├── personal-computer GPU / local LLM
      ├── private cloud model
      └── frontier API + search + tools

3. Apple: From On-Device Intelligence to Private Cloud Intelligence

3.1 Publicly documented Apple Intelligence flow

Apple’s public description is essentially: process on device whenever possible; when a request needs more capacity, send only the relevant data to Private Cloud Compute (PCC).

User request + personal context
             │
             ▼
     On-device orchestration
             │
     Can local model handle it?
          ┌──┴──┐
          │     │
         yes    no / needs more capacity
          │     │
          ▼     ▼
   On-device   Minimal relevant context
   model       + request + model parameters
                        │
                        ▼
             Attested Private Cloud Compute
                        │
                        ▼
                 Larger server model
                        │
                        ▼
             Secure response to the device

Apple states that the device first analyzes whether the task can be completed locally. If a larger server model is required, only data relevant to the request is sent to PCC; the request and response are not retained after processing. See Apple Intelligence & Privacy.

3.2 Layered architecture

┌─────────────────────────────────────────────────────────────┐
│ Product surfaces: Siri, Writing Tools, Photos, Messages,    │
│ notifications, Visual Intelligence, app actions, etc.       │
└──────────────────────────────┬──────────────────────────────┘
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Personal-context orchestration                              │
│ - Understand current activity and app context               │
│ - Select only data necessary for the request                 │
│ - Apply permissions, confirmation, and safety rules           │
└──────────────────────────────┬──────────────────────────────┘
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Model layer                                                 │
│ - On-device Apple Foundation Model                          │
│ - Specialized models and adapters                            │
│ - Private Cloud Compute server model                         │
│ - Optional third-party model with user consent               │
└──────────────────────────────┬──────────────────────────────┘
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Execution and hardware                                      │
│ Apple Silicon CPU / GPU / Neural Engine / unified memory     │
│ Apple Silicon servers for PCC                               │
└──────────────────────────────┬──────────────────────────────┘
                               ▼
┌─────────────────────────────────────────────────────────────┐
│ Privacy and trust                                           │
│ Secure boot, attestation, signed software, statelessness,    │
│ no privileged runtime access, transparency log               │
└─────────────────────────────────────────────────────────────┘

The “personal-context orchestration” layer is an analytical abstraction based on Apple’s public descriptions of data selection, device-side evaluation, and PCC request handling. It is not a public Apple component name.

3.3 Model family: compact local model plus larger server models

Apple’s 2024 technical report described two core language-model categories: an approximately 3B-parameter on-device model and a larger server model running on PCC. See Apple Intelligence Foundation Language Models.

Apple’s 2025 update disclosed several efficiency techniques:

  • approximately 2-bit-per-weight quantization-aware training for on-device decoder weights;
  • ASTC-based compression for server-model weights;
  • low-bit quantization for embeddings and KV cache;
  • adapter training to recover quality after compression.

This shows that the local model is not simply a cloud model copied onto a phone. It is co-designed around memory capacity, bandwidth, power, and latency. See Updates to Apple’s On-Device and Server Foundation Language Models.

In 2026, Apple described a third-generation family spanning on-device and PCC server models, including a sparsely activated on-device model and cloud models intended for complex reasoning and agentic tool use. See Introducing the Third Generation of Apple’s Foundation Models.

3.4 Apple Silicon, unified memory, and model feasibility

          Shared / Unified Memory
        ┌────────────────────────┐
        │ Model weights           │
        │ KV cache                │
        │ Activations             │
        │ OS + app context        │
        └───────────┬────────────┘
                    │ shared address space
        ┌───────────┼───────────┐
        ▼           ▼           ▼
       CPU         GPU       Neural Engine
   orchestration  tensor/    efficient neural
                  graphics   inference

Unified memory can:

  1. reduce copying between CPU RAM and GPU VRAM;
  2. let model weights, KV cache, activations, and application context share a memory pool;
  3. allow the system to allocate work across CPU, GPU, and NPU-like accelerators.

Unified memory does not mean every byte has identical bandwidth or that capacity, thermal, and power constraints disappear. Throughput still depends on bandwidth, access pattern, quantization, cache behavior, and the runtime.

3.5 Foundation Models framework: local models as a platform capability

Apple’s Foundation Models framework lets developers access the on-device large language model at the core of Apple Intelligence for summarization, entity extraction, text understanding, rewriting, structured generation, and tool calling. See Foundation Models.

The application model is shifting from:

App → self-hosted model / cloud API

to:

App → OS-provided AI runtime → on-device model

The 2026 framework updates also describe a LanguageModel protocol that can unify on-device and server models, plus PrivateCloudComputeLanguageModel for larger context and stronger reasoning. See Foundation Models updates.

Apple is therefore turning model execution and model selection from an application infrastructure problem into a system-platform abstraction.

3.6 Private Cloud Compute is not an ordinary cloud API

PCC is designed to extend parts of the device security model into the cloud. Apple’s public requirements include:

  • Stateless computation: personal data is used only to fulfill the current request and is inaccessible afterward;
  • Enforceable guarantees: critical components can be constrained and analyzed;
  • No privileged runtime access: operators cannot bypass privacy guarantees through privileged interfaces;
  • Non-targetability: an attacker cannot selectively target one user’s data without broadly compromising the system;
  • Verifiable transparency: researchers can inspect whether the system matches Apple’s public promises.

See Private Cloud Compute Security Guide and Private Cloud Compute: A new frontier for AI privacy.

Device
  │ validates PCC software / attestation
  │ sends only request-relevant data
  ▼
PCC node
  │ signed and verified software
  │ no privileged access
  │ stateless inference
  ▼
Response returned; personal request data is not retained

3.7 Third-party models and system-level routing

Apple Intelligence can also invoke third-party models, such as ChatGPT, when the user grants permission. This indicates that Apple’s platform role is not simply to provide one Apple model. It controls:

  • which model receives a task;
  • when the user must confirm;
  • which context may be shared;
  • how the result returns to the system experience;
  • which safety and privacy policies apply to external models.

Apple Intelligence is therefore closer to a context-aware model orchestration layer than to a single LLM product.

4. Microsoft: Windows as a Local-AI Development Platform

Apple emphasizes personal context, integrated device/cloud privacy, and system experiences. Microsoft emphasizes a cross-vendor Windows runtime and developer toolchain.

4.1 Three layers of the Windows AI stack

Windows applications
          │
          ├── Windows AI APIs
          │   Built-in models / Phi Silica / OCR / imaging
          │
          ├── Foundry Local
          │   Ready-to-use OSS LLMs, local runtime
          │
          └── Windows ML
              Custom ONNX models
              Execution providers for NPU / GPU / CPU

Microsoft’s Windows AI documentation places on-device models, NPU acceleration, and cloud APIs in one developer framework. Foundry Local runs open-source LLMs locally, while Windows ML deploys custom ONNX models. See Windows AI.

Foundry Local can select hardware-specific model variants from a model alias—for example, QNN NPU on Snapdragon, CUDA on NVIDIA, or CPU fallback. See Get started with Foundry Local.

4.2 NPU as a platform layer

Copilot+ PCs require a 40+ TOPS NPU for the product category’s defining AI experiences and provide developer guidance for local NPU access through ONNX Runtime. See Copilot+ PCs developer guide.

The important signal is not the 40 TOPS number alone. Windows is building a unified execution abstraction across CPU, GPU, and NPU. Windows ML is designed to manage and update hardware execution providers so applications do not need to bind themselves to a single chip vendor. See What is Windows ML?.

5. NVIDIA: Personal Computers as Local AI Supercomputers

NVIDIA’s core advantages remain GPUs, Tensor Cores, CUDA, and a complete software stack. Its personal-compute direction is changing in two ways:

  1. workstation GPUs with larger memory support local generative and agentic AI;
  2. Grace Blackwell systems combine CPU, GPU, and unified memory into desktop AI nodes.

5.1 DGX Spark

NVIDIA DGX Spark uses Grace Blackwell and publicly lists:

  • 20-core Arm CPU;
  • Blackwell GPU and fifth-generation Tensor Cores;
  • 128GB LPDDR5x unified system memory;
  • 273GB/s memory bandwidth;
  • support for local development and inference workloads involving models up to approximately 200B parameters.

Sources: DGX Spark Hardware Overview and DGX Spark System Overview.

Grace CPU + Blackwell GPU
           │
           ▼
   128GB coherent unified memory
           │
           ▼
 CUDA / cuDNN / containers / NGC
           │
           ▼
 Local inference, fine-tuning, agents

The strategic signal is that a PC can become a personal AI compute node with large memory, meaningful local inference, and a complete development stack—not merely a display and input device for cloud AI.

5.2 NVIDIA and Microsoft’s personal-agent direction

NVIDIA and Microsoft’s 2026 RTX Spark materials position Windows PCs as local execution environments for personal agents, emphasizing unified memory, Windows-native agents, and local large-model execution. See NVIDIA and Microsoft Reinvent Windows PCs for the Age of Personal AI.

The model-size and context-length figures in vendor announcements depend heavily on quantization, software, context size, and configuration. They are best read as capability-boundary signals rather than universal user-experience guarantees.

6. AMD: Unified Memory, NPU, and an Open Software Stack

AMD emphasizes AI PCs and developer systems with large shared memory, GPU, NPU, and ROCm across Windows and Linux.

6.1 Ryzen AI Max

AMD has publicly described Ryzen AI Max with:

  • up to 128GB unified memory;
  • up to 96GB available to graphics/AI workloads;
  • RDNA 3.5 GPU;
  • XDNA 2 NPU with up to 50 TOPS.

Source: AMD expanded consumer and commercial AI PC portfolio.

This pushes the boundary between thin AI PCs and local-model workstations: a system does not always need large discrete VRAM to carry a large model working set.

6.2 Ryzen AI Halo

AMD Ryzen AI Halo targets AI developers and agentic AI with:

  • 128GB LPDDR5x unified memory;
  • a positioning around local development and inference for models up to approximately 200B parameters;
  • Windows or Linux;
  • ROCm software;
  • up to 50 TOPS NPU.

Sources: AMD Ryzen AI Halo and AMD Ryzen AI Halo announcement.

AMD’s materials also point toward higher unified-memory ceilings in future products. These should be understood as product-roadmap and capability-boundary signals, not as a guarantee that every AI PC will routinely run a 200B model.

7. Comparison Across the Four Vendors

VendorPrimary control pointEdge / personal-compute directionCloud directionArchitectural signal
AppleOS, devices, personal context, privacy protocolApple Silicon, Neural Engine, Foundation ModelsPCC, Apple Silicon servers, third-party routingLocal-first plus verifiable private cloud
MicrosoftWindows OS, developer runtime, enterprise cloudCopilot+ PC, Windows AI APIs, Foundry Local, Windows MLAzure / Foundry cloud servicesCross-chip unified local-AI platform
NVIDIAGPU, CUDA, AI software ecosystemRTX workstations, DGX Spark, RTX SparkData-center GPUs and AI infrastructureGPU-first plus high-memory personal nodes
AMDCPU/GPU/NPU SoCs, ROCm, OEM platformRyzen AI Max/Halo, 128GB-class unified memoryInstinct GPUs and data-center ecosystemOpen platform plus unified memory and NPU

The common direction is:

AI hardware
  = CPU + GPU + NPU + large memory + high bandwidth
 
AI software
  = shared runtime + model packaging + hardware-aware routing
 
AI product
  = local default + cloud escalation + privacy policy

8. Will Local AI Become Standard?

High-probability developments

  • New phones, PCs, and wearables will continue to include dedicated AI accelerators.
  • Operating systems will provide shared local models and inference APIs.
  • High-frequency, low-latency, privacy-sensitive tasks will move increasingly local.
  • Personal computers will run larger local models and agents.
  • Applications will use a runtime that selects CPU, GPU, NPU, personal compute, or cloud.
  • Personal context will increasingly remain local, with only the minimum relevant projection sent to the cloud.

Claims that should be treated cautiously

  • Not every application will bundle a large independent model.
  • Not every complex task must run in the cloud.
  • NPU TOPS alone does not predict LLM experience.
  • Unified memory does not remove bandwidth, thermal, or power constraints.
  • Local inference will not replace frontier cloud models, search, training, shared knowledge, or centralized governance.
  • A vendor’s maximum model-size claim does not equal a broadly available, high-quality agent experience.

9. Limited Implications for PawsGenie

PawsGenie does not need to make a general-purpose Pet Health foundation model the center of its strategy. A more durable architecture is:

                  PAWSGENIE AI RUNTIME
                         │
                  Task / Policy Router
                         │
       ┌─────────────────┼─────────────────┐
       ▼                 ▼                 ▼
   Device / Edge     Personal Compute      Cloud
   sensor events     private pet context   complex reasoning
   event gating      local summaries       research / tools
   basic classifier  offline interaction   longitudinal synthesis

PawsGenie should focus on owning:

  • Pet Identity;
  • raw Observations and provenance;
  • pet-specific baselines;
  • longitudinal Health Context;
  • Health Episodes;
  • routing policies, evaluation, and safety boundaries.

Models can be replaceable across tiers: device models, OS models, open-source local models, PawsGenie models, and external frontier models.

The durable advantage is not binding the product to one model. It is allowing different models at different tiers to operate on one trusted, continuous, traceable Pet Health Context.

10. Research Conclusions

  1. “Edge + cloud” is the right direction, but the likely end state is hierarchical AI compute, with personal-compute nodes becoming increasingly important.
  2. Apple’s significance is not merely that it has a local model and PCC. It combines model capability, context selection, system APIs, hardware co-design, and verifiable privacy into one platform.
  3. Microsoft is turning local AI into a cross-hardware Windows development infrastructure.
  4. NVIDIA and AMD are turning personal computers into local AI nodes; unified memory is an important way to carry large model working sets.
  5. The core application abstraction will shift from “call a particular cloud model” to “submit a task and let the AI runtime decide where, with which model, and under what privacy policy it should run.”
  6. PawsGenie should prioritize cross-tier context, provenance, policy, and evaluation layers rather than committing prematurely to one model or cloud provider.

11. First-Party Source Index

Apple

Microsoft

NVIDIA

AMD