← AI Engineering

System Check for AI: Know What Your Hardware Can Actually Run

How a browser-based AI hardware advisor uses local device detection and deterministic rules to recommend realistic models, runtimes, and projects for your machine.

·17 min read·
#Local AI#AI Hardware#AI Workloads#Model Inference#On-Device AI#AI Engineering

System Check for AI: Know What Your Hardware Can Actually Run

TL;DR

System Check for AI is a browser-based AI Hardware Workload Advisor that helps you understand what local or on-device AI workloads your hardware can realistically handle.

It detects the hardware information that the browser can safely expose, such as CPU concurrency, memory, GPU details, and platform information, and combines that with a deterministic rules engine to recommend:

  • suitable local AI models and quantization levels
  • approximate memory requirements and expected inference performance
  • tools such as Ollama, llama.cpp, LM Studio, MLX, vLLM, WebLLM, and ComfyUI
  • AI projects that are practical for that hardware
  • the capabilities and limitations of the system

When browsers cannot reliably identify a device, especially on mobile, users can explicitly select their hardware from a local device catalog instead of relying on guesses.

The important architectural choice is that hardware analysis and recommendations do not depend on an LLM.

Detection happens client-side, the recommendation logic is deterministic, and no hardware profile needs to be uploaded to a server.

The core idea is simple:

Hardware → Effective Resource Budget → AI Workload Tier → Models → Tools → Projects

Instead of asking, “Can I somehow run this model?”, the application starts with a better question:

What AI workloads actually make sense on the hardware I already have?


Introduction

Local AI is becoming increasingly accessible.

A few years ago, experimenting with capable AI models usually meant using a cloud API or having access to expensive GPU infrastructure. Today, you can run useful language models, embedding models, computer-vision workloads, speech models, and even image-generation pipelines on laptops, workstations, edge devices, and in some cases directly inside a browser or mobile device.

Tools such as Ollama, llama.cpp, LM Studio, MLX, vLLM, WebLLM, and ComfyUI have made the runtime side much easier.

But another problem has become more visible.

People know that they can run AI locally, but they often don't know what they should run on their particular hardware.

Consider a few very different systems:

Raspberry Pi 5 — 16 GB

NVIDIA Jetson — 8 GB

MacBook Pro — 16 GB unified memory

Mac Studio — 64 GB unified memory

Desktop — NVIDIA GPU with 24 GB VRAM

Cloud instance — multiple high-memory GPUs

All of them can run AI workloads.

But the type of workload, model size, quantization, runtime, expected performance, and practical project scope can be completely different.

A recommendation such as “use a 7B model” is therefore incomplete.

The useful answer depends on several questions:

  • How much memory is actually available for inference?
  • Is that memory system RAM, dedicated VRAM, or unified memory?
  • Is there an accelerator?
  • Which runtime works best with that architecture?
  • What model size fits with enough headroom for context and runtime overhead?
  • Will the model merely load, or will it actually be usable?
  • What kinds of applications make sense within that resource budget?

This is the problem I wanted System Check for AI to address.

The idea was to create a small web application that could inspect the device from the browser and translate hardware specifications into something much more useful:

Your hardware
      ↓
What it can realistically handle
      ↓
Which models fit
      ↓
Which tools to use
      ↓
What you can build

There was also an architectural constraint I wanted to maintain from the beginning:

The application should not invent hardware information and should not require an AI model to make deterministic infrastructure decisions.

Browsers intentionally expose only limited hardware information for privacy and security reasons. So when the browser knows something, the application reports it. When it can only estimate something, that is clearly marked. And when exact hardware information is unavailable, the user can select their device from a known catalog.

The recommendation engine follows the same principle.

Instead of sending the detected specification to an LLM and asking, “What can this computer run?”, the application uses explicit workload rules that can be inspected, tested, versioned, and improved.

That makes System Check for AI not just a hardware-checking utility, but also an interesting example of a broader AI engineering principle:

Building an AI-related product does not mean every decision inside the product needs an AI call.

Sometimes the better architecture is still deterministic software.

The rest of this article explains how the application works, how browser-based hardware detection is handled, how hardware is translated into workload tiers, and why I chose a rules engine rather than an LLM for the core recommendation system.

Running AI locally has become surprisingly easy.

Install Ollama. Download a model. Start chatting.

But that simplicity hides a more practical question:

What can your hardware actually run well?

A model may technically load on your laptop and still be practically unusable. A machine with 8 GB RAM has very different possibilities from a MacBook with unified memory, an NVIDIA workstation, a Jetson board, or a cloud GPU instance.

The same applies beyond LLMs.

Can the machine run:

  • a 1B or 3B language model?
  • a 7B or 14B coding model?
  • Stable Diffusion or ComfyUI?
  • a local RAG application?
  • computer-vision inference?
  • a voice assistant?
  • an embedding service?
  • multiple models simultaneously?
  • a production inference server such as vLLM?

This was the problem behind System Check for AI, an AI Hardware Workload Advisor I built to answer a simple question:

Given this hardware, what AI workloads realistically make sense?

Instead of starting with a model and trying to force it onto the machine, the application starts with the available hardware and works forward.


The problem with local AI hardware recommendations

Most AI model documentation gives us information such as:

  • parameter count
  • quantization
  • approximate memory requirements
  • supported runtimes
  • GPU compatibility

But users usually approach the problem from the opposite direction.

They already own a machine.

For example:

I have a MacBook with 16 GB memory. What can I run?

Or:

I have a Raspberry Pi 5. What meaningful AI experiments can I build?

Or:

I have an NVIDIA Jetson board. Should I run a 7B model, computer vision, or something smaller?

Or:

Which model should I use on this GPU without running out of VRAM?

These questions require more than simply listing popular models.

We need to translate:

Hardware → Resource Budget → Workload Tier → Models → Tools → Projects

That became the core architecture of the application.


A browser-based AI hardware advisor

System Check for AI is a single-page web application that inspects the visitor's device and recommends AI workloads suitable for that machine.

The important part is how it works.

There is:

No account.
No hardware information uploaded.
No server-side profiling.
No LLM deciding what hardware you have.

Hardware detection runs entirely inside the browser.

The recommendation engine is deterministic.

That distinction is intentional.

An AI workload advisor does not itself need AI for every part of the application.


Step 1: Detect what the browser actually knows

Modern browsers expose some useful hardware information.

For example:

navigator.hardwareConcurrency

can provide the number of logical CPU processors available to the browser.

Another useful property is:

navigator.deviceMemory

which can provide an approximate representation of device memory on browsers that support it.

GPU information can sometimes be inferred using WebGL or WebGPU renderer information.

Together these signals can help estimate:

  • operating system
  • device type
  • CPU concurrency
  • available memory
  • GPU vendor
  • graphics architecture
  • approximate GPU capability

But there is an important engineering constraint.

Browser hardware detection is intentionally incomplete

Browsers are privacy boundaries.

They deliberately avoid exposing an exact hardware inventory.

For example, a browser might tell us that a system appears to have an Apple GPU, but it should not be treated like:

Device = MacBook Pro M3 Max
Memory = 64 GB
GPU cores = 40

unless we actually have reliable evidence for those values.

The application therefore attaches confidence information to hardware properties.

A value can be:

detected
estimated
unavailable

This sounds like a small implementation detail.

It is actually an important design principle.

Do not turn uncertainty into fake precision.

If the browser gives us partial information, the application should expose that uncertainty instead of pretending to know more than it does.


Mobile devices make this problem even harder

Mobile browsers expose even less information.

Trying to infer whether someone is using an iPhone 15 Pro, iPhone 16, iPad Pro, or another device based purely on browser signals quickly becomes unreliable.

One approach would be to ask an LLM to guess the device.

That would make the architecture worse, not better.

Instead, System Check for AI uses a known-device catalog.

Users can search for and explicitly select their device.

The catalog can contain devices such as:

  • iPhone models
  • iPad Pro models
  • Windows laptops
  • ASUS and Dell systems
  • Apple desktops
  • Raspberry Pi boards
  • NVIDIA Jetson boards
  • AI workstations
  • AWS GPU instances
  • GCP GPU instances
  • Azure GPU instances
  • hosted GPU environments

The device catalog is versioned and stored locally.

Search uses deterministic fuzzy matching.

This creates a useful separation:

Browser detection
      ↓
Best-effort environment information

Device catalog selection
      ↓
Known hardware specification

The application never silently converts weak browser signals into a supposedly precise device specification.


Step 2: Convert hardware into an effective AI memory budget

Raw specifications alone are not enough.

Consider two machines:

Machine A
16 GB system RAM
Integrated GPU

Machine B
32 GB system RAM
8 GB dedicated GPU

Simply comparing total RAM would produce misleading recommendations.

AI workloads care about different memory domains.

Depending on the architecture, we may need to consider:

  • system RAM
  • dedicated VRAM
  • unified memory
  • model weights
  • KV cache
  • runtime overhead
  • operating-system usage
  • context length
  • batching
  • concurrent applications

System Check therefore calculates an effective workload memory budget rather than treating advertised system memory as entirely available for inference.

That budget becomes the basis for workload classification.


Step 3: Classify the machine into an AI workload tier

Instead of presenting dozens of raw hardware metrics and expecting the user to interpret them, the application classifies the hardware into a workload tier.

The current model uses:

tiny
small
medium
large
xlarge

Each tier represents a practical AI workload envelope.

Conceptually:

Hardware
   ↓
Available memory + compute capability
   ↓
Effective AI budget
   ↓
Workload tier

For example, lower tiers may be appropriate for:

  • tiny language models
  • embeddings
  • lightweight classification
  • basic vision models
  • browser inference
  • experimentation

Higher tiers gradually unlock workloads such as:

  • larger quantized LLMs
  • code assistants
  • multimodal models
  • image generation
  • local RAG
  • larger context windows
  • simultaneous model execution
  • multi-user inference servers

At workstation scale, the architecture can recommend significantly larger models and more demanding inference stacks.

The purpose of the tier is not to say:

Your computer is good or bad.

It says:

This is the realistic workload envelope for this machine.

That is a much more useful distinction.


Step 4: Recommend models that fit the hardware

Once the workload tier is known, the application can recommend specific models.

The rules engine considers factors such as:

  • model size
  • quantization
  • approximate memory requirement
  • expected runtime
  • hardware architecture
  • accelerator availability

A recommendation might look conceptually like:

Model
Llama 3.2 3B

Quantization
Q4

Expected memory
~2–4 GB+

Runtime
llama.cpp / Ollama

Expected experience
Comfortable for lightweight local inference

Larger systems can progressively move toward larger model families and more demanding inference workloads.

The application can also provide an estimated tokens-per-second range.

That number should be treated as an estimate rather than a benchmark guarantee.

Actual inference performance depends on many variables:

  • model architecture
  • quantization format
  • context size
  • runtime
  • GPU offloading
  • memory bandwidth
  • CPU architecture
  • thermal limits
  • operating system
  • concurrent workloads

So the application is deliberately an advisor, not a synthetic benchmark.


Step 5: Recommend the runtime, not just the model

One mistake in local AI discussions is treating model choice as the entire problem.

The runtime matters almost as much.

Different hardware ecosystems favor different tools.

System Check can recommend tools such as:

Ollama

Useful when someone wants a simple local model runtime with straightforward model management and API access.

LM Studio

Useful for people who prefer a graphical interface while still having the option to expose local models through an API.

llama.cpp

A strong choice when you want direct control over GGUF-based inference, quantization, CPU/GPU offloading, and low-level runtime behaviour.

MLX

Particularly interesting for Apple Silicon workloads where unified memory and Apple's compute stack can be used effectively.

vLLM

More appropriate when the objective moves beyond single-user local experimentation toward high-throughput GPU inference.

WebLLM

Interesting for workloads where inference itself can happen inside the browser.

ComfyUI

Relevant when the hardware budget supports diffusion and image-generation workflows.

Mobile runtimes

Mobile devices introduce a different ecosystem involving technologies such as:

  • Core ML
  • MLX Swift
  • mobile-specific local model applications

This creates a second useful mapping:

Hardware
      ↓
Workload
      ↓
Model
      ↓
Runtime

Choosing the right runtime can sometimes make the difference between a frustrating experiment and a useful system.


Step 6: Recommend things to build

Hardware advice becomes more useful when it leads to an experiment.

So each workload tier also includes project ideas.

For example, lower-power systems might start with:

  • text classification
  • embedding generation
  • semantic search
  • lightweight voice commands
  • object detection
  • tiny local assistants

A medium system might move into:

  • local RAG
  • private document assistants
  • code assistants
  • local speech pipelines
  • multimodal experiments
  • development copilots

More capable systems can explore:

  • multi-model pipelines
  • large-context RAG
  • image generation
  • video analysis
  • agentic applications
  • local AI platforms
  • multi-user inference services
  • AI development environments

This changes the question from:

Which model can I download?

to:

What useful system can I build with the compute I already have?

That is the more interesting problem.


Deterministic recommendations instead of an AI call

One architectural decision in this project deserves particular attention.

The application already contains an AI Gateway integration.

But the core hardware recommendation process does not depend on an LLM.

That is deliberate.

There are situations where AI is valuable.

There are also situations where a deterministic function is the better engineering choice.

Hardware classification is one of them.

Given the same input:

RAM
GPU
VRAM
device architecture

the application should ideally produce the same workload classification every time.

A simplified rule might resemble:

if (effectiveMemory < 8) {
  return "tiny";
}

if (effectiveMemory < 16) {
  return "small";
}

if (effectiveMemory < 32) {
  return "medium";
}

if (effectiveMemory < 64) {
  return "large";
}

return "xlarge";

The real rules can of course account for more variables.

But the principle remains:

Known inputs
     ↓
Explicit rules
     ↓
Reproducible recommendation

Instead of:

Known inputs
     ↓
LLM prompt
     ↓
Probabilistic recommendation

For this use case, deterministic behaviour gives us:

  • reproducibility
  • testability
  • explainability
  • lower latency
  • zero inference cost
  • offline capability
  • easier debugging

It also makes model recommendations easier to audit and update.


AI applications do not need AI everywhere

This project reinforced a design principle I increasingly find useful.

Use AI where uncertainty and semantic reasoning justify it. Use deterministic software where rules are sufficient.

It is easy to reach for an LLM simply because we are building an AI-related product.

But many components around an AI system are better implemented conventionally.

For System Check:

Deterministic

  • hardware detection
  • device matching
  • workload classification
  • memory calculations
  • compatibility rules
  • model filtering
  • confidence labels

Potentially AI-assisted

AI can later become useful for things such as:

  • explaining why a particular model fits
  • creating learning paths
  • generating project plans
  • suggesting workload optimizations
  • comparing alternative hardware
  • answering follow-up technical questions

That produces a more disciplined architecture.

AI becomes a reasoning layer rather than an unnecessary dependency.


Privacy was part of the architecture

Hardware fingerprinting deserves careful handling.

The application therefore follows a straightforward model:

Browser
   ↓
Hardware detection
   ↓
Rules engine
   ↓
Recommendation

Nothing needs to leave the device.

There is no requirement to create an account.

Hardware information does not need to be uploaded for classification.

Known device specifications come from the local catalog.

That makes the application particularly suitable for a hardware-detection workflow where users may reasonably prefer not to send device information elsewhere.


Manual override is still necessary

Automatic detection should never become a trap.

Users may know more about their hardware than the browser does.

For example:

  • exact RAM capacity
  • external GPU configuration
  • exact GPU model
  • actual VRAM
  • upgraded hardware
  • custom workstation configuration
  • virtual machines

So detected values can be edited.

The workflow becomes:

Detect
   ↓
Review
   ↓
Correct if required
   ↓
Classify
   ↓
Recommend

That combination of automatic detection and user override usually produces more trustworthy results than either approach alone.


Exporting the recommendation

Users can also export the result as a PDF.

The report contains information such as:

  • detected system specification
  • confidence level
  • hardware workload tier
  • model recommendations
  • quantization information
  • recommended runtimes
  • capabilities
  • limitations

PDF generation happens client-side using jsPDF.

This turns a temporary browser analysis into something that can be kept as a hardware capability report.

It can also be useful when comparing systems before deciding whether a hardware upgrade is actually necessary.


The application architecture

The application is built using:

TanStack Start
React 19
TypeScript
Vite 7
Tailwind CSS 4
shadcn/ui
Bun
jsPDF

The important application logic is intentionally separated into three components.

detect-hardware.ts
       ↓
What does the browser know?

device-catalog.ts
       ↓
What do we know about explicitly selected devices?

hardware-rules.ts
       ↓
What workloads fit those resources?

That separation makes the architecture easier to reason about.

┌───────────────────────────┐
│        Browser UI         │
│     TanStack + React      │
└─────────────┬─────────────┘
              │
              ▼
┌───────────────────────────┐
│    Hardware Detection     │
│ CPU / RAM / GPU / OS      │
└─────────────┬─────────────┘
              │
        uncertain device?
              │
              ▼
┌───────────────────────────┐
│      Device Catalog       │
│ Known hardware profiles   │
└─────────────┬─────────────┘
              │
              ▼
┌───────────────────────────┐
│       Rules Engine        │
│ Effective resource budget │
└─────────────┬─────────────┘
              │
              ▼
┌───────────────────────────┐
│      Workload Tier        │
│ tiny → small → medium →   │
│ large → xlarge            │
└─────────────┬─────────────┘
              │
              ▼
┌───────────────────────────┐
│      Recommendations      │
│ Models                    │
│ Tools                     │
│ Projects                  │
│ Capabilities              │
│ Limitations               │
└───────────────────────────┘

The architecture is deliberately simple.

The complexity lives in the knowledge encoded into the rules and device catalog.


The interesting challenge: keeping recommendations current

Hardware detection itself is only one part of the problem.

The harder long-term problem is maintaining the recommendation knowledge base.

The local AI ecosystem changes quickly.

New developments continuously affect:

  • model sizes
  • quantization formats
  • inference runtimes
  • Apple Silicon support
  • GPU support
  • mobile inference
  • browser inference
  • context-window requirements
  • model efficiency

That means the rules engine and device catalog need to be versioned like any other technical knowledge base.

This is one reason I prefer making these rules explicit.

A rule can be reviewed.

A model entry can be updated.

A device can be added.

A test can validate the expected classification.

The behaviour does not disappear inside a prompt.


Testing the advisor

A deterministic architecture also makes testing much easier.

For example:

Given:
8 GB RAM
CPU-only

Expected:
tiny/small workload class
small quantized model recommendations
no high-memory GPU workload recommendations

Another fixture might represent:

Given:
Apple Silicon
32 GB unified memory

Expected:
larger local models
MLX-compatible recommendations
higher workload class

Another could represent:

Given:
NVIDIA GPU
24 GB VRAM

Expected:
GPU-oriented runtimes
larger quantized models
image-generation capability
higher-throughput inference options

Device catalog matching can be tested independently from hardware classification.

This is important because AI workload advice should be predictable enough to regression-test.


Where I want to take it next

There are several directions in which this can grow.

Better mobile device coverage

Mobile hardware detection is inherently restricted by browsers, so explicit device selection will become increasingly important.

The catalog can expand across:

  • iPhone generations
  • iPads
  • Android devices
  • mobile NPUs

Better accelerator awareness

Future versions can reason more deeply about:

  • Apple Neural Engine
  • NVIDIA Tensor Cores
  • integrated NPUs
  • AMD accelerators
  • Intel AI hardware

Instead of treating everything primarily as CPU/GPU memory.


Context-length-aware recommendations

A 7B model running with a small context window and the same model running with a very large context window are not equivalent workloads.

Future recommendations can incorporate:

Model weights
+
KV cache
+
Context requirement
+
Runtime overhead

to produce a more realistic memory envelope.


Workload-specific modes

Eventually it may be more useful to ask the user what they want to build.

For example:

Local chatbot
Code assistant
RAG
Computer vision
Image generation
Voice agent
AI agent
Embedding server

The advisor could then evaluate the hardware specifically against that workload.


Hardware comparison

Another useful direction is to compare:

My current machine
        vs
Potential upgrade
        vs
Cloud GPU

The application could explain what new AI workload classes become available with each option.

This could help users avoid upgrading simply because a new GPU or laptop appears attractive.


A small project with a broader idea

System Check for AI started from a simple practical problem:

People are experimenting with local AI, but hardware recommendations are fragmented and often confusing.

The larger idea is more interesting.

As AI starts running across increasingly diverse environments—

  • browsers
  • phones
  • laptops
  • Apple Silicon
  • Raspberry Pis
  • Jetson devices
  • workstations
  • cloud GPUs
  • edge systems

—we need better ways to map software capability to available compute.

The common question should no longer be:

What is the most powerful AI model available?

A more useful question is:

What is the right AI workload for the compute available to me?

That is what System Check for AI is trying to answer.

And because all of the core analysis runs locally and deterministically, the system itself follows the same principle it recommends:

Use only as much compute—and as much AI—as the problem actually requires.

Related reading: Understanding GenAI Hardware: CPU, GPU, NPU, Inference, and Model Serving.

Public profile lookup

Ask AI About the Author

Open this query in ChatGPT, Claude, or Perplexity.

Comments

Comments are open to confirmed email subscribers. Use the email you subscribed with. To edit a comment, delete it and post a new one.

0/2000
Verify:

    Subscribe to get the new blogs.

    Field notes from someone who ships before they write about it. Sovereign AI, AI-SDLC, DevOps, and what 59 production deployments teach you. No spam. Unsubscribe anytime.

    Related field notes