System Check for AI: Know What Your Hardware Can Actually Run
How a browser-based AI hardware advisor uses local device detection and deterministic rules to recommend realistic models, runtimes, and projects for your machine.
System Check for AI: Know What Your Hardware Can Actually Run
TL;DR
System Check for AI is a browser-based AI Hardware Workload Advisor that helps you understand what local or on-device AI workloads your hardware can realistically handle.
It detects the hardware information that the browser can safely expose, such as CPU concurrency, memory, GPU details, and platform information, and combines that with a deterministic rules engine to recommend:
- suitable local AI models and quantization levels
- approximate memory requirements and expected inference performance
- tools such as Ollama, llama.cpp, LM Studio, MLX, vLLM, WebLLM, and ComfyUI
- AI projects that are practical for that hardware
- the capabilities and limitations of the system
When browsers cannot reliably identify a device, especially on mobile, users can explicitly select their hardware from a local device catalog instead of relying on guesses.
The important architectural choice is that hardware analysis and recommendations do not depend on an LLM.
Detection happens client-side, the recommendation logic is deterministic, and no hardware profile needs to be uploaded to a server.
The core idea is simple:
Hardware → Effective Resource Budget → AI Workload Tier → Models → Tools → Projects
Instead of asking, “Can I somehow run this model?”, the application starts with a better question:
What AI workloads actually make sense on the hardware I already have?
Introduction
Local AI is becoming increasingly accessible.
A few years ago, experimenting with capable AI models usually meant using a cloud API or having access to expensive GPU infrastructure. Today, you can run useful language models, embedding models, computer-vision workloads, speech models, and even image-generation pipelines on laptops, workstations, edge devices, and in some cases directly inside a browser or mobile device.
Tools such as Ollama, llama.cpp, LM Studio, MLX, vLLM, WebLLM, and ComfyUI have made the runtime side much easier.
But another problem has become more visible.
People know that they can run AI locally, but they often don't know what they should run on their particular hardware.
Consider a few very different systems:
Raspberry Pi 5 — 16 GB
NVIDIA Jetson — 8 GB
MacBook Pro — 16 GB unified memory
Mac Studio — 64 GB unified memory
Desktop — NVIDIA GPU with 24 GB VRAM
Cloud instance — multiple high-memory GPUs
All of them can run AI workloads.
But the type of workload, model size, quantization, runtime, expected performance, and practical project scope can be completely different.
A recommendation such as “use a 7B model” is therefore incomplete.
The useful answer depends on several questions:
- How much memory is actually available for inference?
- Is that memory system RAM, dedicated VRAM, or unified memory?
- Is there an accelerator?
- Which runtime works best with that architecture?
- What model size fits with enough headroom for context and runtime overhead?
- Will the model merely load, or will it actually be usable?
- What kinds of applications make sense within that resource budget?
This is the problem I wanted System Check for AI to address.
The idea was to create a small web application that could inspect the device from the browser and translate hardware specifications into something much more useful:
Your hardware
↓
What it can realistically handle
↓
Which models fit
↓
Which tools to use
↓
What you can build
There was also an architectural constraint I wanted to maintain from the beginning:
The application should not invent hardware information and should not require an AI model to make deterministic infrastructure decisions.
Browsers intentionally expose only limited hardware information for privacy and security reasons. So when the browser knows something, the application reports it. When it can only estimate something, that is clearly marked. And when exact hardware information is unavailable, the user can select their device from a known catalog.
The recommendation engine follows the same principle.
Instead of sending the detected specification to an LLM and asking, “What can this computer run?”, the application uses explicit workload rules that can be inspected, tested, versioned, and improved.
That makes System Check for AI not just a hardware-checking utility, but also an interesting example of a broader AI engineering principle:
Building an AI-related product does not mean every decision inside the product needs an AI call.
Sometimes the better architecture is still deterministic software.
The rest of this article explains how the application works, how browser-based hardware detection is handled, how hardware is translated into workload tiers, and why I chose a rules engine rather than an LLM for the core recommendation system.
Running AI locally has become surprisingly easy.
Install Ollama. Download a model. Start chatting.
But that simplicity hides a more practical question:
What can your hardware actually run well?
A model may technically load on your laptop and still be practically unusable. A machine with 8 GB RAM has very different possibilities from a MacBook with unified memory, an NVIDIA workstation, a Jetson board, or a cloud GPU instance.
The same applies beyond LLMs.
Can the machine run:
- a 1B or 3B language model?
- a 7B or 14B coding model?
- Stable Diffusion or ComfyUI?
- a local RAG application?
- computer-vision inference?
- a voice assistant?
- an embedding service?
- multiple models simultaneously?
- a production inference server such as vLLM?
This was the problem behind System Check for AI, an AI Hardware Workload Advisor I built to answer a simple question:
Given this hardware, what AI workloads realistically make sense?
Instead of starting with a model and trying to force it onto the machine, the application starts with the available hardware and works forward.
The problem with local AI hardware recommendations
Most AI model documentation gives us information such as:
- parameter count
- quantization
- approximate memory requirements
- supported runtimes
- GPU compatibility
But users usually approach the problem from the opposite direction.
They already own a machine.
For example:
I have a MacBook with 16 GB memory. What can I run?
Or:
I have a Raspberry Pi 5. What meaningful AI experiments can I build?
Or:
I have an NVIDIA Jetson board. Should I run a 7B model, computer vision, or something smaller?
Or:
Which model should I use on this GPU without running out of VRAM?
These questions require more than simply listing popular models.
We need to translate:
Hardware → Resource Budget → Workload Tier → Models → Tools → Projects
That became the core architecture of the application.
A browser-based AI hardware advisor
System Check for AI is a single-page web application that inspects the visitor's device and recommends AI workloads suitable for that machine.
The important part is how it works.
There is:
No account.
No hardware information uploaded.
No server-side profiling.
No LLM deciding what hardware you have.
Hardware detection runs entirely inside the browser.
The recommendation engine is deterministic.
That distinction is intentional.
An AI workload advisor does not itself need AI for every part of the application.
Step 1: Detect what the browser actually knows
Modern browsers expose some useful hardware information.
For example:
navigator.hardwareConcurrency
can provide the number of logical CPU processors available to the browser.
Another useful property is:
navigator.deviceMemory
which can provide an approximate representation of device memory on browsers that support it.
GPU information can sometimes be inferred using WebGL or WebGPU renderer information.
Together these signals can help estimate:
- operating system
- device type
- CPU concurrency
- available memory
- GPU vendor
- graphics architecture
- approximate GPU capability
But there is an important engineering constraint.
Browser hardware detection is intentionally incomplete
Browsers are privacy boundaries.
They deliberately avoid exposing an exact hardware inventory.
For example, a browser might tell us that a system appears to have an Apple GPU, but it should not be treated like:
Device = MacBook Pro M3 Max
Memory = 64 GB
GPU cores = 40
unless we actually have reliable evidence for those values.
The application therefore attaches confidence information to hardware properties.
A value can be:
detected
estimated
unavailable
This sounds like a small implementation detail.
It is actually an important design principle.
Do not turn uncertainty into fake precision.
If the browser gives us partial information, the application should expose that uncertainty instead of pretending to know more than it does.
Mobile devices make this problem even harder
Mobile browsers expose even less information.
Trying to infer whether someone is using an iPhone 15 Pro, iPhone 16, iPad Pro, or another device based purely on browser signals quickly becomes unreliable.
One approach would be to ask an LLM to guess the device.
That would make the architecture worse, not better.
Instead, System Check for AI uses a known-device catalog.
Users can search for and explicitly select their device.
The catalog can contain devices such as:
- iPhone models
- iPad Pro models
- Windows laptops
- ASUS and Dell systems
- Apple desktops
- Raspberry Pi boards
- NVIDIA Jetson boards
- AI workstations
- AWS GPU instances
- GCP GPU instances
- Azure GPU instances
- hosted GPU environments
The device catalog is versioned and stored locally.
Search uses deterministic fuzzy matching.
This creates a useful separation:
Browser detection
↓
Best-effort environment information
Device catalog selection
↓
Known hardware specification
The application never silently converts weak browser signals into a supposedly precise device specification.
Step 2: Convert hardware into an effective AI memory budget
Raw specifications alone are not enough.
Consider two machines:
Machine A
16 GB system RAM
Integrated GPU
Machine B
32 GB system RAM
8 GB dedicated GPU
Simply comparing total RAM would produce misleading recommendations.
AI workloads care about different memory domains.
Depending on the architecture, we may need to consider:
- system RAM
- dedicated VRAM
- unified memory
- model weights
- KV cache
- runtime overhead
- operating-system usage
- context length
- batching
- concurrent applications
System Check therefore calculates an effective workload memory budget rather than treating advertised system memory as entirely available for inference.
That budget becomes the basis for workload classification.
Step 3: Classify the machine into an AI workload tier
Instead of presenting dozens of raw hardware metrics and expecting the user to interpret them, the application classifies the hardware into a workload tier.
The current model uses:
tiny
small
medium
large
xlarge
Each tier represents a practical AI workload envelope.
Conceptually:
Hardware
↓
Available memory + compute capability
↓
Effective AI budget
↓
Workload tier
For example, lower tiers may be appropriate for:
- tiny language models
- embeddings
- lightweight classification
- basic vision models
- browser inference
- experimentation
Higher tiers gradually unlock workloads such as:
- larger quantized LLMs
- code assistants
- multimodal models
- image generation
- local RAG
- larger context windows
- simultaneous model execution
- multi-user inference servers
At workstation scale, the architecture can recommend significantly larger models and more demanding inference stacks.
The purpose of the tier is not to say:
Your computer is good or bad.
It says:
This is the realistic workload envelope for this machine.
That is a much more useful distinction.
Step 4: Recommend models that fit the hardware
Once the workload tier is known, the application can recommend specific models.
The rules engine considers factors such as:
- model size
- quantization
- approximate memory requirement
- expected runtime
- hardware architecture
- accelerator availability
A recommendation might look conceptually like:
Model
Llama 3.2 3B
Quantization
Q4
Expected memory
~2–4 GB+
Runtime
llama.cpp / Ollama
Expected experience
Comfortable for lightweight local inference
Larger systems can progressively move toward larger model families and more demanding inference workloads.
The application can also provide an estimated tokens-per-second range.
That number should be treated as an estimate rather than a benchmark guarantee.
Actual inference performance depends on many variables:
- model architecture
- quantization format
- context size
- runtime
- GPU offloading
- memory bandwidth
- CPU architecture
- thermal limits
- operating system
- concurrent workloads
So the application is deliberately an advisor, not a synthetic benchmark.
Step 5: Recommend the runtime, not just the model
One mistake in local AI discussions is treating model choice as the entire problem.
The runtime matters almost as much.
Different hardware ecosystems favor different tools.
System Check can recommend tools such as:
Ollama
Useful when someone wants a simple local model runtime with straightforward model management and API access.
LM Studio
Useful for people who prefer a graphical interface while still having the option to expose local models through an API.
llama.cpp
A strong choice when you want direct control over GGUF-based inference, quantization, CPU/GPU offloading, and low-level runtime behaviour.
MLX
Particularly interesting for Apple Silicon workloads where unified memory and Apple's compute stack can be used effectively.
vLLM
More appropriate when the objective moves beyond single-user local experimentation toward high-throughput GPU inference.
WebLLM
Interesting for workloads where inference itself can happen inside the browser.
ComfyUI
Relevant when the hardware budget supports diffusion and image-generation workflows.
Mobile runtimes
Mobile devices introduce a different ecosystem involving technologies such as:
- Core ML
- MLX Swift
- mobile-specific local model applications
This creates a second useful mapping:
Hardware
↓
Workload
↓
Model
↓
Runtime
Choosing the right runtime can sometimes make the difference between a frustrating experiment and a useful system.
Step 6: Recommend things to build
Hardware advice becomes more useful when it leads to an experiment.
So each workload tier also includes project ideas.
For example, lower-power systems might start with:
- text classification
- embedding generation
- semantic search
- lightweight voice commands
- object detection
- tiny local assistants
A medium system might move into:
- local RAG
- private document assistants
- code assistants
- local speech pipelines
- multimodal experiments
- development copilots
More capable systems can explore:
- multi-model pipelines
- large-context RAG
- image generation
- video analysis
- agentic applications
- local AI platforms
- multi-user inference services
- AI development environments
This changes the question from:
Which model can I download?
to:
What useful system can I build with the compute I already have?
That is the more interesting problem.
Deterministic recommendations instead of an AI call
One architectural decision in this project deserves particular attention.
The application already contains an AI Gateway integration.
But the core hardware recommendation process does not depend on an LLM.
That is deliberate.
There are situations where AI is valuable.
There are also situations where a deterministic function is the better engineering choice.
Hardware classification is one of them.
Given the same input:
RAM
GPU
VRAM
device architecture
the application should ideally produce the same workload classification every time.
A simplified rule might resemble:
if (effectiveMemory < 8) {
return "tiny";
}
if (effectiveMemory < 16) {
return "small";
}
if (effectiveMemory < 32) {
return "medium";
}
if (effectiveMemory < 64) {
return "large";
}
return "xlarge";
The real rules can of course account for more variables.
But the principle remains:
Known inputs
↓
Explicit rules
↓
Reproducible recommendation
Instead of:
Known inputs
↓
LLM prompt
↓
Probabilistic recommendation
For this use case, deterministic behaviour gives us:
- reproducibility
- testability
- explainability
- lower latency
- zero inference cost
- offline capability
- easier debugging
It also makes model recommendations easier to audit and update.
AI applications do not need AI everywhere
This project reinforced a design principle I increasingly find useful.
Use AI where uncertainty and semantic reasoning justify it. Use deterministic software where rules are sufficient.
It is easy to reach for an LLM simply because we are building an AI-related product.
But many components around an AI system are better implemented conventionally.
For System Check:
Deterministic
- hardware detection
- device matching
- workload classification
- memory calculations
- compatibility rules
- model filtering
- confidence labels
Potentially AI-assisted
AI can later become useful for things such as:
- explaining why a particular model fits
- creating learning paths
- generating project plans
- suggesting workload optimizations
- comparing alternative hardware
- answering follow-up technical questions
That produces a more disciplined architecture.
AI becomes a reasoning layer rather than an unnecessary dependency.
Privacy was part of the architecture
Hardware fingerprinting deserves careful handling.
The application therefore follows a straightforward model:
Browser
↓
Hardware detection
↓
Rules engine
↓
Recommendation
Nothing needs to leave the device.
There is no requirement to create an account.
Hardware information does not need to be uploaded for classification.
Known device specifications come from the local catalog.
That makes the application particularly suitable for a hardware-detection workflow where users may reasonably prefer not to send device information elsewhere.
Manual override is still necessary
Automatic detection should never become a trap.
Users may know more about their hardware than the browser does.
For example:
- exact RAM capacity
- external GPU configuration
- exact GPU model
- actual VRAM
- upgraded hardware
- custom workstation configuration
- virtual machines
So detected values can be edited.
The workflow becomes:
Detect
↓
Review
↓
Correct if required
↓
Classify
↓
Recommend
That combination of automatic detection and user override usually produces more trustworthy results than either approach alone.
Exporting the recommendation
Users can also export the result as a PDF.
The report contains information such as:
- detected system specification
- confidence level
- hardware workload tier
- model recommendations
- quantization information
- recommended runtimes
- capabilities
- limitations
PDF generation happens client-side using jsPDF.
This turns a temporary browser analysis into something that can be kept as a hardware capability report.
It can also be useful when comparing systems before deciding whether a hardware upgrade is actually necessary.
The application architecture
The application is built using:
TanStack Start
React 19
TypeScript
Vite 7
Tailwind CSS 4
shadcn/ui
Bun
jsPDF
The important application logic is intentionally separated into three components.
detect-hardware.ts
↓
What does the browser know?
device-catalog.ts
↓
What do we know about explicitly selected devices?
hardware-rules.ts
↓
What workloads fit those resources?
That separation makes the architecture easier to reason about.
┌───────────────────────────┐
│ Browser UI │
│ TanStack + React │
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ Hardware Detection │
│ CPU / RAM / GPU / OS │
└─────────────┬─────────────┘
│
uncertain device?
│
▼
┌───────────────────────────┐
│ Device Catalog │
│ Known hardware profiles │
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ Rules Engine │
│ Effective resource budget │
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ Workload Tier │
│ tiny → small → medium → │
│ large → xlarge │
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ Recommendations │
│ Models │
│ Tools │
│ Projects │
│ Capabilities │
│ Limitations │
└───────────────────────────┘
The architecture is deliberately simple.
The complexity lives in the knowledge encoded into the rules and device catalog.
The interesting challenge: keeping recommendations current
Hardware detection itself is only one part of the problem.
The harder long-term problem is maintaining the recommendation knowledge base.
The local AI ecosystem changes quickly.
New developments continuously affect:
- model sizes
- quantization formats
- inference runtimes
- Apple Silicon support
- GPU support
- mobile inference
- browser inference
- context-window requirements
- model efficiency
That means the rules engine and device catalog need to be versioned like any other technical knowledge base.
This is one reason I prefer making these rules explicit.
A rule can be reviewed.
A model entry can be updated.
A device can be added.
A test can validate the expected classification.
The behaviour does not disappear inside a prompt.
Testing the advisor
A deterministic architecture also makes testing much easier.
For example:
Given:
8 GB RAM
CPU-only
Expected:
tiny/small workload class
small quantized model recommendations
no high-memory GPU workload recommendations
Another fixture might represent:
Given:
Apple Silicon
32 GB unified memory
Expected:
larger local models
MLX-compatible recommendations
higher workload class
Another could represent:
Given:
NVIDIA GPU
24 GB VRAM
Expected:
GPU-oriented runtimes
larger quantized models
image-generation capability
higher-throughput inference options
Device catalog matching can be tested independently from hardware classification.
This is important because AI workload advice should be predictable enough to regression-test.
Where I want to take it next
There are several directions in which this can grow.
Better mobile device coverage
Mobile hardware detection is inherently restricted by browsers, so explicit device selection will become increasingly important.
The catalog can expand across:
- iPhone generations
- iPads
- Android devices
- mobile NPUs
Better accelerator awareness
Future versions can reason more deeply about:
- Apple Neural Engine
- NVIDIA Tensor Cores
- integrated NPUs
- AMD accelerators
- Intel AI hardware
Instead of treating everything primarily as CPU/GPU memory.
Context-length-aware recommendations
A 7B model running with a small context window and the same model running with a very large context window are not equivalent workloads.
Future recommendations can incorporate:
Model weights
+
KV cache
+
Context requirement
+
Runtime overhead
to produce a more realistic memory envelope.
Workload-specific modes
Eventually it may be more useful to ask the user what they want to build.
For example:
Local chatbot
Code assistant
RAG
Computer vision
Image generation
Voice agent
AI agent
Embedding server
The advisor could then evaluate the hardware specifically against that workload.
Hardware comparison
Another useful direction is to compare:
My current machine
vs
Potential upgrade
vs
Cloud GPU
The application could explain what new AI workload classes become available with each option.
This could help users avoid upgrading simply because a new GPU or laptop appears attractive.
A small project with a broader idea
System Check for AI started from a simple practical problem:
People are experimenting with local AI, but hardware recommendations are fragmented and often confusing.
The larger idea is more interesting.
As AI starts running across increasingly diverse environments—
- browsers
- phones
- laptops
- Apple Silicon
- Raspberry Pis
- Jetson devices
- workstations
- cloud GPUs
- edge systems
—we need better ways to map software capability to available compute.
The common question should no longer be:
What is the most powerful AI model available?
A more useful question is:
What is the right AI workload for the compute available to me?
That is what System Check for AI is trying to answer.
And because all of the core analysis runs locally and deterministically, the system itself follows the same principle it recommends:
Use only as much compute—and as much AI—as the problem actually requires.
Related reading: Understanding GenAI Hardware: CPU, GPU, NPU, Inference, and Model Serving.
Ask AI About the Author
Open this query in ChatGPT, Claude, or Perplexity.
Comments
Comments are open to confirmed email subscribers. Use the email you subscribed with. To edit a comment, delete it and post a new one.
Subscribe to get the new blogs.
Field notes from someone who ships before they write about it. Sovereign AI, AI-SDLC, DevOps, and what 59 production deployments teach you. No spam. Unsubscribe anytime.