on-premise AI AI-native builds principal-engineered

Private, on-premise AI infrastructure and applications, built within your data boundary.

I help teams evaluate, build and operate open-weight AI on their hardware or in their cloud account. Each engagement ends with measured results, documented limits and a handover your engineers can use.

24+ yrsin IT, software & product development — architecture, engineering and security
0cloud calls in the reference architecture
1 + AIthe team on my current engagement — eighteen months and counting

Three ways to work together

offer 01 · private infrastructure

Private AI infrastructure

Private, on-premise AI workflows that never send your data to a cloud API — for legal, health, finance, and anyone whose compliance team says no to SaaS LLMs.

  • Audit: where AI actually pays in your business
  • Reference build on your hardware — or your VPC
  • Domain adaptation: fine-tune open models on your data — data and weights never leave
  • Handover as a named deliverable — runbooks and a measured baseline; your team runs it without me by month three
YOUR BUILDING your data docs · mail · code GPU box models + agents answers stay inside ☁ cloud AI — never called , zero egress, zero tokens billed

Fixed-scope engagement. You keep the keys, the weights and the numbers.

offer 02 · inside your walls

AI features, inside your walls

0→1 product builds and AI features in the app you already have — architected for private infrastructure from the first commit. Three-plus years of AI-leveraged delivery: principal engineer, direct with the client, AI doing the heavy lifting.

  • Principal on every commit — no bench, no juniors on your budget
  • Senior specialists added on demand, never in advance
  • Shipping cadence measured in weeks, not quarters

Same boundary rules as offer 01 — the models underneath run where you say they run.

offer 03 · fractional CTO

Fractional CTO & advisory

Senior technical leadership without the full-time hire: architecture and vendor decisions, an AI roadmap that survives model churn, and post-handover stewardship — each model refresh re-benched against your baseline, with independent verdicts that include don't upgrade. Not a new label — I've held the role for the past eighteen months.

Quarterly, fixed fee — pause any quarter. Pairs with offers 01 and 02, or stands alone.

Who this is for

Data-boundary AI isn't for everyone — if your data can go to the cloud, a good enterprise tier probably serves you. It's for the teams whose constraint isn't budget; it's where the data is allowed to go.

Finance & paymentsPrivate AI for payment-telephony workflows, drawing on my PCI payment-systems experience. Scope and controls are agreed with your security team.
Logistics & last-mileDriver onboarding, document pipelines and dispatch AI. I built a platform in this exact vertical — the workflows aren't new to me.
TelecomsTwenty years in and around carrier networks — protocol stacks to platform. Network-ops and subscriber-data AI on your own infrastructure.
Corporates under NDATenders, M&A, R&D pipelines. Your contracts already said no to the cloud.
Legal firmsCase files, contracts, privilege. Drafting and review that never leaves chambers.
HealthcareNotes, triage, admin automation. HIPAA follows the data — running inference on your hardware takes the AI vendor out of your BAA chain.
Public sector & sovereigntyResidency designed into the architecture first — so the paperwork describes something true.
Retail & e-commerceCustomer data and purchase intelligence. Personalization that doesn't leak your competitive edge.

The machine room

I don't sell AI readiness from slideware. I run an on-premise production LLM stack on my own hardware — self-hosted, with fork-level depth in llama.cpp, vLLM and SGLang — and every claim I make about private AI comes with a number attached.

wasif@club-3090:~$ nvidia-smi --query-gpu=name --format=csv,noheader
NVIDIA GeForce RTX 3090
NVIDIA GeForce RTX 3090
wasif@club-3090:~$ bash scripts/bench.sh  # 3 warm-ups + 5 measured runs
run  2026-09-15 · dedicated boot
model  Qwen3.8-27B · Frozenlock AutoRound INT4 W4A16
engine  SGLang v0.5.19 · TP=2 · native MTP · KV fp8_e4m3
narrative decode  101 tok/s  · mean accept 3.142
code decode      135 tok/s
prefill @10K     1,224 tok/s
configured context  262,144 tokens — decode measured at short prompt
wasif@club-3090:~$ ▊
2× 3090consumer GPUs — the same method runs on A100/H100 class
262Ktoken context, zero cloud
0hosted-model token charges — operating costs, licences and upgrades documented
The rig is open source. club-3090 — ★ 2,283 — an open-source stack for local AI on consumer GPUs: working configs, quant recipes and measured benchmarks across vLLM, llama.cpp and SGLang. Full measured row (2026-09-15) pinned in the repo — clone it before you buy anything.
★ 2,283 on GitHub

The agent room

The machine room is the floor. One layer up is where the new practices are evolving right now: agents that plan, subagents that work in parallel, harnesses that keep them honest. I build and operate these harnesses daily — this site was assembled by one, as a labelled internal demonstration. On your stack, they run the same way: inside your walls.

Agents & subagents

One agent plans; subagents execute in parallel, each with its own context and a job small enough to finish. Decomposition, delegation and verification loops — so the output is checked, not vibes-checked.

Orchestration & optimisation

The right model for each task, context budgets that don't silently bleed, routing and caching, cost and latency measured per run. The difference between a demo and a system is measurement at this layer.

Harnesses that hold the line

Planning, tools, memory, guardrails, human checkpoints. The harness is what separates an agent that helps from an agent that improvises with your business — and it is engineered, not improvised.

Every modality, one boundary

Code with tests as guardrails. Documents, audio transcription, video indexed and searchable. Each modality needs its own pipeline — and every one of them can run inside the same boundary as the models underneath.

How I work

  1. Audit before architecture. Every engagement starts by measuring your actual workload — not estimating it.
  2. Fixed-scope offers, fixed price bands. Outcomes, not timesheets. I left the man-day business on purpose.
  3. Principal on every commit. No bench, no juniors learning on your budget. Senior capacity joins on demand, never in advance.
  4. Everything measured. Claims arrive with numbers attached — or they don't arrive.

— Wasif Basharat, founder & fractional CTO, AppScience

Questions I get

Can my team use AI tools with client data?

Depends entirely on where the model runs. Consumer chat tools generally can't promise your inputs stay private. Your options: enterprise tiers with zero-retention agreements, or models running on your own hardware — where no copy ever exists outside your walls. Which answer fits is a function of your data and your regulator; determining that is the first deliverable of an engagement.

Is my data used to train AI models?

With consumer tiers: often yes by default, under terms that have changed before and will change again. Provider commitments differ by product — check training use, retention, access and residency against the specific service and contract. Models on your own hardware: no vendor whose policy changes to track, no telemetry.

What does "GDPR-compliant AI" actually mean?

No AI tool is automatically compliant. GDPR cares where personal data goes, who processes it, and on what legal basis; in the US, HIPAA's business-associate agreements and SOC 2 expectations land the same way. Running open-weight models on infrastructure you control removes the hosted model provider from that data flow. You still need an appropriate legal basis, applicable contracts — and a DPIA where the processing is likely high risk — but there are fewer parties to account for.

Do we need to buy expensive GPUs to use AI privately?

Not always. Some workloads fit a single machine; others fit a small box in your office or a private slice of cloud. Sizing against your actual workload — not vendor spec sheets — is what the audit is for. My reference rig runs production-shaped workloads on two consumer GPUs.

Are open-source models as good as ChatGPT?

For a large share of business workloads — retrieval, summarization, drafting, classification, agentic tool use — yes, and the gap keeps narrowing. For the hardest reasoning tasks, frontier cloud models still lead. A good build uses local where it counts and is honest about the rest.

Can we fine-tune a model on our own data — privately?

Yes. Open-weight models can be fine-tuned on hardware you control, so both the training data and the resulting weights stay inside your boundary — no vendor sees your corpus, and the adapted model is your asset. The engineering question is usually whether fine-tuning is even the right tool: often retrieval (RAG) or better prompting solves the problem cheaper. Answering that honestly is what the audit is for; when fine-tuning is right, it runs on the same private stack.

How does contracting and procurement work?

Through AppScience LLC — a Delaware-registered US company, operating for five years. US contracts, USD invoicing, security questionnaires and data-processing agreements answered directly. Delivery engineering runs from Lahore under the same boundary rules: access is agreed before work begins — approved accounts, environments and tools, with production-data access restricted and logged.

Can the AI run fully air-gapped — no internet at all?

Yes, for inference on the reference architecture: model weights, storage and telemetry all stay local, and nothing requires an outbound call. Air-gapped installs are validated against your workload before handover — offline is a constraint we design for, not an afterthought.

Are AI agents safe to use with business data?

An agent is AI with standing access — mail, files, tools — acting on your behalf. That's where the productivity is, and also where the blast radius is. Agents on infrastructure you control, with deliberate permission boundaries, are manageable. Agents piping your systems through a remote model deserve scrutiny before, not after, deployment.

Tell me what you're building — or what your data isn't allowed to do.

AppScience LLC — a Delaware-registered US company, operating for five years. Contracts and invoicing run through the LLC; engineering is Lahore-based, working North America & UK hours. First call is 30 minutes and costs you nothing but the agenda. Capacity is deliberately small — a handful of concurrent engagements, never a bench.