← Back to Work
Case Study 03 · Systems & AI Architecture

Argus: Private AI Search Without Sending Data to the Cloud

Most modern AI systems require users to upload sensitive information to external providers. For many organizations and individuals, that creates a trust problem.

Search SpeedSub-10msIn-memory similarity retrieval.
ConcurrencyWorker PoolsEfficient async file parsing.
Security ModelAES-256 CTRZero-trust, disk-level data isolation.
Privacy ModelPrivate On-Device AIZero data sent to third-party cloud servers.

Context

Modern search tools create a difficult tradeoff: privacy vs functionality. Most solutions force users to choose one. Organizations cannot safely run cloud-based semantic search tools across proprietary documentation without risk of leakage. I wanted to explore whether both high-precision search and total user data ownership could coexist.

The question behind Argus was simple: Can modern AI-powered search remain useful without requiring users to surrender ownership of their data?

Approach

I designed and built a local-first retrieval system that processes, indexes, encrypts, and searches documents entirely on local hardware, requiring zero external APIs.

The system combines a background crawling daemon to parse documents asynchronously, a local similarity index, and direct Server-Sent Events to stream conceptual matches safely to the user client interface.

Core Engineering

1. High-Speed Processing Without Resource Spikes

Parsing multi-gigabyte directories containing diverse document types requires I/O efficiency and resource containment. I developed the core ingestion system to manage computational workloads smoothly.

Controlled Processing Boundaries
By implementing controlled worker pools and structured, synchronized task queues, I isolated heavy parsing tasks from file crawling. The system processes documents concurrently without slowing down the machine. Limiting active workers ensures stable CPU and memory overhead, allowing the computer to remain fast and responsive.

2. Sub-10ms Answers via Vector Search

To support semantic search queries (searching for conceptual ideas rather than exact keyword matches), document segments are transformed into vector embeddings. I integrated an embedded pipeline mapping text to vectors, loaded directly into an in-memory vector index. High-dimensional calculations are executed instantly, delivering semantic results in under 10 milliseconds.

3. Zero-Trust Storage Security & Live Streaming

To maintain strict privacy, plain text document segments are never written directly to disk. The parsed texts are encrypted using AES-256 CTR mode and matched with specific vector indices.

When a search is submitted, it is vectorized and sent to the local similarity index. The system retrieves only the corresponding encrypted files, decrypts them in secure temporary memory buffers, and streams them instantly to the user interface using Server-Sent Events. Plain text fragments are never saved, giving users instant, word-by-word live feedback.

The Stack

Utilizing low-latency, highly specialized systems languages and optimized mathematical engines to secure absolute privacy:

Go (Golang)PythonFAISS Vector IndexingAES-256-CTRServer-Sent Events (SSE)Next.js (TypeScript)

Outcome

Argus demonstrated that high-quality semantic search does not require cloud infrastructure. The system achieved sub-10ms retrieval, zero cloud dependencies, encrypted storage, and a local-first architecture.

The project became less about search and more about a broader question: What would privacy-first AI infrastructure actually look like?

Have a technical problem?
Let's talk.

Bring the problem. We'll clarify the outcome, define the right technical approach, and build a system you can rely on.