Product & Technical Proposal

Hesabman (حساب‌من)

A resilient, voice-first intelligent accounting system engineered for Persian speakers. Seamlessly bridges conversational speech, zero-shot intent categorization, local language models, and double-entry personal accounting.

Project Codename
Hesabman Voice
Target Language / Locale
Persian / Farsi (fa-IR)
Accounting Backend
Actual Budget (SQLite API)
AI Architecture
Jev + Local LLM + Whisper

1. Executive Summary & Problem Statement

Manual bookkeeping apps suffer from high cognitive friction: users must remember to open an app, navigate nested category trees, manually input amounts, and select account ledgers. Consequently, personal budgeting habit retention drops by over 80% after two weeks.

Hesabman redesigns the personal accounting interaction around natural human speech: a user taps a single button, speaks colloquially in Persian (e.g. «پنجاه هزار تومن اسنپ دادم» or «حقوق این ماه واریز شد»), and the application instantly parses, categorizes, validates, and commits the ledger transaction.

2. System Architecture & Information Pipeline

Hesabman employs a robust multi-tiered pipeline that separates high-frequency audio ingestion, semantic intent categorization, parameter extraction, and financial transactional persistence.

1

Voice Capture & Transcribe

In-browser Web Audio stream analyzer + MediaRecorder. Fast chunk upload to OpenAI-compatible Deepgram Nova-3 Whisper endpoint enforcing digit formatting.

2

System One (Jev) Classifier

Codiv SystemOne API (DiffusionGemma/OpenJev) computes probability distributions for Intent, Transaction Type (Income vs Expense), and Financial Category.

3

LLM Function Caller

Prompt augmented with Jev's classified intent. Local LLM extracts structured function arguments (`add_transaction` or `get_financial_report`).

4

Actual Budget Engine

Server-side `@actual-app/api` reconciles SQLite transactions, enforces double-entry rules, and recalculates real-time ledger balances.

Application Visual Interface

Mobile-first Persian (RTL) Design System
Unified Voice Card & Live Balance
Voice prompt and ledger card morph into a single bubble with real-time balance metrics.
Slide-Over Account Drawer
Seamless account switching for context-specific finances (e.g. Trips, Roommates, Businesses).
Offline Voice Persistence & Audio Review
Audios are cached in local IndexedDB. Users can play back or delete recordings before online syncing.

3. Technical Innovations & Key Features

Non-Blocking Recording Queue

Unlike conventional voice bots that lock the user while processing speech, Hesabman decouples recording from inference. Users can record back-to-back voice items rapidly; each spawned item creates an isolated parallel processing thread.

Web Audio API Async Threads AbortController

Deterministic Semantic Classifier (Jev)

Utilizes Codiv's OpenJev SystemOne model for low-latency probability readout across multiple questions in a single request: Intent classification, Expense vs Income detection, and Automatic Categorization.

OpenJev 0.1 SystemOne API Probabilistic Calibration

Double-Entry Actual Budget Ledger

All recorded entities are persisted directly into an embedded SQLite engine powered by Actual Budget (`@actual-app/api`). Guarantees zero data loss, strict ledger auditability, and instant account balance recalculation.

@actual-app/api SQLite Ledger Envelope Budget

Jalali (Solar Hijri) Native Calendar

Built-in Persian calendar intelligence calculates precise month, week, and custom date bounds based on Shamsi leap years, preventing discrepancies common in standard Gregorian accounting software.

Intl.DateTimeFormat fa-IR-u-ca-persian Date Filtration

In-Place Interactive Editing

Recognized entities are never static. Users can tap title, amount, type (+/-), or category badge to perform inline mutations. The backend immediately reconciles the edit and refreshes the live account balance.

Inline Mutation Optimistic UI Zero-Layout-Shift

Zero-Loss Offline Caching

Audio blobs recorded during network outages are securely staged inside IndexedDB (`idb-keyval`). The client displays interactive pending cards with voice playback and automatically uploads them upon reconnection.

IndexedDB idb-keyval Auto-Sync Queue

4. Technology Stack Comparison

Component Technology Choice Rationale & Advantage
Frontend Architecture Next.js 14 App Router, React 18, Tailwind CSS Server-rendered shell, fast client-side reactivity, zero CSS runtime overhead.
Typography & Localization Google Fonts Vazirmatn, RTL CSS Flex/Grid First-class Persian typography, crisp numeral rendering, native RTL flow.
Speech Transcription Deepgram Nova-3 / OpenAI Whisper API Accurate Persian language model with digit-first phonetic transcriptions.
Intent Classification Codiv SystemOne (OpenJev 0.1) Sub-second probabilistic classifier eliminates brittle regex pipelines.
Parameter Extraction Local LLM (OpenAI-compatible / Gemini Flash) Zero-shot schema extraction and polite Persian assistant messaging.
Accounting Engine Actual Budget Core API (`@actual-app/api`) Open-source, privacy-first local SQLite double-entry ledger.
Client Storage IndexedDB (`idb-keyval`) High-capacity binary storage for audio blobs and offline chat histories.

5. Local AI & Privacy-First Alternatives

For fully sovereign, air-gapped, or privacy-strict deployments, the cloud-based AI dependencies can be entirely replaced by highly optimized local models running on consumer hardware.

Fast Decisions (System One)

Replace Codiv's OpenJev API with near-instant local probabilistic classification. Laya (322M–421M params, built on ModernBERT) handles Jev-style typed decisions without autoregressive overhead. Alternatively, Kev (0.8B+, Qwen architecture) matches the System One interface locally, or Von (~395M) offers rapid non-autoregressive option-scoring.

Laya Kev Von

Speech-to-Text (STT)

Match Deepgram's extreme speed without melting local GPUs. Faster-Whisper (CTranslate2 implementation) runs up to 4x faster than the original Whisper with significantly lower VRAM. For CPU-bound or low-end hardware, Distil-Whisper is 6x faster and 49% smaller while maintaining a near-identical word error rate.

Faster-Whisper Distil-Whisper

Tool-Calling LLMs

Replace cloud LLM APIs with resource-friendly models built for structured data workflows. Qwen3 8B comfortably fits on an 8GB laptop and excels at multilingual tool calling. For highly constrained devices, Phi-4-mini (3.8B) or Gemma 3 4B provide heavily optimized reasoning and function calling on standard CPU/RAM setups.

Qwen3 8B Phi-4-mini Gemma 3 4B

Unified Management Gateway

To orchestrate this purely local multi-model stack efficiently, LocalAI acts as a complete drop-in REST API replacement. It can simultaneously host the tool-calling LLMs, Faster-Whisper (STT), and audio workflows under one unified, OpenAI-compatible /v1 endpoint.

LocalAI OpenAI-Compatible API

6. Future Roadmap & Strategic Next Steps

Milestone 1
Receipt & Invoice OCR

Camera capture allowing instant receipt photo extraction alongside voice notes.

Milestone 2
Bank SMS Parsing

Automatic webhook integration for Persian banking debit/credit SMS alerts.

Milestone 3
PWA & Biometric Security

Installable Progressive Web App with FaceID/Fingerprint biometric app lock.

Get in Touch

Contact & Collaboration

Interested in implementing or expanding the Hesabman voice-first accounting platform? Let's connect.