A resilient, voice-first intelligent accounting system engineered for Persian speakers. Seamlessly bridges conversational speech, zero-shot intent categorization, local language models, and double-entry personal accounting.
Manual bookkeeping apps suffer from high cognitive friction: users must remember to open an app, navigate nested category trees, manually input amounts, and select account ledgers. Consequently, personal budgeting habit retention drops by over 80% after two weeks.
Hesabman redesigns the personal accounting interaction around natural human speech: a user taps a single button, speaks colloquially in Persian (e.g. «پنجاه هزار تومن اسنپ دادم» or «حقوق این ماه واریز شد»), and the application instantly parses, categorizes, validates, and commits the ledger transaction.
Hesabman employs a robust multi-tiered pipeline that separates high-frequency audio ingestion, semantic intent categorization, parameter extraction, and financial transactional persistence.
In-browser Web Audio stream analyzer + MediaRecorder. Fast chunk upload to OpenAI-compatible Deepgram Nova-3 Whisper endpoint enforcing digit formatting.
Codiv SystemOne API (DiffusionGemma/OpenJev) computes probability distributions for Intent, Transaction Type (Income vs Expense), and Financial Category.
Prompt augmented with Jev's classified intent. Local LLM extracts structured function arguments (`add_transaction` or `get_financial_report`).
Server-side `@actual-app/api` reconciles SQLite transactions, enforces double-entry rules, and recalculates real-time ledger balances.
Unlike conventional voice bots that lock the user while processing speech, Hesabman decouples recording from inference. Users can record back-to-back voice items rapidly; each spawned item creates an isolated parallel processing thread.
Utilizes Codiv's OpenJev SystemOne model for low-latency probability readout across multiple questions in a single request: Intent classification, Expense vs Income detection, and Automatic Categorization.
All recorded entities are persisted directly into an embedded SQLite engine powered by Actual Budget (`@actual-app/api`). Guarantees zero data loss, strict ledger auditability, and instant account balance recalculation.
Built-in Persian calendar intelligence calculates precise month, week, and custom date bounds based on Shamsi leap years, preventing discrepancies common in standard Gregorian accounting software.
Recognized entities are never static. Users can tap title, amount, type (+/-), or category badge to perform inline mutations. The backend immediately reconciles the edit and refreshes the live account balance.
Audio blobs recorded during network outages are securely staged inside IndexedDB (`idb-keyval`). The client displays interactive pending cards with voice playback and automatically uploads them upon reconnection.
| Component | Technology Choice | Rationale & Advantage |
|---|---|---|
| Frontend Architecture | Next.js 14 App Router, React 18, Tailwind CSS | Server-rendered shell, fast client-side reactivity, zero CSS runtime overhead. |
| Typography & Localization | Google Fonts Vazirmatn, RTL CSS Flex/Grid | First-class Persian typography, crisp numeral rendering, native RTL flow. |
| Speech Transcription | Deepgram Nova-3 / OpenAI Whisper API | Accurate Persian language model with digit-first phonetic transcriptions. |
| Intent Classification | Codiv SystemOne (OpenJev 0.1) | Sub-second probabilistic classifier eliminates brittle regex pipelines. |
| Parameter Extraction | Local LLM (OpenAI-compatible / Gemini Flash) | Zero-shot schema extraction and polite Persian assistant messaging. |
| Accounting Engine | Actual Budget Core API (`@actual-app/api`) | Open-source, privacy-first local SQLite double-entry ledger. |
| Client Storage | IndexedDB (`idb-keyval`) | High-capacity binary storage for audio blobs and offline chat histories. |
For fully sovereign, air-gapped, or privacy-strict deployments, the cloud-based AI dependencies can be entirely replaced by highly optimized local models running on consumer hardware.
Replace Codiv's OpenJev API with near-instant local probabilistic classification. Laya (322M–421M params, built on ModernBERT) handles Jev-style typed decisions without autoregressive overhead. Alternatively, Kev (0.8B+, Qwen architecture) matches the System One interface locally, or Von (~395M) offers rapid non-autoregressive option-scoring.
Match Deepgram's extreme speed without melting local GPUs. Faster-Whisper (CTranslate2 implementation) runs up to 4x faster than the original Whisper with significantly lower VRAM. For CPU-bound or low-end hardware, Distil-Whisper is 6x faster and 49% smaller while maintaining a near-identical word error rate.
Replace cloud LLM APIs with resource-friendly models built for structured data workflows. Qwen3 8B comfortably fits on an 8GB laptop and excels at multilingual tool calling. For highly constrained devices, Phi-4-mini (3.8B) or Gemma 3 4B provide heavily optimized reasoning and function calling on standard CPU/RAM setups.
To orchestrate this purely local multi-model stack efficiently, LocalAI acts as a complete drop-in REST API replacement. It can simultaneously host the tool-calling LLMs, Faster-Whisper (STT), and audio workflows under one unified, OpenAI-compatible /v1 endpoint.
Interested in implementing or expanding the Hesabman voice-first accounting platform? Let's connect.