HushMap: AI Services API
The central nervous system linking physical M5GO devices, external Computer Vision tensors, and Conversational NLP APIs synchronously.
Setup Instructions
Prerequisites
- Python 3.10+ is strictly recommended to support asynchronous typing paradigms.
- FFmpeg must be successfully registered onto your OS PATH environments. This engine handles the core conversions decoding MP3 output arrays into 16-bit, 16kHz Mono arrays natively required for browser contexts:
- Ubuntu/Debian:
sudo apt install ffmpeg - macOS:
brew install ffmpeg - Windows: Install globally via the FFmpeg website.
- Ubuntu/Debian:
Environment Initialization
Bootstrap the virtual environment and initialize project dependencies:
cd backend
python -m venv venv
source venv/bin/activate # Windows: .\venv\Scripts\activate
pip install -r requirements.txt
Configuration Tokens
Provide runtime keys securely targeting TerpAI context queues, Gemini Fallback, and ElevenLabs synthesized avatars within a .env dotfile:
ELEVENLABS_API_KEY=sk_...
ELEVENLABS_VOICE_ID=JBFqnCBsd6RMkjVDRZzb
TERP_AI_BEARER_TOKEN=eyJhbGciOiJSUz...
TERP_AI_CONVERSATION_ID=37fa27cc-...
GEMINI_API_KEY=AIza...
MONGODB_URI=mongodb+srv://...
USE_DB=false
Note: USE_DB controls whether the application connects to MongoDB (true) or uses on-the-fly generated in-memory data for demonstrations (false).
To invoke the engine, simply execute Uvicorn across your 0.0.0.0 loopback:
uvicorn server:app --host 0.0.0.0 --port 8000
Gateway Pipelines
Full-Duplex Subroutines (/ws/voice)
This WebSocket proxy establishes a fully integrated multi-turn communication bridge seamlessly interacting between Edge node Hardware APIs (ESP32/M5GO/Browsers) and NLP architectures.
- Int16 Byte Array Exchange: Devices connect to
ws://<server_ip>:8000/ws/voiceand push raw binary frames asynchronously over the socket. - Contextual Augmentation: The server waits for the
"stop_listening"payload event to signify a completed audio snippet. That float array is cast throughfaster-whisperand combined seamlessly with real-time decibel tracking telemetry parameters natively attached into the AI user conversation chunk. We utilize Terp AI with an automatic, seamless fallback to Gemini 2.5 Flash if the primary Terp service is unavailable. - TTS Pipeline Rendering: Output predictions are caught instantly, forwarded natively into the
ElevenLabsTTS interface renderingpcm_16000wav codecs, and alerted back down to clients using atts_readydispatcher.
Tensor Vision Endpoints (/api/vision/room-status)
Leveraging OpenCV bindings layered beneath a YOLOv8-driven bounding box topology detector, this POST API analyzes raw camera image buffers returning capacity logic natively.
Note
This API calculates euclidean distances algorithmically detecting adjacent proximities between "person" classifiers and untaken "chair" bounding frames to accurately diagnose available seats inside crowded architectures!
Response Output Protocol:
{
"room_status": "full",
"counts": {
"people": 2,
"chairs": 2
},
"pairs": [
{
"person_index": 0,
"chair_index": 1,
"distance": 150.5
}
]
}
Database Registries
GET /api/study-rooms/history: Pulls the active global repository of logged architectural noise measurements captured universally within the preceding 24 hours. (Uses MongoDB or in-memory generated data based on theUSE_DBflag).GET /api/study-rooms: Pulls generic unstructured noise lists.
Important
The browser frontend strictly configures standard Web Audio API's
ScriptProcessorNodeinterfaces routing data synchronously to this backend! Wait to close down pipelines until after all WS queues have successfully been delivered.