Real-time sign-language recognition that turns live ASL gestures into text and speech for accessible communication.
01Overview
An accessibility-focused computer-vision system that recognizes American Sign Language gestures from a live camera feed, assembles detected signs into text, and converts the resulting text into speech. The project combines object detection, temporal smoothing, word construction and text-to-speech to support real-time communication between sign-language users and non-signers. Developed as my undergraduate thesis project.
02What it does
- 1Live-camera ASL gesture detection with YOLO
- 2Temporal smoothing for stable predictions
- 3Sign-to-text word construction
- 4Speech synthesis of assembled text
03The challenge
Keeping detection stable and responsive on a live feed required temporal smoothing over noisy per-frame predictions and inference optimization across ONNX and TensorFlow Lite.
04The outcome
A working end-to-end pipeline from camera input to spoken output, demonstrating real-time accessible communication.
05Built with
- Python
- Ultralytics YOLO
- OpenCV
- PyTorch
- ONNX
- TensorFlow Lite
- Text-to-Speech
Next project — N°03
GeoInsight
A real-time location-intelligence dashboard tracking 100+ simulated vehicles across Dhaka with live maps, clustering and heatmaps.
