Front obstacle
- Front sensor detects an obstacle
- Immediate STOP
- Read fresh rear distance
- Reverse briefly if the rear is safe
- STOP and recheck
- Turn, check the front, then resume
A smartphone-powered autonomous companion robot combining embedded electronics, Android software, physical safety, expressive interaction, and Realtime AI.

AMSAN is a companion robot I designed and built around an Android smartphone and an ESP32. The phone handles the face, voice, memory, and high-level behaviour. The ESP32 controls movement, sensors, arms, RGB, and physical safety.
I wanted to learn robotics by making one complete physical system, not only testing individual sensors or pieces of code. Commercial companion robots can be expensive and hard to customise or repair, so I tried combining accessible electronics with a phone I could use as the display, microphone, speaker, processor, storage, and internet interface.
The result is a stable functional companion-robot prototype. It is not commercial hardware or a certified consumer product.
What V1 had to do
System architecture
The division is intentional. Android can decide what AMSAN wants to do, but the ESP32 can refuse an unsafe movement request.
FaceRig · microphone · speaker · Realtime voice · Life Engine · games · memory · settings · Developer Hub
Motors · front/rear distance · edge sensors · touch · arms · RGB · telemetry · watchdog · STOP rules
Motor driver
Differential drive
Physical feedback
Hardware
The blue rectangular body carries a horizontally mounted Android phone, two front ultrasonic modules, an RGB ring, padded side arms, two powered wheels, a caster, and the internal controller and power wiring.


The INA219 voltage/current sensor is provisioned at address 0x40 with firmware retry logic. The robot’s main movement and safety functions do not depend on it starting successfully.
Motors AIN1 16 · AIN2 17 · PWMA 18 · BIN1 19 · BIN2 23 · PWMB 25
Ultrasonic Front TRIG 26 / ECHO 34 · Rear TRIG 32 / ECHO 35
Edge Front 27 · Rear 4
Servos Left 14 · Right 13
RGB 33 · 12 pixels
Touch Left 36 · Right 39
INA219 SDA 21 · SCL 22 · 0x40
Autonomy
AMSAN has one local character: curious, playful, mischievous, expressive, and sometimes a little stubborn. The behaviour engine coordinates movement, gaze, expressions, arms, RGB, and pauses so it feels coherent rather than random.
A front obstacle makes forward movement unsafe, not automatically backward movement. A rear obstacle blocks reverse, not automatically forward movement.
The two HC-SR04 sensors use an approximate 18 cm obstacle threshold.
Two touch sensors can trigger local face, RGB, and arm reactions without an API call.
The RGB engine is non-blocking so it does not delay sensors, Bluetooth, watchdog refreshes, or STOP commands.
Safety hierarchy
If drive refreshes stop for about 1500 ms while the motors are active, the ESP32 stops them. Android refreshes active commands roughly every 400–500 ms.
Three front edge sensors are combined to GPIO27 and two rear sensors to GPIO4. Edge events outrank ordinary roaming and obstacle recovery.
Two roller/limit switches add a safety path through the TB6612 STBY arrangement, so motor safety does not depend entirely on Android.


Android, voice, and local behaviour
AMSAN used gpt-realtime-2.1-mini during development for intentional voice conversations. Obstacles, recovery, roaming, touch reactions, expressions, arms, RGB, games, and mood changes remain local.
Voice can request actions such as wave, dance, celebrate, arms home, happy, surprised, mischievous, sleepy, wake, or a game. The model cannot send unrestricted motor commands; requests are mapped into typed local actions and still pass through ESP32 safety.
The app was tested in English, Urdu, and mixed Urdu/English. The phone microphone supports voice activity handling and barge-in where available, so a person can interrupt AMSAN while it is speaking.
Autonomous events do not start Realtime turns. If nobody is intentionally talking to AMSAN, conversational API use should be zero. The Realtime session closes after about 25 seconds of inactivity.
The FaceRig uses large cyan eyes, pupils, highlights, eyebrows, eyelids, blinking, gaze, mouth movement, cheeks, and roughly 30 expressions and animations. Local reactions also include stylised animal sounds, birthday behaviour, and a short original tune or chant.

A local 3×3 game with touch input, validation, winner and draw checks, minimax strategy, rematch, exit, sounds, and robot reactions.
Local choice and scoring with rematch, exit, sounds, expressions, arms, and RGB reactions.
Profile memory on the phone can retain AMSAN’s identity, Hassan’s profile, useful facts, and preferences across app restarts. It is not unlimited human-like memory.
Development problems
Moving the arms made the ESP32 unstable. I revised the 5V distribution and common-ground arrangement instead of treating the resets as a software problem.
The ESP32 stops the motors after about 1500 ms without a valid drive refresh. Android was refreshing too slowly, so I changed it to send active commands about every 400–500 ms and kept the watchdog enabled.
The robot could detect an obstacle and stop, but sometimes stayed stopped or tried the blocked direction again. Fresh telemetry was clearing a temporary Android stop marker before recovery ran. I kept the directional obstacle state until recovery finished.
During development I had to separate simulator behaviour from the real Bluetooth transport. A screen that looked connected was not enough; commands and fresh sensor telemetry had to agree.
Realtime speech first used the phone's call-style earpiece path. I moved it to media audio and set a clear order: wired or USB output, intentional Bluetooth media, then the phone loudspeaker.
Tic-Tac-Toe and Rock-Paper-Scissors first existed only inside the Developer Hub. I connected validated voice actions to the local game engines without letting the remote model decide the rules.
CameraX, ML Kit, FaceNet, TensorFlow Lite, face enrollment, pose detection, and gestures were tested near the end. Automated tests worked, but the physical phone became less reliable. I removed the entire camera and ML runtime from V1.

I learned that a feature is not useful if it makes the complete system less reliable.
V1 result
Camera-based social awareness, local face detection, recognition of Hassan, local face embeddings, gesture recognition, on-device ML, richer adaptation, and more sensors are possible future experiments. They are not V1 features.
ESP32 · Arduino/C++ · Bluetooth Classic · Adafruit NeoPixel · INA219 support
Kotlin · Jetpack Compose · Material 3 · Android audio APIs · DataStore
TB6612FNG · TT motors · HC-SR04 · TCRT5000 · TTP223 · servos · RGB ring
Android Life Engine · validated actions · ESP32 safety · JSON protocol · OpenAI Realtime
Final demonstration
A real demonstration of the completed AMSAN V1 prototype.