Sign-to-Text uses artificial intelligence developed by Google DeepMind to interpret American Sign Language (ASL) through the front-facing camera and generate English text in near real time.
Google is bringing artificial intelligence to a new form of interaction. With the Pixel 11, the company introduced Sign-to-Text, a feature capable of interpreting American Sign Language (ASL) through the smartphone’s front-facing camera and converting the message into English text in near real time.
The feature is designed to facilitate communication between people who use ASL and those who do not understand the language, with the smartphone acting as an intermediary.
Behind the technology is Google DeepMind, which developed a model called SL2T (Sign Language-to-Text) specifically to interpret sign languages and generate written text.
An AI That Looks Beyond Hand Movements
One of the main challenges is that sign language cannot be understood simply by identifying specific hand movements.
Sign languages are fully developed languages with their own structures and grammar. Meaning also depends on elements such as facial expressions, body position and movement. SL2T therefore analyzes multiple visual cues to interpret the message more comprehensively.
This distinction is crucial: the goal is not to match each gesture with an English word, but to develop models capable of understanding a visual language and transforming it into text.
How Does Sign-to-Text Work on the Pixel 11?
Another notable aspect of the technology is the way it processes the information captured by the camera.
Google uses MediaPipe Holistic to analyze video directly on the device and identify reference points across the hands, face and body. These points are converted into geometric data that can then be processed by the model.
As a result, the system does not need to rely on the complete video throughout the entire process. According to Google DeepMind, the original footage is immediately discarded once the necessary information has been extracted.
This approach is intended to reduce the amount of visual data the system must process while addressing one of the most sensitive concerns surrounding the constant use of a camera during an interaction: privacy.
More Than 100,000 Hours of Sign-Language Data Used to Train the AI
To develop SL2T, Google DeepMind used more than 100,000 hours of data spanning over 50 sign languages. This scale allowed the team to improve the model’s ability to understand patterns in visual communication.
However, this does not mean that the Pixel 11 can currently translate more than 50 sign languages.
The initial release of Sign-to-Text focuses on American Sign Language (ASL) and English text generation. The inclusion of other sign languages in the training data should therefore not be confused with the languages currently supported by the feature.
Artificial Intelligence and Accessibility
Sign-to-Text points to another potential direction for artificial intelligence on personal devices: using multimodal models not only to recognize objects, images or voices, but also to interpret forms of communication that rely on movement and physical expression.
It also presents a considerable challenge. Sign languages are complex linguistic systems that vary across countries and communities. Extending this technology to other sign languages will therefore require far more than simply recognizing new gestures.
With Sign-to-Text, Google turns the smartphone’s front-facing camera into a communication interface. But the most compelling development lies behind the feature: training artificial intelligence to understand forms of language that traditional digital interfaces were historically never designed to interpret.


