Intelligent glasses for deaf-mute communication
The smart glasses feature built-in vision and speech translation modules, enabling bidirectional translation of speech and sign language between deaf and hearing individuals, thus overcoming communication barriers and achieving barrier-free communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 深圳目渡科技有限公司
- Filing Date
- 2026-05-27
- Publication Date
- 2026-07-07
AI Technical Summary
Communication barriers between deaf and hearing people, especially the communication difficulties caused by hearing people's lack of knowledge of sign language, and the problem of deaf people being unable to communicate normally.
Design a smart glasses system that incorporates a visual translation module, a speech translation module, and a call translation module. It acquires information through a sign language acquisition unit and a sound acquisition unit, and utilizes a sign language recognition model and human voice-environmental sound separation technology to achieve bidirectional real-time translation between speech and sign language. The system transmits information through a near-eye display and a Bluetooth communication module.
It enables two-way communication between deaf and mute individuals and others, freeing their hands, reducing communication costs, reconstructing the portability of barrier-free communication, and preserving the original expressive advantages of sign language.
Smart Images

Figure CN122347896A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart glasses technology, and more particularly to a smart glasses for communication by deaf and mute people. Background Technology
[0002] Currently, deaf and mute individuals are unable to communicate normally through language due to physical reasons, such as hearing loss or inability to speak. They primarily rely on sign language for communication. However, this presents several problems. When deaf and mute individuals use sign language to communicate with hearing individuals, some hearing individuals do not understand the meaning of the sign language and are unable to comprehend what the deaf and mute individuals are trying to express. Conversely, some hearing individuals are unable to use sign language to communicate, and since the deaf and mute individuals cannot hear spoken language, communication becomes impossible. Furthermore, deaf and mute individuals are also unable to engage in normal phone conversations, thus hindering their participation in normal social activities. Summary of the Invention
[0003] The purpose of this invention is to provide a smart glasses for communication by deaf and mute people, which realizes the functions of speech-to-text and sign language-to-speech, thereby enabling deaf and mute people to communicate with others or make phone calls, and allowing deaf and mute people to participate equally in normal social life.
[0004] The technical solution adopted by the smart glasses for communication by deaf and mute people disclosed in this invention is as follows: A smart glasses for communication by deaf and mute people includes: a glasses body, wherein the glasses body is provided with a visual translation module, a voice translation module, a call translation module and a processing unit, wherein the processing unit is electrically connected to the visual translation module, the voice translation module and the call translation module; The speech translation module includes a sign language acquisition unit and a speech playback unit. The sign language acquisition unit acquires the wearer's sign language actions, converts them into natural speech through a sign language recognition model, and plays them out through the speech playback unit, thereby realizing bidirectional real-time translation between speech and sign language. The visual translation module includes a sound acquisition unit and a near-eye display unit. The sound acquisition unit acquires external sound, and after human voice and ambient sound separation and noise reduction processing, it is converted into subtitle information or environmental prompt symbols, which are then projected and displayed in real time through the near-eye display unit. The call translation module includes a Bluetooth communication module, a miniature vibration motor, and a call button. The Bluetooth communication module is connected to a mobile terminal. When a call comes in, pressing the call button converts the call audio into subtitles through the voice translation module and the visual translation module. The wearer's sign language is converted into speech and transmitted back to the other party, thus realizing two-way call translation.
[0005] As a preferred embodiment, the external sounds include human voices and ambient sounds. Human voices are processed and converted into text subtitles or standard sign language patterns for display, while ambient sounds are recognized and converted into preset symbols or text prompts for display.
[0006] As a preferred embodiment, the sign language acquisition unit includes a wide-angle camera and a facial electromyography (EMG) sensor. The wide-angle camera is installed at the front of the frame with a shooting angle of 90°-100°, and the facial EMG sensor is installed on the inside of the nose pad in contact with the skin to collect facial micro-expressions.
[0007] As a preferred embodiment, the sign language semantic recognition weights are: 70% for wide-angle camera information and 30% for facial electromyography information.
[0008] As a preferred option, the sign language recognition model is a lightweight edge-side model trained based on the national common sign language dictionary, which supports offline real-time inference and has a recognition latency of ≤200ms.
[0009] As a preferred embodiment, the near-eye display unit includes a micro-optical engine and a waveguide lens, with the micro-optical engine located in the middle of the frame and the waveguide lens integrated with the eyeglass lens.
[0010] As a preferred option, an online assistance unit is also included, which includes a help button and a remote transmission module. The help button is installed on the surface of the temple. After the help button is pressed, the sign language image and ambient sound are uploaded to the backend server in real time, where they are translated and transmitted back for display or playback with human assistance.
[0011] The beneficial effects of the smart glasses for communication used by deaf and mute people disclosed in this invention are: External sounds are acquired through the sound acquisition unit, and after human voice and environmental sound separation and noise reduction processing, they are converted into subtitle information or environmental prompts and displayed in real time through the near-eye display unit. The sign language acquisition unit acquires the wearer's sign language movements, which are then converted into natural speech through the sign language recognition model and played out through the voice playback unit. This achieves two-way real-time translation between speech and sign language, enabling better information acquisition and content transmission. It also frees up the wearer's hands, making it easier for deaf and mute people to perform sign language operations during communication. At the same time, it recognizes the content of sign language movements, reduces communication costs, and, more importantly, reconstructs the portability of barrier-free communication while retaining the advantages of the original expression of sign language. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the structure of a smart glasses for communication by deaf and mute people according to the present invention.
[0013] Figure 2 This is a schematic diagram of the speech translation module structure of a smart glasses for communication by deaf and mute people according to the present invention.
[0014] Figure 3 This is a schematic diagram of the visual translation module structure of a smart glasses for communication by deaf and mute people according to the present invention. Detailed Implementation
[0015] The present invention will be further described and illustrated below with reference to specific embodiments and the accompanying drawings: Please refer to Figure 1 A smart glasses for communication by deaf and mute individuals includes: a glasses body, which contains a visual translation module, a voice translation module, a call translation module, and a processing unit 3. The processing unit 3 is electrically connected to the visual translation module, the voice translation module, and the call translation module. The processing unit 3 is an end-side main control chip located inside the temple, supporting offline real-time reasoning. It also includes a power supply unit 12, which uses a lithium battery located inside the temple. A magnetic charging port 2 is located at the end of the temple for charging the smart glasses.
[0016] Please refer to Figure 2 The speech translation module includes a sign language acquisition unit and a speech playback unit 10. The sign language acquisition unit acquires the wearer's sign language movements, converts them into natural speech through the sign language recognition model, and plays them out through the speech playback unit 10, realizing bidirectional real-time translation between speech and sign language.
[0017] The sign language acquisition unit includes a wide-angle camera 6 and a facial electromyography sensor 8. The wide-angle camera 6 is installed at the front of the frame with a shooting angle of 90°-100°. The facial electromyography sensor 8 is installed on the inside of the nose pad in contact with the skin to collect facial micro-expressions. It is fused with visual information to improve the accuracy of semantic recognition. The voice playback unit 10 uses a bone conduction speaker or a miniature speaker and is installed at the front of the frame.
[0018] In the sign language semantic recognition weighting, vision is the primary factor and electromyography (EMG) is secondary. For example, wide-angle camera information accounts for approximately 70%, while facial EMG information accounts for approximately 30%. Improving the accuracy of sign language semantic recognition involves combining sign language gestures with the wearer's facial expressions and emotions, thereby enhancing the accuracy of sign language interpretation and enabling the expression of meanings more closely aligned with the wearer's intended message.
[0019] Furthermore, the sign language recognition model is a lightweight edge-side model trained based on the national common sign language dictionary, which supports offline real-time inference and has a recognition latency of ≤200ms.
[0020] The speech translation module addresses the bottleneck of expression by capturing sign language movements with a wide-angle camera. The AI sign language recognition model decodes these movements into natural speech, converting the spatial grammar of sign language into speech for playback. This enables wearers to initiate conversations in scenarios such as shopping, medical consultations, and communication, allowing them to express themselves even when others do not understand sign language.
[0021] The specific workflow of the speech translation module is as follows: wide-angle camera 6 collects sign language + facial electromyography sensor 8 collects facial expressions → data fusion processing → end-side sign language model inference → conversion to natural language → TTS speech synthesis → output by speech playback unit 10.
[0022] Please refer to Figure 3 The visual translation module includes a sound acquisition unit and a near-eye display unit. The sound acquisition unit acquires external sound, and after human voice and ambient sound separation and noise reduction processing, it is converted into subtitle information or environmental prompt symbols, which are then projected and displayed in real time through the near-eye display unit.
[0023] External sounds include human voices and ambient sounds. Human voices are processed and converted into text subtitles or standard sign language patterns for display. Ambient sounds are recognized and converted into preset symbols or text prompts for display, such as horns, doorbells, telephone rings, alarms, etc.
[0024] Sound acquisition unit: Dual microphone matrix 4, installed on the outside of the left and right temples, to achieve directional sound pickup and environmental noise reduction. Near-eye display unit includes miniature optical engine 7 and optical waveguide lens 9. Miniature optical engine 7 is placed in the middle of the frame, and optical waveguide lens 9 is integrated with the eyeglass lens. Subtitle information and sign language translation text are displayed synchronously for the wearer to check.
[0025] The specific workflow of the visual translation module is as follows: microphone matrix 4 collects sound → human voice / ambient sound separation algorithm → noise reduction processing → speech to text → display by optical waveguide lens 9; among which, ambient sound is converted into text or icon prompts.
[0026] The call translation module includes a Bluetooth communication module, a miniature vibration motor 1, and a call button 5, all integrated on one side of the temple. The Bluetooth communication module connects to a mobile terminal. When a call comes in, pressing the call button 5 converts the call audio into subtitles through the voice translation module and the visual translation module. The wearer's sign language is converted into voice and transmitted back to the other party, thus realizing two-way call translation.
[0027] In actual use, the call translation module uses a Bluetooth communication module to connect the smart glasses to the mobile phone, thereby feeding back the call request received on the mobile phone to the miniature vibration motor 1 of the smart glasses, reminding the wearer that there is a call request.
[0028] By translating the voice content of a call into subtitles, the wearer can easily understand the content of the call. At the same time, the wide-angle camera 6 translates the wearer's sign language into voice content and plays it to the caller, thereby realizing two-way communication and enabling the wearer to communicate with others in different ways.
[0029] The specific workflow of the call translation module is as follows: Bluetooth connection to the mobile phone → incoming call triggers micro-vibration → press the call button 5 to answer the call → the other party's voice is converted into subtitles and displayed → the wearer's sign language is converted into voice and transmitted back to the other party.
[0030] It also includes an online assistance unit, which comprises an assistance button 11 and a remote transmission module. The assistance button 11 is installed on the surface of the temple. After pressing the assistance button 11, sign language images and ambient sounds are uploaded to the backend server in real time, where they are translated and transmitted back for display or playback with human assistance.
[0031] In special circumstances, such as when sign language gestures cannot be effectively recognized or when the collected sound cannot be converted into accurate content, the wearer can actively trigger the help button 11 to transmit the information data collected by the wide-angle camera 6 and microphone matrix 4 to the human system in real time. The human customer service in the background processes the transmitted data, converts the captured sign language gestures into text, and then plays it through the voice playback unit 10; or after manually processing the collected sound, the corresponding information is presented to the wearer through the optical waveguide lens 9, thereby achieving more accurate content translation and avoiding ambiguity.
[0032] The specific workflow of the online assistance unit is as follows: Press the help button 11 → Upload video and audio to the backend server in real time → Manual translation assistance → Translation result feedback → Display or playback.
[0033] This invention provides smart glasses for communication by deaf and mute individuals. It acquires external sound through a sound acquisition unit, processes it through human voice and ambient sound separation and noise reduction, and converts it into subtitle information or environmental prompts. This information is then projected and displayed in real-time through a near-eye display unit. A sign language acquisition unit acquires the wearer's sign language gestures, which are then converted into natural speech through a sign language recognition model and played back through a speech playback unit. This achieves bidirectional real-time translation between speech and sign language, improving information acquisition and content delivery. It also frees the wearer's hands, making it easier for deaf and mute individuals to use sign language during communication. Furthermore, it recognizes sign language gestures, reducing communication costs. More importantly, while retaining the advantages of native sign language expression, it reconstructs the portability of barrier-free communication.
[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A pair of smart glasses for communication by deaf and mute individuals, characterized in that: include: The glasses body contains a visual translation module, a voice translation module, a call translation module, and a processing unit. The processing unit is electrically connected to the visual translation module, the voice translation module, and the call translation module. The speech translation module includes a sign language acquisition unit and a speech playback unit. The sign language acquisition unit acquires the wearer's sign language actions, converts them into natural speech through a sign language recognition model, and plays them out through the speech playback unit, thereby realizing bidirectional real-time translation between speech and sign language. The visual translation module includes a sound acquisition unit and a near-eye display unit. The sound acquisition unit acquires external sound, and after human voice and ambient sound separation and noise reduction processing, it is converted into subtitle information or environmental prompt symbols, which are then projected and displayed in real time through the near-eye display unit. The call translation module includes a Bluetooth communication module, a miniature vibration motor, and a call button. The Bluetooth communication module is connected to a mobile terminal. When a call comes in, pressing the call button converts the call audio into subtitles through the voice translation module and the visual translation module. The wearer's sign language is converted into speech and transmitted back to the other party, thus realizing two-way call translation.
2. The smart glasses for communication by deaf and mute individuals as described in claim 1, characterized in that, The external sounds include human voices and ambient sounds. Human voices are processed and converted into text subtitles or standard sign language patterns for display. Ambient sounds are recognized and converted into preset symbols or text prompts for display.
3. The smart glasses for communication by deaf and mute individuals as described in claim 1, characterized in that, The sign language acquisition unit includes a wide-angle camera and a facial electromyography (EMG) sensor. The wide-angle camera is installed at the front of the frame with a shooting angle of 90°-100°. The facial EMG sensor is installed on the inside of the nose pad in contact with the skin and is used to collect facial micro-expressions.
4. The smart glasses for communication by deaf and mute people as described in claim 3, characterized in that, The sign language recognition model is a lightweight edge-side model trained based on the national common sign language dictionary, which supports offline real-time inference and has a recognition latency of ≤200ms.
5. The smart glasses for communication by deaf and mute individuals as described in claim 1, characterized in that, The near-eye display unit includes a micro-optical engine and a waveguide lens. The micro-optical engine is placed in the middle of the frame, and the waveguide lens is integrated with the eyeglass lens.
6. The smart glasses for communication by deaf and mute individuals as described in claim 1, characterized in that, It also includes an online assistance unit, which includes a help button and a remote transmission module. The help button is installed on the surface of the temple. After the help button is pressed, the sign language image and ambient sound are uploaded to the backend server in real time, where they are translated and transmitted back for display or playback by human assistance.