Real-time glasses system and device based on neural network machine translation

By adopting a neural network-based machine translation real-time glasses system in translation devices, integrating multiple modules to support multilingual hybrid processing and privacy protection, the challenges of existing translation devices in real-time, scene adaptability and user experience are solved, and translation coherence and privacy security are achieved in complex scenarios.

CN119987034AInactive Publication Date: 2025-05-13CHANGCHUN INST OF ELECTRONIC TECH

Patent Information

Application Number
CN202510482095.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing translation devices have significant challenges in real-time, scenario adaptability and user experience, especially in network-free environments, and are difficult to deal with professional terms and semantic coherence of multiple rounds of dialogues, and cannot meet the localized processing needs of sensitive information in business and other scenarios.

Method used

It adopts a machine translation real-time glasses system based on neural networks, integrates end-to-end neural network processing pipelines, dynamic context caching modules, user sight tracing modules and privacy protection modules, and supports multilingual hybrid processing, real-time response and localized data processing.

Benefits of technology

It realizes accurate voice acquisition and visual text positioning, ensures translation coherence in complex scenarios, and achieves real-time response, multilingual hybrid processing and privacy and security effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987034A_ABST
    Figure CN119987034A_ABST
Patent Text Reader

Abstract

The invention relates to the field of translation glasses, in particular to a real-time glasses system and device based on neural network machine translation. The technical problem to be solved by the invention is to provide accurate voice acquisition and visual text positioning to ensure translation coherence in a complex scene. The neural network machine translation-based real-time glasses system and device can realize real-time response, multi-language mixed processing and privacy security. The invention discloses a real-time glasses system and device based on neural network machine translation. The real-time glasses system comprises a beam forming microphone array, a micro transparent display screen, a bone conduction loudspeaker, a front camera, an AI processor, an end-to-end neural network processing assembly line, a dynamic context caching module, a user sight tracking module, a privacy protection module and a user interaction module. According to the invention, accurate voice acquisition and visual text positioning are achieved, and translation coherence in a complex scene is ensured; and the effects of real-time response, multi-language mixed processing and privacy security are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of translation glasses, and in particular to a real-time glasses system and device based on neural network machine translation. Background Art

[0002] As the process of globalization accelerates, the demand for cross-language communication has shown explosive growth. Traditional translation equipment faces significant challenges in real-time performance, scene adaptability and user experience. In the existing technology, handheld translation machines rely on cloud processing, resulting in response delays and limited functions in a network-free environment. Although smart glasses have AR display capabilities, they have problems such as text blocking the field of view and insufficient multi-language mixed processing capabilities. More importantly, traditional voice translation systems lack understanding of dialogue context, and have difficulty processing professional terms and the semantic coherence of multiple rounds of dialogues, and have low recognition accuracy in noisy environments. At the same time, in terms of privacy protection, existing solutions mostly use cloud data desensitization processing, which cannot meet the needs of sensitive information localization processing in business and other scenarios. Therefore, it is urgent to develop a precise voice acquisition and visual text positioning to ensure translation coherence in complex scenarios; to achieve real-time response, multi-language mixed processing and privacy security based on neural network machine translation real-time glasses system and device. Summary of the invention

[0003] In order to overcome the shortcomings of existing voice translation systems and devices, such as text blocking the field of view, insufficient multi-language mixed processing capabilities, lack of understanding of dialogue context, difficulty in processing professional terms and semantic coherence of multiple rounds of dialogues; and in terms of privacy protection, the technical problem to be solved by the present invention is to provide a precise voice collection and visual text positioning to ensure translation coherence in complex scenarios; to achieve real-time response, multi-language mixed processing and privacy security based on neural network machine translation real-time glasses system and device.

[0004] The present invention is achieved by the following specific technical means: A real-time glasses system based on neural network machine translation, comprising Real-time translation processing system, including: End-to-end neural network processing pipeline, integrating speech recognition model ASR, neural network machine translation model NMT, and speech synthesis model TTS; Dynamic context cache module, used to store conversation history to improve the coherence of multiple rounds of translation; automatic clearing strategy based on the timeliness of conversation topics, setting the maximum storage time to no longer than the current conversation period; The user gaze tracking module identifies the visual text area to be translated based on eye movement; the pupil corneal reflection method and the micro infrared LED array form a three-dimensional gaze positioning system; Privacy protection module, used to ensure user data security; The user interaction module is used to receive user input instructions.

[0005] Furthermore, the neural network machine translation model NMT adopts a bidirectional encoder-decoder structure based on the Transformer architecture and is compressed into a lightweight model through knowledge distillation technology; The neural network machine translation model NMT supports offline mode operation, has a built-in multi-language translation parameter library, and covers the mutual translation function of at least 12 languages; The training data of the neural network machine translation model NMT includes a dialect and accent enhancement data set, and an adversarial training strategy is used to improve robustness.

[0006] Furthermore, the privacy protection module includes: The local voice data encryption storage unit uses the AES-256 algorithm to encrypt the original voice stream; Automatic erasure mechanism after data processing is completed, ensuring that voice and text data reside in memory for less than 1 second; The federated learning interface supports updating model parameters through encrypted gradients, prohibits uploading raw data to the cloud, and adopts a gradient aggregation method with differential privacy encryption.

[0007] Furthermore, the real-time translation processing system also includes: Real-time display optimization module, which dynamically calculates the text display area based on the attention mechanism to avoid blocking the user's key field of view; Multi-language mixed processing module, supporting recognition and translation of two languages ​​mixed in a single sentence; double-layer LSTM network structure integrating language classifier and word segmentation boundary detection; Emergency mode automatically highlights and repeats specific keywords when they are detected.

[0008] Furthermore, the user interaction module includes: The gesture recognition unit uses the infrared sensor integrated in the temple to recognize swipe and click gestures to switch translation modes; Voice command set, supporting natural language commands; The tactile feedback module prompts the translation status change through the vibration of the temples.

[0009] Furthermore, the neural network machine translation model NMT is deployed and optimized in the following ways: Use the TensorRT engine to quantize the model and convert floating-point calculations into 8-bit integer operations; Use layer fusion technology to merge adjacent neural network layers to reduce the number of memory accesses; The processor utilization rate is improved through dynamic batch processing technology, and the end-to-end delay is maintained in a degraded mode of ≤450ms under 80dB environmental noise.

[0010] Furthermore, the training data enhancement method of the neural network machine translation model includes: Use the voice conversion network VC to generate accented speech training data; Improve the model's robustness to noise and speech variations through adversarial sample generator; A curriculum learning strategy is adopted to train multilingual translation models in stages according to language difficulty.

[0011] Furthermore, a real-time eyewear device based on neural network machine translation includes Smart glasses device, comprising: The beamforming microphone array integrated into the temple is used to collect user voice signals in a directionally manner; The micro transparent display screen on the inside of the lens supports augmented reality (AR) overlay display; The bone conduction speaker located at the rear end of the temple, opposite to the back of the ear, is used to play the translated voice signal; A front-facing camera located on the front side of the frame, used to capture text or image information within the user's field of view; The AI ​​processor embedded in the temple is used to locally perform speech recognition, machine translation and speech synthesis tasks.

[0012] Furthermore, the smart glasses device also includes: Environmental noise suppression module, which separates the target speech from the environmental noise through an adaptive filtering algorithm; A multi-modal sensor array, including a 9-axis IMU inertial measurement unit and a distance sensor, is used to dynamically adjust the position of the display text to prevent visual jitter; The battery module is replaceable, supports hot swapping and has a battery life of greater than or equal to 8 hours.

[0013] Furthermore, the augmented reality display function of the micro transparent display screen includes: Text rendering engine that automatically adjusts display contrast and font color according to ambient lighting; Spatial anchoring technology, which binds the translated text to real-world objects; Multi-language parallel display mode supports simultaneous display of translation results in less than or equal to 3 languages.

[0014] Compared with the prior art, the present invention has the following beneficial effects: The present invention achieves accurate voice collection and visual text positioning, ensuring translation coherence in complex scenarios; and realizes the effects of real-time response, multi-language mixed processing and privacy security.

[0015] The technical solution of the present invention adopts a hardware-algorithm collaborative innovation system; constructs an end-to-end lightweight NMT model to achieve low-latency localized reasoning on an embedded processor; develops a multimodal perception fusion architecture that integrates gaze tracking, environmental perception, and dynamic display optimization algorithms; and uses a hybrid encryption privacy protection mechanism to achieve full life cycle security management of voice data "generation-processing-destruction" to enable cross-language communication to break through the limitations of the physical environment and network conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a three-dimensional structural diagram of the smart glasses device of the present invention.

[0017] Figure 2 It is a hierarchical topology diagram of the system architecture of the present invention.

[0018] Figure 3 This is a structural diagram of the user interaction module of the present invention.

[0019] Figure 4 This is a structural diagram of the privacy protection module of the present invention.

[0020] Figure 5 It is a schematic diagram of the data processing flow of the present invention.

[0021] The markings in the attached figure are: 1-microphone array, 2-transparent display, 3-bone conduction speaker, 4-front camera, 5-AI processor. DETAILED DESCRIPTION

[0022] The present invention is further described below in conjunction with the accompanying drawings: Example

[0023] A real-time glasses system and device based on neural network machine translation, such as Figure 1-5 As shown, including: Smart glasses device, comprising: The beamforming microphone array 1 integrated in the temple is used to directionally collect the user's voice signal; The micro transparent display 2 located inside the lens supports augmented reality (AR) overlay display; A bone conduction speaker 3 located at the rear end of the temple, opposite to the back of the ear, is used to play the translated voice signal; A front camera 4 located on the front side of the frame, used to capture text or image information within the user's field of view; AI processor 5 embedded in the temple for localized voice recognition, machine translation and speech synthesis tasks; Real-time translation processing system, including: End-to-end neural network processing pipeline, integrating speech recognition model ASR, neural network machine translation model NMT, and speech synthesis model TTS; Dynamic context cache module, used to store conversation history to improve the coherence of multiple rounds of translation; automatic clearing strategy based on the timeliness of conversation topics, setting the maximum storage time to no longer than the current conversation period; The user gaze tracking module identifies the visual text area to be translated based on eye movement; the pupil corneal reflection method and the micro infrared LED array form a three-dimensional gaze positioning system; Privacy protection module, used to ensure user data security; A user interaction module, used to receive user input instructions; The privacy protection module further includes: a local encrypted storage unit for voice data, which uses the AES-256 algorithm to encrypt the original voice stream; an automatic erasure mechanism after data processing is completed to ensure that the voice and text data reside in the memory for less than 1 second; a federated learning interface that supports updating model parameters through encrypted gradients and prohibits uploading original data to the cloud; The smart glasses device also includes: an environmental noise suppression module, which separates the target voice from the environmental noise through an adaptive filtering algorithm; a multimodal sensor array, including a 9-axis IMU inertial measurement unit and a distance sensor, which is used to dynamically adjust the position of the text on the display screen to prevent visual jitter; a replaceable battery module, which supports hot swapping and has a battery life of greater than or equal to 8 hours; The real-time translation processing system also includes: a real-time display optimization module, which dynamically calculates the text display area based on the attention mechanism to avoid blocking the user's key field of view; a multi-language mixing processing module, which supports the recognition and translation of two languages ​​mixed in a single sentence; an emergency mode, which automatically highlights and repeats specific keywords when they are detected; The user interaction module includes: a gesture recognition unit, which recognizes swiping and clicking gestures to switch translation modes through infrared sensors integrated in the temples; a voice command set, which supports natural language commands; and a tactile feedback module, which prompts translation status changes through temple vibrations.

[0024] Example 1: Basic real-time translation scenario Hardware startup and voice collection: After the user puts on the smart glasses, the beamforming microphone array is automatically activated and locks the sound source within 1.5 meters in front of it through an adaptive beamforming algorithm (the signal-to-noise ratio is improved by 15dB). The front camera scans the field of view at a frame rate of 30fps, and when it detects that the user continues to stare at a foreign language logo for more than 2 seconds, the OCR recognition mode is triggered. Embedded AI processor 5 performs: The ASR module converts the collected English speech into text (recognition delay 120ms, accuracy 98.2%); NMT model performs English-to-Chinese translation: It uses an 8-layer Transformer structure (4 layers for encoder / 4 layers for decoder) and handles long-distance dependencies through a multi-head attention mechanism; The TTS module generates Chinese speech, and the bone conduction speaker 3 outputs at a sound pressure level of 45dB, while the vibration frequency of the ear contact point matches the 300-3000Hz speech frequency band; The micro transparent display screen 2 projects the translation result on the upper right side of the lens; Voice dialogue scene: display Chinese and English subtitles (font size automatically adapts to 6-12pt); Visual text translation: Replace the "EXIT" sign with "Exit" in real time, and keep the text aligned with the real scene 1:1 through spatial anchoring technology; The tactile feedback module generates a short vibration prompt of 50ms when the translation is completed.

[0025] Example 2: Multi-language conference scenario Enhanced processing in complex environments: In a noisy conference room, the environmental noise suppression module is activated, and the NLMS algorithm is used to eliminate the air conditioning noise (noise reduction of 20dB); the multimodal sensor detects that the user frequently turns his head, and the IMU data triggers the display anti-shake algorithm (the position correction frequency is increased to 120Hz); Mixed language real-time translation: Processing a sentence mixed in Chinese and English: "We need to align the timelines of various departments"; the multilingual mixed input module identifies language boundaries, and the NMT model performs mixed translation, translating "align" into "alignment" and "timeline" into "timeline"; and dynamically caches the context for reference in subsequent translations; Privacy protection mechanism operation: The AES-256 encryption unit encrypts the voice stream in real time (256-bit key, CBC mode); after data processing is completed, the memory erasure module immediately clears the relevant buffer (erasing speed reaches 5GB / s); the federated learning interface uploads encrypted gradients (homomorphic encryption parameter δ=1e-5) to update the model without leaking the original data; Example 3: Offline travel scenario Optimized operation without network: Enable offline mode in remote areas and load lightweight NMT models (compress model size to 150MB, pre-store parameter libraries for 12 languages); enable INT8 quantized computing for the TensorRT engine (inference speed increased by 2.3 times, power consumption reduced by 40%); Dynamic display optimization: When browsing the French menu under strong light, the text rendering engine automatically switches to high contrast mode (background transparency is reduced to 20%, and the font is changed to pure white); the attention mechanism analyzes the visual focus and intelligently avoids the translated text from the dish image (the occlusion area is reduced by 70%); the Chinese, Japanese and Korean translations are displayed in parallel in multiple languages ​​(the font spacing is adaptively adjusted to 1.2 times the line height); Emergency Mode Response: When the local language "fire" is detected, the emergency mode is automatically activated and the display switches to a red border warning; the translation result "Fire emergency" is displayed in the center with a font size of 200%; the bone conduction speaker (3) plays the alarm voice in a loop at the maximum volume (repetition frequency 3 times / second).

[0026] Although the present disclosure has been described in detail with reference to the exemplary embodiments, the present disclosure is not limited thereto, and it will be apparent to those skilled in the art that various modifications and changes may be made thereto without departing from the scope of the present disclosure.

Claims

1. A real-time glasses system based on neural network machine translation, characterized in that: Included Real-time translation processing system, including: End-to-end neural network processing pipeline, integrating speech recognition model ASR, neural network machine translation model NMT, and speech synthesis model TTS; Dynamic context cache module, used to store conversation history to improve the coherence of multiple translation rounds; A user gaze tracking module that identifies the visual text area to be translated based on eye movements; Privacy protection module, used to ensure user data security; The user interaction module is used to receive user input instructions.

2. The real-time glasses system based on neural network machine translation according to claim 1, characterized in that: The neural network machine translation model adopts a bidirectional encoder-decoder structure based on the Transformer architecture and is compressed into a lightweight model through knowledge distillation technology; The neural network machine translation model supports offline mode operation, has a built-in multi-language translation parameter library, and covers the translation function of at least 12 languages; The training data of the neural network machine translation model includes a dialect and accent enhancement data set, and an adversarial training strategy is used to improve robustness.

3. The real-time glasses system based on neural network machine translation according to claim 1, characterized in that: The privacy protection module further comprises: The local voice data encryption storage unit uses the AES-256 algorithm to encrypt the original voice stream; Automatic erasure mechanism after data processing is completed, ensuring that voice and text data reside in memory for less than 1 second; The federated learning interface supports updating model parameters through encrypted gradients and prohibits uploading raw data to the cloud.

4. The real-time glasses system based on neural network machine translation according to claim 1, characterized in that: The real-time translation processing system also includes: Real-time display optimization module, which dynamically calculates the text display area based on the attention mechanism to avoid blocking the user's key field of view; Multi-language mixed processing module, supporting recognition and translation of two languages ​​mixed in a single sentence; Emergency mode automatically highlights and repeats specific keywords when they are detected.

5. The real-time glasses system based on neural network machine translation according to claim 1, characterized in that: The user interaction module further comprises: The gesture recognition unit uses the infrared sensor integrated in the temple to recognize swipe and click gestures to switch translation modes; Voice command set, supporting natural language commands; The tactile feedback module prompts the translation status change through the vibration of the temples.

6. The real-time glasses system based on neural network machine translation according to claim 1, characterized in that: The neural network machine translation model is deployed and optimized in the following ways: Use the TensorRT engine to quantize the model and convert floating-point calculations into 8-bit integer operations; Use layer fusion technology to merge adjacent neural network layers to reduce the number of memory accesses; The processor utilization rate is improved through dynamic batch processing technology, and the end-to-end delay is maintained in a degraded mode of ≤450ms under 80dB environmental noise.

7. The real-time glasses system based on neural network machine translation according to claim 2, characterized in that: The training data enhancement method of the neural network machine translation model includes: Use the voice conversion network VC to generate accented speech training data; Improve the model's robustness to noise and speech variations through adversarial sample generator; A curriculum learning strategy is adopted to train multilingual translation models in stages according to language difficulty.

8. A real-time eyewear device based on neural network machine translation, characterized in that: Included Smart glasses device, comprising: A beamforming microphone array (1) integrated into the temples of the glasses is used to directionally collect user voice signals; A micro transparent display (2) located inside the lens supports augmented reality (AR) overlay display; A bone conduction speaker (3) located at the rear end of the temple, relative to the back of the ear, for playing the translated speech signal; A front camera (4) located on the front side of the frame, used to capture text or image information within the user's field of view; The AI ​​processor (5) embedded in the temple is used to locally perform speech recognition, machine translation and speech synthesis tasks.

9. The real-time glasses device based on neural network machine translation according to claim 8, characterized in that: The smart glasses device further comprises: Environmental noise suppression module, which separates the target speech from the environmental noise through an adaptive filtering algorithm; A multi-modal sensor array, including a 9-axis IMU inertial measurement unit and a distance sensor, is used to dynamically adjust the position of the display text to prevent visual jitter; The battery module is replaceable, supports hot swapping and has a battery life of greater than or equal to 8 hours.

10. The real-time glasses device based on neural network machine translation according to claim 8, characterized in that: The augmented reality display function of the micro transparent display screen (2) includes: Text rendering engine that automatically adjusts display contrast and font color according to ambient lighting; Spatial anchoring technology, which binds the translated text to real-world objects; Multi-language parallel display mode supports simultaneous display of translation results in less than or equal to 3 languages.

Citation Information

Patent Citations

  • Portable offline machine translation intelligent box

    CN114444523A

  • Real-time translation and prompt AR / VR glasses and working method thereof

    CN115617179A

  • Multi-modal dialogue translation method and device, electronic equipment and storage medium

    CN116663574A

  • Real-time translation method and system of intelligent wearable device and medium

    CN116911323A

  • Intelligence translation glasses

    CN204679734U

Cited By

  • A head-mounted real-time speech translation control method and system based on smart glasses

    CN122635381A