Intelligent glasses system and method based on natural language processing (NLP) and multi-scene fusion
By introducing high-resolution displays, multi-microphone array noise reduction, ARM architecture chips and deep learning NLP models into smart glasses, the problems of single function and insufficient NLP application of smart glasses have been solved, multi-scenario adaptability and efficient interaction have been achieved, and user experience and efficiency have been improved.
Patent Information
- Application Number
- CN202510268185.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-09-19
AI Technical Summary
Existing smart glasses have single functions and cannot meet user needs in multiple scenarios. The insufficient application of natural language processing (NLP) technology leads to unsmooth voice interaction and inaccurate understanding.
It adopts high-resolution micro-display, multi-microphone array noise reduction technology, ARM architecture high-performance chip, deep learning NLP model, combined with scene recognition and multi-scene application modules to achieve multi-scene adaptability and efficient interaction.
It realizes intelligent and convenient interactive experience in different scenarios, improves the practicality and efficiency of learning, work and life, and enhances the accuracy of speech recognition and semantic understanding.
Smart Images

Figure CN120670567A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart wearable devices, and in particular to a smart glasses system and method based on natural language processing (NLP) and multi-scene fusion Background Art
[0002] With the continuous advancement of technology, smart wearable devices are becoming increasingly popular. Smart glasses, as a key branch of this category, have broad application prospects. However, existing smart glasses have relatively limited functions and cannot meet the diverse needs of users in different scenarios. In learning scenarios, they cannot effectively assist users with language learning; in work scenarios, they cannot integrate well with the office environment to provide convenient information interaction services; and in daily life scenarios such as shopping and travel, they also lack intelligent support. In addition, the application of natural language processing technology in smart glasses is not mature enough, resulting in problems such as unsmooth voice interaction and inaccurate comprehension. Therefore, it is of great practical significance to develop a smart glasses system and method that can integrate multi-scenario applications and achieve efficient and intelligent interaction using natural language processing (NLP) technology. Summary of the Invention
[0003] The purpose of the present invention is to provide a smart glasses system and method based on natural language processing (NLP) and multi-scenario fusion to solve the problems of existing smart glasses having single functions, being unable to meet the needs of multiple scenarios, and insufficient application of natural language processing (NLP), so as to realize intelligent and convenient interactive experience for users in different scenarios.
[0004] The hardware structure includes a display module, an audio module, a processing module, and a sensor module. The display module utilizes a high-resolution, low-power micro-display using MicroLED technology, ensuring clear display in all lighting conditions and providing intuitive visual feedback to the user. The audio module is equipped with a high-sensitivity microphone that utilizes the latest noise reduction algorithms to accurately distinguish user voice from ambient noise, enabling clear voice capture even in noisy environments such as subways and shopping malls. The high-sensitivity microphone utilizes multi-microphone array technology. The device described herein has multiple microphones, allowing simultaneous capture of both ambient noise and target sounds. However, the intensity and phase of the noise and target sounds received by different microphones will vary. By analyzing these differences and utilizing signal processing algorithms, the noise can be separated and suppressed from the mixed signal. The sound source location is determined based on the time difference (TDOA) or phase difference (PDOA) of the sound arriving at different microphones. The noise reduction algorithm is an inverse phase localization algorithm, which uses an array of multiple microphones to collect noise signals and transmits the collected noise signals to a digital signal processor (DSP) or computer for processing. Using algorithms such as Fourier transforms, the time-domain noise signal is converted to the frequency domain for clearer analysis of its frequency components and phase characteristics. Based on the analyzed noise phase information, a signal with an opposite phase is generated. This generated anti-phase signal is played through a speaker. In the smart glasses of this invention, the speaker is located close to the ear, allowing the anti-phase sound waves it emits to fully interfere with the noise in the target area. This allows the anti-phase sound waves to cancel out the incoming noise within the ear canal, achieving noise reduction. The processing module is based on a high-performance chip with ARM architecture. Its multi-core processing capabilities are powerful, and it features an advanced cache management mechanism that allows for rapid access to all types of data required by the NLP algorithm, significantly reducing computation time. For multi-scenario data processing, the chip utilizes a specially optimized parallel computing architecture, enabling efficient and simultaneous processing of data from multiple modules, including sensors, audio, and display. The sensor module includes a GPS module and an ambient light sensor. The GPS module is used for positioning, providing location information for travel, navigation, and other scenarios; the ambient light sensor automatically adjusts the display brightness based on ambient light.
[0005] The software modules include an NLP engine, a scene recognition module, and a multi-scenario application module. The NLP engine, based on a deep learning framework, constructs an advanced NLP model. This model is capable of performing operations such as speech recognition, semantic understanding, and text generation. The scene recognition module determines the user's context by analyzing sensor data, user behavior patterns, and environmental information. For example, in a learning scenario, if the user is detected in an environment such as a library or classroom and the device frequently receives learning-related commands, it is considered a learning scenario. In a travel scenario, GPS data and the user's use of the navigation function are used to identify the travel scenario. The multi-scenario application module includes learning scenario applications, learning scenario applications, and life scenario applications. The learning scenario application provides language learning assistance functions, including real-time translation. When reading foreign books or watching foreign language videos, users can use a camera to record or directly recognize speech, and use NLP technology to achieve instant translation. A word query and memory function allows users to speak or enter a word, and the system provides detailed explanations and examples. It also creates a personalized learning plan based on the memory curve to help users deepen their memory. The work scenario application integrates with office software to receive and display information such as email reminders and schedules. It also supports voice conferencing, enabling clear voice communication through a microphone and speaker, while leveraging NLP technology for real-time transcription and summary generation of meeting content. In the travel scenario, the smart glasses, combined with a high-precision GPS positioning module, can accurately obtain the user's location information and provide real-time road navigation, bus and subway transfer information, and more.
[0006] The interaction method consists of four phases: initialization, scene recognition, interaction processing, and feedback and update. In the initialization phase (S1), after the smart glasses are powered on, the hardware module performs a self-check, ensuring that all sensors, display modules, and audio modules are ready. The software module loads the NLP engine, scene recognition module, and multi-scenario application module, while simultaneously connecting to the network to ensure access to the latest language model data and scene-related information. In the scene recognition phase (S2), the sensor module continuously collects data, which the scene recognition module analyzes and processes. The GPS module determines movement speed and direction. If the speed is high and the direction continuously changes, the user may be traveling; if the ambient light is low and the sound environment is quiet, the user may be studying or resting indoors. Simultaneously, the user's operational behavior and voice commands are analyzed to comprehensively determine the user's current scenario. In the interaction processing phase (S3), when the user issues a voice command, the audio module collects the voice signal and transmits it to the NLP engine for speech recognition and semantic understanding. If the user is identified as being in a learning context and the command is a word query, the NLP engine queries a local or online vocabulary to obtain word definitions and examples, plays them back to the user via the audio module, and displays the relevant information on the display module. If the user is in a work context and the command is to check emails, the system calls the office software interface to obtain the email list and displays it on the display module. The user can then use voice commands to read and reply to emails. During the interaction process, the NLP engine continuously optimizes its understanding of the user's speech and semantics, improving the accuracy and fluency of the interaction. In the feedback and update phase S4, the system optimizes and updates the NLP model and scene recognition model based on the user's operations and feedback. For example, if the user is dissatisfied with the translation result, the system records the relevant information for subsequent optimization of the translation model. If a scene recognition error occurs, the parameters of the scene recognition algorithm are adjusted based on the actual situation. Simultaneously, the system regularly obtains the latest language data and scene feature data from the server to update the local model and improve system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 Schematic diagram of the smart glasses structure.
[0008] Figure 2 It is a functional flow chart. DETAILED DESCRIPTION
[0009] A smart glasses system and method based on natural language processing (NLP) and multi-scene fusion. Its hardware modules include an input module, a data processing and interaction module, an output and display module, and a power management module. The input module includes a microphone, a preamplifier, and an analog-to-digital converter (ADC); the data processing and interaction module includes a speech preprocessing chip, a digital signal processor (DSP), a main processor (CPU + GPU), memory (RAM), a storage module (ROM), and a network communication module (Wi-Fi / 4G / 5G). A positioning module; the output and display module includes a display driver chip and display, a speech synthesis chip, a digital-to-analog converter (DAC), and a speaker; and the power management module includes a battery and a power management chip.
[0010] Specifically, the output module includes a microphone, a preamplifier, and an analog-to-digital converter (ADC). The microphone is used to collect voice signals from the surrounding environment in all directions. In the present invention, the smart glasses are equipped with multiple high-precision microphones to form a microphone array. In a noisy environment, the microphone array uses beamforming technology to focus on sounds in a specific direction, suppress noise interference from other directions, and accurately capture the user's voice commands. The preamplifier amplifies the voice signal collected from the microphone and increases the signal strength to a level range suitable for further processing. At the same time, it also performs preliminary filtering on the signal to remove some high-frequency or low-frequency noise, providing a purer voice signal for subsequent processing. The analog-to-digital converter (ADC) converts the analog voice signal processed by the preamplifier into a digital signal. Only digital signals can be recognized and processed by the digital processing system of the smart glasses.
[0011] Specifically, the data processing and interaction module includes a voice preprocessing chip, a digital signal processor (DSP), a main processor (CPU + GPU), memory (RAM), a storage module (ROM), a network communication module (Wi-Fi / 4G / 5G), and a positioning module. The voice preprocessing chip performs a series of preprocessing operations on the digital voice signal. Using a noise suppression algorithm, it further reduces the impact of ambient noise on the voice signal, improving voice clarity. It also employs voice enhancement technology to highlight key features in the voice, enhancing subsequent voice recognition accuracy. The chip may also perform frame processing on the voice signal, dividing the continuous voice signal into multiple short frames for subsequent voice feature extraction. The digital signal processor (DSP) uses a specific algorithm to extract parameters representing voice features from the voice signal. These feature parameters reflect the acoustic characteristics of the voice, such as pitch and timbre, providing key data for subsequent interaction with the large model. The extracted voice feature data is organized and prepared for transmission to the main processor. The main processor (CPU + GPU) is responsible for overall system control and task scheduling, ensuring smooth data transmission and collaborative operation between the various hardware modules. The GPU, with its powerful parallel computing capabilities, takes on the complex computation and analysis of voice feature data. The RAM is used to temporarily store various data generated during device operation. During the voice input phase, collected voice data is temporarily stored in the RAM pending processing. During data processing, intermediate calculation results and model parameters are also stored in the RAM, allowing the main processor to access and process them at any time. The storage module stores important data, including the smart glasses' operating system, basic models related to AI dialogue, and configuration files. The operating system provides an operating environment for the smart glasses' hardware and software, enabling the various hardware modules to work together. The basic models are the foundation for AI dialogue. Although they are integrated with a large cloud-based model during actual dialogue, the local basic models can perform some preliminary processing and analysis, reducing network reliance and improving response speed. The network communication module (Wi-Fi / 4G / 5G) uploads the locally processed voice feature data to the cloud server through the network communication module. The cloud server is equipped with a powerful large-scale model. After receiving the data, the large-scale model uses its powerful language understanding and generation capabilities, trained on massive amounts of data, to deeply analyze and understand the speech characteristics and generate appropriate responses. The response data is then transmitted back to the smart glasses via the network communication module. The positioning module receives signals from multiple satellites and uses the principle of triangulation to calculate the approximate location of the smart glasses. The positioning module continuously tracks satellite signals to ensure accurate location data in real time.Components from other modules also assist in positioning. The Wi-Fi module scans for nearby Wi-Fi hotspots and utilizes a database of known Wi-Fi hotspot locations for positioning, providing relatively accurate location information. The Bluetooth module interacts with nearby Bluetooth beacons to achieve precise positioning at close range. Furthermore, the inertial measurement unit (IMU), which includes an accelerometer and gyroscope, can sense the motion and directional changes of the smart glasses, assisting in positioning by inferring their trajectory. This allows for continued location updates, especially when GPS signals are lost.
[0012] The output and display module includes lenses, a speech synthesis chip, a digital-to-analog converter (DAC), and a speaker. The lenses serve as the display. The main processor transmits text response data received from the cloud to the display driver chip. The display driver chip converts this data into signals suitable for display on the display, driving the display to visually present the results of the AI conversation. Users can intuitively access the responses from the smart glasses by viewing the text on the display. The main processor also sends the text response data to the speech synthesis chip. The speech synthesis chip uses a specific algorithm to convert the text into a digital speech signal that simulates human speech intonation. The digital speech signal is then converted to an analog signal by a digital-to-analog converter (DAC), which drives the speaker. Finally, the speaker converts the analog signal into sound, transmitting the AI conversation response to the user in the form of voice, enabling natural and smooth voice interaction.
[0013] The power management module includes a battery and a power management chip. The battery provides power to the various hardware modules in the smart glasses. The power management chip manages the battery's charging and discharging processes, ensuring safe and efficient battery use. It precisely distributes voltage and current based on the operating requirements of each hardware module, ensuring that each module operates stably under appropriate power conditions. Furthermore, the power management chip monitors and manages the battery's charge level to optimize battery life.
[0014] In a specific embodiment 1, in a real-time translation conversation scenario, the large model combines the previous conversation content to understand the current sentence, avoiding translation errors caused by multiple meanings. For example, the large model can accurately determine whether "bank" means "bank" or "riverbank" in different contexts. In real-time translation, the model accurately identifies and processes input with accents or special language variants, improving translation accuracy, such as accurately translating English with an Indian accent. The intelligent agent intelligently determines and selects appropriate source and target languages based on the user's usage habits, current location, and historical translation history. For example, if the user is abroad and frequently communicates with locals, the intelligent agent can automatically set the source language to the local language and the target language to the user's native language. If it detects that the user frequently communicates with people from a specific country, it can also intelligently switch to the corresponding language pair to improve translation efficiency. The intelligent agent combines contextual information, such as the location (conference room, restaurant, tourist attraction, etc.), to optimize the translation content. In business meetings, the translation style is more formal and professional; at tourist attractions, the translation is more understandable and includes cultural background information. For example, when visiting historical sites, the agent not only translates the guide's explanation but also provides additional historical and cultural information to help users better understand. If a user repeatedly modifies the translation of a specific word or sentence, the agent will adjust the corresponding translation rules to ensure that subsequent translations better meet the user's needs.
[0015] In a specific embodiment 2, in AI conversation scenarios, the AI large model can understand vague, colloquial, or even incomplete user expressions. When conversing with smart glasses, users may express themselves casually. The large model, leveraging its powerful language comprehension capabilities, can infer precise intent from ambiguous expressions. For example, if a user says, "Help me find a place nearby to eat," the large model can interpret this as a search for nearby restaurants. The large model remembers the conversation context and understands referential relationships and implicit information within the conversation. In multiple rounds of conversation, it accurately understands the meaning of the current user's utterance based on the preceding context, enabling coherent and natural conversation. For example, if a user first asks, "What's the weather like tomorrow?" and then asks, "Is it suitable for hiking?" the large model, based on the preceding context, understands that "that" refers to tomorrow and accurately answers whether it is suitable for hiking. The intelligent agent, through long-term interaction with users, collects information about the user's interests, hobbies, occupation, and daily concerns, building a detailed user profile. Based on this profile, the agent provides answers that better suit the user's needs and interests. For example, if the user is a technology enthusiast, the agent will try to explain general questions from a technological perspective and proactively push relevant technology information. The intelligent agent analyzes the user's language style, including vocabulary and tone, and responds in a similar manner. If the user's language is humorous, the intelligent agent will also respond in a light-hearted and humorous manner, enhancing the conversation's approachability and interest, making it more natural and fluid. The intelligent agent integrates data from sensors such as the smart glasses' accelerometer and gyroscope to understand the user's behavioral state (e.g., walking, running, or standing still). During the conversation, the agent adjusts its response style and content based on the user's state. For example, when a user asks for route information while exercising, the intelligent agent provides clear and concise navigation instructions, allowing the user to quickly access information.
[0016] In a specific example 3, a user wears smart glasses and enters a library. After the smart glasses are powered on and initialized, the scene recognition module begins operating. The sensor module detects that the ambient light is low, the sound environment is quiet, and the user's head movement is relatively stable. At the same time, the user opens a book and begins reading. Based on this information, the scene recognition module determines that the user is in a learning scenario. The user encounters an English word they don't recognize and says, "Search for [word name]." The audio module captures the speech, and the NLP engine performs speech recognition and semantic understanding, identifying the user's intended search for the word. The NLP engine queries the local vocabulary to obtain the word's explanation, pronunciation, and example sentences. The pronunciation is played through the audio module, and detailed information is displayed on the display module. The user can further ask, "What are the synonyms of this word?" The NLP engine, based on semantic understanding, queries the vocabulary again and provides an answer. The user is confused about an example sentence in the word explanation and raises a question. The system records the user's feedback and optimizes the relevant explanation text to better serve the user in the future. Simultaneously, the user's learning data is uploaded to the server to update the personalized learning model.
[0017] Specific Example 4 involves a work scenario application. A user is working in an office, with their smart glasses turned on and connected to the office network. The sensor module detects that the user is in a relatively fixed position, with characteristic office sounds such as keyboard tapping and conversations. The user is also frequently checking the smart glasses and performing work-related operations. The scene recognition module determines that the user is in a work scenario. The smart glasses receive a new email notification, and the display module pops up a notification. The user issues a voice command "Read email." The audio module captures the voice. The NLP engine recognizes the command, calls the office software interface, retrieves the email content, and reads it aloud to the user via the audio module. After listening to the email, the user says, "Reply to email with the content [reply content]." The NLP engine converts the voice to text, generates an email reply, and sends it. In a meeting scenario, the user activates the voice conferencing feature. The microphone captures the user's voice, and the speakers play the voices of other participants. The NLP engine transcribes the meeting content in real time and displays it on the smart glasses' display module for easy viewing. During the voice conference, if the user finds the audio unclear, they adjust the volume using voice commands. The system records the user's operating habits and automatically adjusts the volume to an appropriate level in similar scenarios. At the same time, the meeting transcription content is proofread, and if errors are found, the transcription model of the NLP engine is optimized.
[0018] Beneficial effects: (1) Strong adaptability to multiple scenarios: Through the scene recognition module, it can accurately judge the user's learning, work, life and other scenarios, and provide targeted functional services according to the needs of different scenarios, greatly improving the practicality and user experience of smart glasses; (2) Efficient NLP interaction: The advanced NLP engine achieves accurate speech recognition, semantic understanding, and text generation, enabling users to interact smoothly with smart glasses through natural language without complex operations, thus lowering the user threshold; (3) Personalized learning assistance: In the learning scenario, based on the user's learning behavior and feedback, personalized learning plans are formulated for the user and precise learning assistance is provided, which helps to improve the user's learning efficiency and effectiveness; (4) Improve work efficiency: In work scenarios, integration with office software and voice interaction functions make information acquisition and processing more convenient, such as email processing and meeting participation, effectively improving user work efficiency; (5) Continuously optimize performance: By collecting and analyzing user feedback and usage data, we continuously optimize the NLP model and scene recognition model to continuously improve the performance and functions of smart glasses to better meet user needs.
Claims
1. A smart glasses system based on NLP and multi-scene fusion, characterized by: include: Hardware modules, including a display module, an audio module, a processing module, and a sensor module; Software modules, including an NLP engine, a scene recognition module, and a multi-scene application module; Among them, the sensor module is used to collect data, the scene recognition module determines the user's scene based on the data collected by the sensor module and the user's behavior pattern and environmental information, the NLP engine recognizes and understands the user's voice, and the multi-scene application module provides corresponding functional services based on the scene recognition results and user instructions.
2. The smart glasses system based on NLP and multi-scene fusion according to claim 1, characterized in that: The display module adopts a high-resolution, low-power micro display screen, the audio module includes a high-sensitivity microphone and a high-quality speaker, the processing module is a high-performance processor, and the sensor module includes an accelerometer, a gyroscope, a GPS module and an ambient light sensor.
3. The smart glasses system based on NLP and multi-scene fusion according to claim 1, characterized in that: The NLP engine is built on a deep learning framework and is capable of performing speech recognition, semantic understanding, and text generation operations.
4. The smart glasses system based on NLP and multi-scene fusion according to claim 1, characterized in that: The multi-scenario application module includes learning scenario applications, work scenario applications and life scenario applications. The learning scenario application provides language learning assistance functions, the work scenario application is integrated with office software and supports voice conferencing functions, and the life scenario application provides corresponding services in scenarios such as shopping and travel.
5. A smart glasses method based on NLP and multi-scene fusion, characterized in that: The following steps are involved: During the initialization phase, after the smart glasses are powered on, the hardware module performs a self-test, the software module loads the NLP engine, scene recognition module, and multi-scene application module, and connects to the network. During the scene recognition phase, the sensor module continuously collects data, and the scene recognition module analyzes and processes the data to determine the scene the user is in; During the interactive processing phase, when a user issues a voice command, the audio module collects the voice signal and transmits it to the NLP engine for voice recognition and semantic understanding. The multi-scenario application module then provides corresponding functional services based on the scene recognition results and user commands. During the feedback and update phase, the system optimizes and updates the NLP model and scene recognition model based on user operations and feedback, and regularly obtains the latest data from the server to update the local model.
6. The smart glasses method based on NLP and multi-scene fusion according to claim 5, characterized in that: During the scene recognition stage, the GPS data is used to determine the speed and direction of movement, while the ambient light, sound and other information as well as the user's operating behavior and voice commands are analyzed to comprehensively determine the user's current scene.
7. The smart glasses method based on NLP and multi-scene fusion according to claim 5, characterized in that: During the interactive processing phase, if the user is in a learning scenario, the NLP engine queries the vocabulary and performs translation based on the user's instructions, and provides feedback through the audio and display modules. If in a work scenario, the system calls the office software interface to obtain and process information, and supports functions such as voice conferencing and email processing.
8. The smart glasses method based on NLP and multi-scene fusion according to claim 5, characterized in that: During the feedback and update stage, the system records user feedback on translation results, scene recognition, etc., optimizes the NLP model and scene recognition model parameters, and uploads user data to the server for updating personalized models.