AI-driven multi-mode sensing intelligent glasses device and data processing method

The AI-driven multimodal perception smart glasses device integrates a high-definition camera, LiDAR, and microphone array. It combines visual SLAM and LiDAR point cloud fusion algorithms to construct a 3D map, parse user voice commands, and achieve cross-language translation. This solves the positioning and interaction problems of multimodal perception smart glasses in complex environments and improves data analysis and application effects.

CN121784973APending Publication Date: 2026-04-03BEIJING ZHISHANG ZONGHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing multimodal sensing smart glasses suffer from inaccurate positioning and unclear structure in complex environments. Their speech recognition and translation technologies perform poorly in noisy environments and lack effective cross-language support, resulting in poor data analysis and application performance.

Method used

The AI-driven multimodal perception smart glasses device integrates a high-definition camera, LiDAR, microphone array, and motion sensor. It combines dynamic noise reduction algorithms to output synchronized multimodal data, constructs a 3D map through visual SLAM+LiDAR point cloud fusion algorithms, uses a GPT-4 level model to parse voice commands and achieve real-time translation between Chinese, English, Japanese, and Korean, generates AR images using the Transformer-XL architecture, and updates the knowledge base through 5G/Wi-Fi 6 transmission and edge computing.

Benefits of technology

It achieves high-precision spatial understanding and content association, improves the accuracy and robustness of 3D map construction, supports natural language interaction and cross-language communication, lowers the user interaction threshold, and promotes global team collaboration and communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121784973A_ABST
    Figure CN121784973A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-driven multi-mode sensing intelligent glasses device and a data processing method, and relates to the technical field of multi-mode sensing intelligent glasses. The lens comprises a sunglasses lens, a full-color perspective display screen and at least one glasses leg, the full-color perspective display screen is superposed on the inner side of the sunglasses lens, and the glasses leg is connected with the glasses frame; the shell is integrated on one side of the glasses legs and comprises an image engine for generating an optical image and projecting the optical image to the full-color perspective display screen, a CPU (Central Processing Unit) for executing data processing and a built-in battery for providing power for the device; the functional assembly comprises a high-definition camera which is arranged in front of the glasses frame and is used for environment image acquisition, an LED lamp which is integrated on the glasses frame and is used for environment light supplement, a noise reduction microphone which is arranged on the inner side of the glasses leg and is used for voice input, and a touch panel which is arranged on the outer side of the glasses leg and is used for gesture control input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal sensing smart glasses technology, and in particular to AI-driven multimodal sensing smart glasses devices and data processing methods. Background Technology

[0002] Multimodal sensing smart glasses technology is a wearable device that integrates multiple sensors and technologies, aiming to provide users with a platform for comprehensively perceiving their surroundings and engaging in efficient interaction. Therefore, how to utilize advanced technologies to improve the intelligence and security of multimodal sensing smart glasses has become one of the most pressing issues to be addressed.

[0003] In the field of multimodal sensing smart glasses, existing technologies often suffer from time delays or inconsistencies in the data collected by different sensors, making it difficult to achieve accurate multimodal data analysis and application. Furthermore, traditional 3D modeling methods are prone to problems such as inaccurate positioning and structural ambiguity in complex environments, affecting the subsequent application results. At the same time, speech recognition and translation technologies often fail to accurately understand the user's intent, especially performing poorly in noisy environments, and lack effective cross-language support. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides an AI-driven multimodal perception smart glasses device and data processing method to solve the problems of inaccurate positioning and structural ambiguity that traditional 3D modeling methods are prone to in complex environments, affecting the subsequent application effect. At the same time, speech recognition and translation technologies often fail to accurately understand the user's intentions, especially in noisy environments, and lack effective cross-language support.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides an AI-driven multimodal sensing smart glasses device, comprising: A pair of glasses frames; At least one lens integrated into the eyeglass frame, the lens comprising: Sunglasses lens, a full-color transparent display screen superimposed on the inner side of the sunglasses lens, and at least one temple connected to the eyeglasses frame; An enclosure is integrated on one side of the temple, the enclosure including an image engine for generating optical images and projecting them onto the full-color transparent display screen, a CPU for performing data processing, and a built-in battery for providing power to the device; The functional components include a high-definition camera arranged in front of the eyeglasses frame for capturing environmental images, an LED light integrated into the eyeglasses frame for ambient lighting, a noise-canceling microphone located on the inside of the temple for voice input, a touchpad located on the outside of the temple for gesture control input, a speaker embedded inside the temple for audio output, an indicator light located in a prominent position on the eyeglasses frame for status display, and a Type-C interface located at the end of the temple for charging and data transmission. Interactive components include a power button located on the outside of the temple and a nose pad connected to the middle of the eyeglass frame; The communication module includes an antenna layer disposed near the lens, and one or more wireless transceivers for near-field communication.

[0007] Secondly, the present invention provides a data processing method for an AI-driven multimodal sensing smart glasses device, comprising: Environmental data is collected using cameras, LiDAR, microphone arrays, and motion sensors, and combined with dynamic noise reduction algorithms to output synchronized multimodal data; Based on synchronized multimodal data, a dynamic 3D map is constructed using a visual SLAM+LiDAR point cloud fusion algorithm, and AR content is associated through a cross-modal retrieval engine to generate a semantically annotated 3D environment model. It integrates a GPT-4 level model to parse user voice commands, generates operation intentions based on context, and uses the Transformer-XL architecture to achieve real-time translation between Chinese, English, Japanese and Korean, outputting interactive commands with semantic tags. By matching semantic commands with 3D models, AR images with a 40° ultra-wide field of view are generated through diffractive waveguide + DLP optomechanical technology. AR images are transmitted using 5G / Wi-Fi 6, and feedback is annotated. The knowledge base is dynamically updated at the edge nodes to output AR guidance content adapted to the current scene; The SOP engine generates operation reports, records the collaboration process, and forms a structured knowledge base for recording and execution in a closed loop.

[0008] As a preferred embodiment of the data processing method for the AI-driven multimodal perception smart glasses device of the present invention, the steps of acquiring environmental data using a camera, LiDAR, microphone array, and motion sensor, and outputting synchronized multimodal data by combining a dynamic noise reduction algorithm are as follows: The camera continuously captures images of the surrounding environment to obtain image data; The surrounding environment was scanned using a LiDAR device to obtain point cloud data at the same point in time. Audio data is obtained by capturing sound signals in the environment using a microphone array; Accelerometer motion sensors are used to monitor user movement and generate motion data; Add timestamps to each of the four types of data mentioned above; A timestamp-based synchronization algorithm is used to synchronize image data with timestamps. Point cloud data Audio data and sports data Perform synchronous processing to form a unified data stream; Finally, a fusion function is used to integrate the synchronized and denoised multimodal data to generate a synchronized multimodal dataset.

[0009] As a preferred embodiment of the data processing method for the AI-driven multimodal perception smart glasses device of the present invention, the steps of constructing a dynamic 3D map based on synchronized multimodal data using a visual SLAM+LiDAR point cloud fusion algorithm, and generating a semantically annotated 3D environment model by associating AR content through a cross-modal retrieval engine are as follows: Visual SLAM algorithm is used to synchronize image data The data is processed to extract keyframes and feature points, yielding preliminary spatial location information and camera pose.

[0010] Using LiDAR point cloud fusion algorithm to synchronize point cloud data The point cloud data is then processed and fused with the results of visual SLAM to obtain a three-dimensional spatial structure.

[0011] For the fused 3D spatial data, machine learning models are used to identify objects and their categories in the scene, and semantic labels are added to each identified object to generate a 3D environment model containing semantic information.

[0012] As a preferred embodiment of the data processing method for the AI-driven multimodal perception smart glasses device described in this invention, the method involves: integrating a GPT-4 level model to parse user voice commands, generating operational intentions based on context, and employing a Transformer-XL architecture to achieve real-time translation between Chinese, English, Japanese, and Korean, outputting interactive commands with semantic tags. The specific steps are as follows: Audio data is acquired from the microphone array and preprocessed to reduce the impact of background noise; The denoised audio signal is converted into text format, and the text content is analyzed using a GPT-4 level model to extract the user's intent and contextual information. Leveraging the contextual understanding capabilities of the GPT-4 level model, based on historical dialogue records With the current voice command Generate specific operational intentions ; The generated operational intent is translated in real time using the Transformer-XL architecture to obtain the translation result. Based on the translation results, appropriate semantic tags are added to each operation instruction through semantic analysis to form interactive instructions with semantic tags.

[0013] As a preferred embodiment of the data processing method for the AI-driven multimodal perception smart glasses device of the present invention, the specific steps of matching semantic instructions with a three-dimensional model and generating an AR image with a 40° ultra-wide field of view through diffractive waveguide + DLP optical engine technology are as follows: Interactive commands with semantic tags are used to parse user needs; Based on the content of the semantic instructions, locate the corresponding object or location in the previously constructed 3D environment model containing semantic information, and set the matching function; In the three-dimensional environment model Add semantic instructions Specified content ; The output is a visual feedback that integrates the original environment view with augmented reality content, directly presented to the user so that they can intuitively see and understand important information in the surrounding environment.

[0014] As a preferred embodiment of the data processing method for the AI-driven multimodal perception smart glasses device of the present invention, the steps of transmitting AR images based on 5G / Wi-Fi 6, annotating feedback, dynamically updating the knowledge base at edge nodes, and outputting AR guidance content adapted to the current scene are as follows: The final visual feedback is transmitted from the smart glasses device to a remote server or collaborating terminal using high-speed wireless communication technologies such as 5G / Wi-Fi 6. During transmission, the AR screen is annotated and fed back in real time based on the user's immediate actions and changes in the environment to enhance the interactive experience. Based on user interaction and on-site conditions, the knowledge base is dynamically updated on edge nodes close to the user using edge computing capabilities. The updated knowledge base KBKB is analyzed to identify elements relevant to the current scene, and new or updated AR guidance content is generated accordingly. ; Newly generated or updated AR guidance content The data is then transmitted back to the smart glasses device via a 5G / Wi-Fi 6 network.

[0015] As a preferred embodiment of the data processing method for the AI-driven multimodal perception smart glasses device of the present invention, the specific steps of generating operation reports based on the SOP engine, recording the collaboration process, and forming a structured knowledge base recording and execution closed loop are as follows: The SOP engine tracks and records all user actions in the augmented reality (AR) environment in real time. The operational behaviors include the user's voice commands, gestures, and interactions with virtual objects; According to the monitoring function The result is that the SOP engine is used to generate an operation report containing timestamps and specific operation details; While the user is performing the task, the system records audio and video information throughout the entire collaboration process; The recorded audio and video information is combined with the previously generated operation report to form an operation record containing multimedia information; The comprehensive operation record generated above will be uploaded to the structured knowledge base.

[0016] Thirdly, the data processing system of the AI-driven multimodal sensing smart glasses device includes, in which the CPU for performing data processing comprises: Data acquisition module, synchronous processing module, 3D modeling module, and voice parsing module; The data acquisition module is used to acquire image, point cloud, audio, and motion data; The synchronization processing module is used to add timestamps and synchronize multi-source data to generate a unified data stream; The 3D modeling module is used to construct a dynamic 3D map and generate a semantically annotated environment model; The voice parsing module is used to recognize user voice commands, generate operation intentions, and perform multilingual translation.

[0017] As a preferred embodiment of the data processing system for the AI-driven multimodal perception smart glasses device of the present invention, the CPU used for performing data processing further includes: AR matching module, transmission feedback module, and knowledge update module; The AR matching module is used to match semantic instructions with three-dimensional models to generate AR images; The transmission feedback module is used to transmit AR images via 5G / Wi-Fi 6 and update the knowledge base at the edge node; The knowledge update module is used to generate operation reports, record the collaboration process, and form a structured knowledge base record.

[0018] The beneficial effects of this invention are as follows: By constructing a dynamic 3D map based on synchronized multimodal data using a visual SLAM+LiDAR point cloud fusion algorithm, and associating AR content through a cross-modal retrieval engine, a semantically labeled 3D environment model is generated, achieving high-precision spatial understanding and content association. By combining visual SLAM and LiDAR point cloud fusion technologies, the accuracy and robustness of 3D map construction are significantly improved. Simultaneously, the semantic labeling function enables the system to better understand user intent and provide targeted guidance. The highly customized interactive experience greatly improves work efficiency, especially in scenarios requiring rapid decision-making and support. The integrated GPT-4 level model parses user voice commands, generates operational intent based on context, and uses the Transformer-XL architecture to achieve real-time translation between Chinese, English, Japanese, and Korean, outputting semantically labeled interactive commands. This enables natural language interaction and cross-language communication. This function not only lowers the interaction threshold between users and devices, making it easy for non-professional users to use, but also breaks down language barriers through multilingual support, promoting cross-border cooperation and exchange. This is particularly important for global team collaboration, significantly improving communication efficiency and workflow. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the AI-driven multimodal perception smart glasses device in Example 1.

[0021] Figure 2 This is a flowchart of the data processing method for the AI-driven multimodal sensing smart glasses device in Example 2.

[0022] Figure 3 This is a schematic diagram of the data processing system of the AI-driven multimodal perception smart glasses device in Example 2.

[0023] Figure 4 This is a schematic diagram of the multimodal sensing smart camera in Example 1. Detailed Implementation

[0024] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0025] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0026] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0027] Example 1, referring to Figure 1 and Figure 4 This is the first embodiment of the present invention, which provides an AI-driven multimodal perception smart glasses device, including: A pair of glasses frames; At least one lens integrated into the eyeglass frame, the lens comprising: Sunglasses lens, a full-color transparent display screen superimposed on the inside of the sunglasses lens, and at least one temple connected to the eyeglass frame; An enclosure is integrated on one side of the temple, and the enclosure includes an image engine for generating optical images and projecting them onto a full-color transparent display, a CPU for performing data processing, and a built-in battery for providing power to the device; Functional components include a high-definition camera positioned at the front of the eyeglasses frame for capturing environmental images, an LED light integrated into the eyeglasses frame for ambient lighting, a noise-canceling microphone located on the inside of the temples for voice input, a touchpad located on the outside of the temples for gesture control input, a speaker embedded inside the temples for audio output, an indicator light located in a prominent position on the eyeglasses frame for status display, and a Type-C interface located at the end of the temples for charging and data transfer. Interactive components include a power button located on the outside of the temples and a nose pad connected to the middle of the eyeglass frame; The communication module includes an antenna layer disposed near the lens and one or more wireless transceivers for near-field communication.

[0028] Furthermore, the full-color transparent display screen adopts binocular diffractive waveguide technology, combined with a DLP optical engine to achieve a 40° ultra-wide field of view display with a light transmittance of over 80%, ensuring the natural integration of AR content with the real environment; the CPU is a dedicated AI processing chip, integrating a neural network acceleration unit, supporting real-time processing of multimodal sensor data and executing 3D target recognition, voice recognition, and gesture recognition algorithms; the wireless transceiver includes a terahertz communication module, which achieves high-speed near-field data transmission through a photoconductive antenna, while being compatible with 5G / Wi-Fi 6 multimode communication protocols to adapt to different network environments; It should be noted that in this embodiment, the sunglasses lenses are made of variable transmittance material, which can automatically adjust the transmittance according to the ambient light intensity; the noise-canceling microphone adopts an array design, supporting accurate voice acquisition in an industrial noise environment of 90dB; the touchpad supports multi-touch and pressure sensing, and can recognize complex gesture commands; the built-in battery uses graphene composite material, and works with a power management chip to achieve intelligent power consumption adjustment, supporting more than 8 hours of continuous operation on a full charge; all electronic components are connected through flexible circuit boards, and a nano-coating is used to achieve IP67 waterproof and dustproof rating, ensuring the reliability and durability of the equipment in industrial environments.

[0029] In summary, by integrating multimodal input components such as a high-definition camera, noise-canceling microphone array, and touchpad, along with a full-color transparent display screen and speaker output, multi-channel natural interaction involving vision, voice, and touch is achieved, significantly improving the efficiency and ease of operation of human-machine interaction in industrial scenarios. The combination design of variable transmittance sunglasses lenses and intelligent LED supplementary lighting enables the device to automatically adapt to various working environments from strong light to low light. IP67 protection and industrial-grade noise reduction ensure stable operation under harsh working conditions. By distributing and integrating core components such as the battery and CPU into the temples and frame, and using lightweight materials (titanium alloy frame + graphene components), wearing comfort is achieved while ensuring functional integrity, making it suitable for long-term work.

[0030] Example 2, refer to Figure 2 and Figure 3 This is a second embodiment of the present invention, which provides a data processing method for an AI-driven multimodal sensing smart glasses device, including the following steps: S1. Uses cameras, LiDAR, microphone arrays and motion sensors to collect environmental data, and combines dynamic noise reduction algorithms to output synchronized multimodal data; Furthermore, a camera is used to continuously capture images of the surrounding environment to obtain image data; The surrounding environment was scanned using a LiDAR device to obtain point cloud data at the same point in time. Audio data is obtained by capturing sound signals in the environment using a microphone array; Accelerometer motion sensors are used to monitor user movement and generate motion data; Add timestamps to each of the four types of data mentioned above; A timestamp-based synchronization algorithm is used to synchronize image data with timestamps. Point cloud data Audio data and sports data Perform synchronous processing to form a unified data stream; Finally, a fusion function is used to integrate the synchronized and denoised multimodal data to generate a synchronized multimodal dataset, expressed as: ; in, It is the sampling time period. , , , These are the weighting coefficients corresponding to image, point cloud, audio, and motion data, respectively. It should be noted that, in this process, the application of dynamic noise reduction algorithms not only significantly reduced the impact of background noise on audio data, but also improved the accuracy of speech recognition. In addition, through precise timestamp synchronization technology, it was ensured that data from different sources could be accurately aligned in the time dimension, which is crucial for subsequent multimodal data analysis and fusion. The high-precision time synchronization mechanism laid the foundation for building a high-quality three-dimensional environment model.

[0031] S2. Based on synchronized multimodal data, a dynamic 3D map is constructed using a visual SLAM+LiDAR point cloud fusion algorithm, and AR content is associated through a cross-modal retrieval engine to generate a semantically annotated 3D environment model. Furthermore, a visual SLAM algorithm is used to synchronize image data. The data is processed to extract keyframes and feature points, yielding preliminary spatial location information and camera pose.

[0032] Using LiDAR point cloud fusion algorithm to synchronize point cloud data The point cloud data is then processed and fused with the results of visual SLAM to obtain a three-dimensional spatial structure.

[0033] For the fused 3D spatial data, machine learning models are used to identify objects and their categories in the scene, and semantic labels are added to each identified object to generate a 3D environment model containing semantic information. It should be noted that the combination of visual SLAM and LiDAR point cloud fusion algorithms fully utilizes the advantages of both technologies, improving both positioning accuracy and environmental perception. Visual SLAM provides high-resolution spatial information, while LiDAR provides accurate distance measurement. The two complement each other, making the generated 3D map more accurate and stable. In addition, the use of machine learning models for scene understanding and semantic annotation further enhances the system's ability to understand complex environments and improves the interactive experience.

[0034] S3 integrates a GPT-4 level model to parse user voice commands, generates operation intentions based on context, and uses the Transformer-XL architecture to achieve real-time translation between Chinese, English, Japanese and Korean, outputting interactive commands with semantic tags; Furthermore, audio data is acquired from the microphone array and preprocessed to reduce the impact of background noise; The denoised audio signal is converted into text format, and the text content is analyzed using a GPT-4 level model to extract the user's intent and contextual information. The expression is as follows: ; in, This is the text for voice commands. This is a speech recognition function; Leveraging the contextual understanding capabilities of the GPT-4 level model, based on historical dialogue records With the current voice command Generate specific operational intentions ; The generated operational intent is translated in real time using the Transformer-XL architecture to obtain the translation result. Based on the translation results, appropriate semantic tags are added to each operation instruction through semantic analysis to form interactive instructions with semantic tags; It should be noted that the powerful natural language processing capabilities of the GPT-4 level model enable it to accurately understand the user's intent and generate specific operation instructions based on the context. The introduction of the Transformer-XL architecture enables real-time translation between Chinese, English, Japanese, and Korean, greatly expanding the system's applicability and user base. Cross-language support not only improves the user experience but also plays an important role in a global collaborative environment, promoting efficient communication across different language backgrounds.

[0035] S4. Match semantic instructions with the 3D model and generate an AR image with a 40° ultra-wide field of view through diffractive waveguide + DLP optical engine technology; Furthermore, interactive instructions with semantic tags are used to analyze user needs; Based on the content of the semantic instructions, locate the corresponding object or location in the previously constructed 3D environment model containing semantic information, and set the matching function; In the three-dimensional environment model Add semantic instructions Specified content The expression is: ; in, The function is responsible for integrating new content into the 3D model; The output is a visual feedback that integrates the original environment view with augmented reality content, which is directly presented to the user, enabling them to intuitively see and understand important information in the surrounding environment; It should be noted that the 40° ultra-wide field of view AR image generated by diffractive waveguide + DLP optical engine technology not only provides a wide field of view, but also ensures the clarity of the image and the realism of the colors, greatly enhancing the user's immersion. The design of the matching function ensures that semantic commands can accurately correspond to objects or positions in the three-dimensional environment model, thereby achieving accurate display of augmented reality content. This step is of great significance for improving user experience, enhancing interactivity and ease of operation.

[0036] S5 transmits AR images based on 5G / Wi-Fi 6, adds annotations and feedback, dynamically updates the knowledge base at edge nodes, and outputs AR guidance content adapted to the current scene; Furthermore, by utilizing high-speed wireless communication technologies such as 5G / Wi-Fi 6, the final generated visual feedback can be sent from the smart glasses device to a remote server or collaborating terminal. During transmission, the AR screen is annotated and fed back in real time based on the user's immediate actions and changes in the environment to enhance the interactive experience. Based on user interaction and on-site conditions, the knowledge base is dynamically updated on edge nodes close to the user using edge computing capabilities. The updated knowledge base KBKB is analyzed to identify elements relevant to the current scene, and new or updated AR guidance content is generated accordingly. ; Newly generated or updated AR guidance content The data is then transmitted back to the smart glasses device via a 5G / Wi-Fi 6 network. It should be noted that by utilizing 5G / Wi-Fi 6 high-speed wireless communication technology, not only is the real-time performance and stability of AR image transmission guaranteed, but the image can also be annotated and fed back in real time during transmission, enhancing the user's interactive experience. The dynamic knowledge base update mechanism of the edge computing nodes enables the method to quickly adjust and optimize the AR guidance content based on the latest user interactions and on-site conditions, ensuring that the guidance provided is always up-to-date and relevant, significantly improving the efficiency and effectiveness of remote collaboration and support.

[0037] S6. Generate operation reports based on the SOP engine, record the collaboration process, and form a structured knowledge base recording and execution closed loop; Furthermore, the SOP engine tracks and records all user actions in the augmented reality (AR) environment in real time; User actions include voice commands, gestures, and interactions with virtual objects; According to the monitoring function The result is that the SOP engine is used to generate an operation report containing timestamps and specific operation details; While the user is performing the task, the system records audio and video information throughout the entire collaboration process; The recorded audio and video information is combined with the previously generated operation report to form an operation record containing multimedia information; Upload the comprehensive operation record generated above to the structured knowledge base; It should be noted that the operation reports generated by the SOP engine record every step of the user's operation and its results in detail, which is of great value for subsequent analysis and improvement. At the same time, the audio and video information recorded by the system provides intuitive multimedia materials for the operation process, which is convenient for review and learning. Integrating information into a structured knowledge base not only helps the accumulation and inheritance of knowledge, but also continuously improves the performance of methods and service quality through continuous learning and optimization. The closed-loop knowledge management and execution process design ensures that the system can adapt to constantly changing needs and technological advancements.

[0038] This embodiment also provides a data processing system for an AI-driven multimodal perception smart glasses device, including: The system includes a data acquisition module, a synchronization processing module, a 3D modeling module, a voice parsing module, an AR matching module, a transmission feedback module, and a knowledge update module. The data acquisition module is used to acquire image, point cloud, audio, and motion data; The synchronization processing module is used to add timestamps and synchronize multi-source data to generate a unified data stream; The 3D modeling module is used to build dynamic 3D maps and generate semantically annotated environment models; The voice parsing module is used to recognize user voice commands, generate operation intentions, and perform multilingual translations. The AR matching module is used to match semantic commands with 3D models to generate AR images; The transmission feedback module is used to transmit AR images via 5G / Wi-Fi 6 and update the knowledge base at the edge nodes; The knowledge update module is used to generate operation reports, record collaboration processes, and form a structured knowledge base record.

[0039] In summary, this invention constructs a dynamic 3D map based on synchronized multimodal data using a visual SLAM+LiDAR point cloud fusion algorithm. It then associates AR content through a cross-modal retrieval engine to generate a semantically annotated 3D environment model, achieving high-precision spatial understanding and content association. By combining visual SLAM and LiDAR point cloud fusion technologies, the accuracy and robustness of 3D map construction are significantly improved. Simultaneously, the semantic annotation function enables the system to better understand user intent and provide targeted guidance. The highly customized interactive experience greatly improves work efficiency, especially in scenarios requiring rapid decision-making and support. It integrates a GPT-4 level model to parse user voice commands, generates operational intent based on context, and uses the Transformer-XL architecture to achieve real-time translation between Chinese, English, Japanese, and Korean, outputting semantically tagged interactive commands. This enables natural language interaction and cross-language communication. This function not only lowers the interaction threshold between users and devices, making it easy for non-professional users to use, but also breaks down language barriers through multilingual support, promoting cross-border cooperation and exchange. This is particularly important for global team collaboration, significantly improving communication efficiency and workflow.

[0040] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An AI-driven multimodal sensing smart glasses device, characterized in that: include: A pair of glasses frames; At least one lens integrated into the eyeglass frame, the lens comprising: Sunglasses lens, a full-color transparent display screen superimposed on the inner side of the sunglasses lens, and at least one temple connected to the eyeglasses frame; An enclosure is integrated on one side of the temple, the enclosure including an image engine for generating optical images and projecting them onto the full-color transparent display screen, a CPU for performing data processing, and a built-in battery for providing power to the device; The functional components include a high-definition camera arranged in front of the eyeglasses frame for capturing environmental images, an LED light integrated into the eyeglasses frame for ambient lighting, a noise-canceling microphone located on the inside of the temple for voice input, a touchpad located on the outside of the temple for gesture control input, a speaker embedded inside the temple for audio output, an indicator light located in a prominent position on the eyeglasses frame for status display, and a Type-C interface located at the end of the temple for charging and data transmission. Interactive components include a power button located on the outside of the temple and a nose pad connected to the middle of the eyeglass frame; The communication module includes an antenna layer disposed near the lens, and one or more wireless transceivers for near-field communication.

2. A data processing method for an AI-driven multimodal sensing smart glasses device, based on the AI-driven multimodal sensing smart glasses device of claim 1, characterized in that: The processing method of the CPU for performing data processing includes: Environmental data is collected using cameras, LiDAR, microphone arrays, and motion sensors, and combined with dynamic noise reduction algorithms to output synchronized multimodal data; Based on synchronized multimodal data, a dynamic 3D map is constructed using a visual SLAM+LiDAR point cloud fusion algorithm, and AR content is associated through a cross-modal retrieval engine to generate a semantically annotated 3D environment model. It integrates a GPT-4 level model to parse user voice commands, generates operation intentions based on context, and uses the Transformer-XL architecture to achieve real-time translation between Chinese, English, Japanese and Korean, outputting interactive commands with semantic tags. By matching semantic commands with 3D models, AR images with a 40° ultra-wide field of view are generated through diffractive waveguide + DLP optomechanical technology. AR images are transmitted using 5G / Wi-Fi 6, and feedback is annotated. The knowledge base is dynamically updated at the edge nodes to output AR guidance content adapted to the current scene; The SOP engine generates operation reports, records the collaboration process, and forms a structured knowledge base for recording and execution in a closed loop.

3. The data processing method for the AI-driven multimodal sensing smart glasses device as described in claim 2, characterized in that: The process involves using cameras, LiDAR, microphone arrays, and motion sensors to collect environmental data, combined with dynamic noise reduction algorithms, to output synchronized multimodal data. The specific steps are as follows: The camera continuously captures images of the surrounding environment to obtain image data; The surrounding environment was scanned using a LiDAR device to obtain point cloud data at the same point in time. Audio data is obtained by capturing sound signals in the environment using a microphone array; Accelerometer motion sensors are used to monitor user movement and generate motion data; Add timestamps to each of the four types of data mentioned above; A timestamp-based synchronization algorithm is used to synchronize image data with timestamps. Point cloud data Audio data and sports data Perform synchronous processing to form a unified data stream; Finally, a fusion function is used to integrate the synchronized and denoised multimodal data to generate a synchronized multimodal dataset.

4. The data processing method for the AI-driven multimodal sensing smart glasses device as described in claim 3, characterized in that: The process involves constructing a dynamic 3D map based on synchronized multimodal data using a visual SLAM+LiDAR point cloud fusion algorithm, and generating a semantically annotated 3D environment model by associating AR content with a cross-modal retrieval engine. The specific steps are as follows: Visual SLAM algorithm is used to synchronize image data The process involves extracting keyframes and feature points to obtain preliminary spatial location information and camera pose. Using LiDAR point cloud fusion algorithm to synchronize point cloud data The point cloud data is processed and fused with the results of visual SLAM to obtain a three-dimensional spatial structure. For the fused 3D spatial data, machine learning models are used to identify objects and their categories in the scene, and semantic labels are added to each identified object to generate a 3D environment model containing semantic information.

5. The data processing method for the AI-driven multimodal sensing smart glasses device as described in claim 4, characterized in that: The integrated GPT-4 level model parses user voice commands, generates operation intentions based on context, and uses the Transformer-XL architecture to achieve real-time translation between Chinese, English, Japanese, and Korean, outputting interactive commands with semantic tags. The specific steps are as follows: Audio data is acquired from the microphone array and preprocessed to reduce the impact of background noise; The denoised audio signal is converted into text format, and the text content is analyzed using a GPT-4 level model to extract the user's intent and contextual information. Leveraging the contextual understanding capabilities of the GPT-4 level model, based on historical dialogue records With the current voice command Generate specific operational intentions ; The generated operational intent is translated in real time using the Transformer-XL architecture to obtain the translation result. Based on the translation results, appropriate semantic tags are added to each operation instruction through semantic analysis to form interactive instructions with semantic tags.

6. The data processing method for the AI-driven multimodal sensing smart glasses device as described in claim 5, characterized in that: The specific steps for matching semantic commands with a 3D model and generating an AR image with a 40° ultra-wide field of view using diffractive waveguide + DLP optomechanical technology are as follows: Interactive commands with semantic tags are used to parse user needs; Based on the content of the semantic instructions, locate the corresponding object or location in the previously constructed 3D environment model containing semantic information, and set the matching function; In the three-dimensional environment model Add semantic instructions Specified content ; The output is a visual feedback that integrates the original environment view with augmented reality content, directly presented to the user so that they can intuitively see and understand important information in the surrounding environment.

7. The data processing method for the AI-driven multimodal perception smart glasses device as described in claim 6, characterized in that: The process of transmitting AR images based on 5G / Wi-Fi 6, annotating feedback, dynamically updating the knowledge base at edge nodes, and outputting AR guidance content adapted to the current scene involves the following steps: The final visual feedback generated is sent from the smart glasses device to a remote server or collaborating device using high-speed wireless communication technology 5G / Wi-Fi 6. During transmission, the AR screen is annotated and fed back in real time based on the user's immediate actions and changes in the environment to enhance the interactive experience. Based on user interaction and on-site conditions, the knowledge base is dynamically updated on edge nodes close to the user using edge computing capabilities. The updated knowledge base KBKB is analyzed to identify elements relevant to the current scene, and new or updated AR guidance content is generated accordingly. ; Newly generated or updated AR guidance content The data is then transmitted back to the smart glasses device via a 5G / Wi-Fi 6 network.

8. The data processing method for the AI-driven multimodal sensing smart glasses device as described in claim 7, characterized in that: The steps for generating operation reports based on the SOP engine, recording the collaboration process, and forming a structured knowledge base record and execution loop are as follows: The SOP engine tracks and records all user actions in the augmented reality (AR) environment in real time. The operational behaviors include the user's voice commands, gestures, and interactions with virtual objects; According to the monitoring function The result is that the SOP engine is used to generate an operation report containing timestamps and specific operation details; While the user is performing the task, the system records audio and video information throughout the entire collaboration process; The recorded audio and video information is combined with the previously generated operation report to form an operation record containing multimedia information; The comprehensive operation record generated above will be uploaded to the structured knowledge base.

9. A data processing system for an AI-driven multimodal sensing smart glasses device, based on the AI-driven multimodal sensing smart glasses device and data processing method according to any one of claims 1 to 8, characterized in that: The CPU used for performing data processing includes: Data acquisition module, synchronous processing module, 3D modeling module, and voice parsing module; The data acquisition module is used to acquire image, point cloud, audio, and motion data; The synchronization processing module is used to add timestamps and synchronize multi-source data to generate a unified data stream; The 3D modeling module is used to construct a dynamic 3D map and generate a semantically annotated environment model; The voice parsing module is used to recognize user voice commands, generate operation intentions, and perform multilingual translation.

10. The AI-driven multimodal sensing smart glasses device and data processing method as described in claim 9, characterized in that: The CPU used for performing data processing also includes: AR matching module, transmission feedback module, and knowledge update module; The AR matching module is used to match semantic instructions with three-dimensional models to generate AR images; The transmission feedback module is used to transmit AR images via 5G / Wi-Fi 6 and update the knowledge base at the edge node; The knowledge update module is used to generate operation reports, record the collaboration process, and form a structured knowledge base record.