Multi-modal wearable device real-time interaction system and method based on light-wing visual field large model

By building a multimodal wearable device system through the Light Wing Vision large model, the problem of high hardware configuration of traditional devices is solved, and an efficient and convenient multimodal interactive experience is achieved, which is suitable for multi-scenario applications.

CN120670902APending Publication Date: 2025-09-19TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510737723.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional smart wearable devices use a large visual language model, which leads to high hardware configuration requirements, limiting the portability and interactive experience of the device, and making it difficult to meet users' needs for efficient interaction.

Method used

Using the Light Wing Vision large model and combining it with multimodal processing technology, we build a multimodal perception module, parallel processing module, multi-scenario application module and task management module, including target detection, speech recognition, language translation and music recognition functions, and reduce hardware performance requirements through lightweight design.

Benefits of technology

It significantly improves the device's performance in tasks such as image recognition, translation, and music recognition, provides an efficient and convenient multimodal interactive experience, adapts to diverse application scenarios, and achieves a balance between portability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670902A_ABST
    Figure CN120670902A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal wearable device real-time interaction system and method based on a light-wing visual field large model, and the system comprises a multi-modal sensing module which is used for obtaining sensing data through a sensing part, and a data collection module collects the obtained sensing data; the parallel processing module is used for performing feature extraction on the data acquired by the acquisition module and performing task allocation processing on the data based on a preset function model to obtain a corresponding processing result; the multi-scene application module is used for applying the processing result to the corresponding multi-application scene application module; and the task management module is used for cooperative operation of each functional module in the multi-scene application. The system is composed of a plurality of core technology modules including a target detection module, a voice recognition module and a cross-modal path aggregation network. According to the method, the image and language information is fused by using the context sensing technology through the cross-modal path aggregation network, so that the precision and robustness of the multi-modal task are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a multimodal wearable device real-time interaction system and method based on a light wing vision large model. Background Art

[0002] With the rapid development of artificial intelligence and multimodal technologies, smart wearable devices are gradually expanding beyond traditional functions to integrate more advanced sensing and interactive capabilities. Their influence is significantly increasing, creating broad application prospects for consumers and businesses. In this wave of technological innovation, smart wearable devices not only carry basic communication and entertainment functions, but also demonstrate significant value in areas such as information acquisition, real-time translation, and environmental recognition. However, traditional smart wearable devices have limited multimodal processing capabilities, especially in resource-constrained personal devices, where their performance often fails to meet users' demands for efficient interaction.

[0003] Mainstream smart wearable devices currently on the market typically utilize large visual language models (VLMs), such as GPT-4V, Qwen-VL, and LlaVA. While these models can demonstrate impressive performance in multimodal processing tasks, their tens of billions of parameters place extremely high demands on hardware configuration, making them prohibitive for average users. The resource consumption of large models during on-device inference limits the device's portability and interactive experience, significantly hindering its widespread adoption and application. To address this issue, lightweight visual language models have become a key area of ​​technological exploration in recent years. The Light Wing Vision large model series was developed to address this need. With its miniaturization, low resource consumption, and multimodal processing capabilities, it has brought disruptive breakthroughs to the field of smart wearable devices.

[0004] A multimodal real-time interaction system for smart wearable devices has been built based on the LightWing Vision large model. This system leverages the lightweight advantages of the LightWing Vision model and combines it with multimodal processing technology to significantly improve the device's performance in tasks such as image recognition, translation, and music recognition. Summary of the Invention

[0005] In view of the above problems existing in the prior art, the purpose of the present invention is to provide a real-time interactive system for multimodal wearable devices based on the light wing vision large model.

[0006] To solve the above problems, the present invention adopts the following technical solutions: a multi-modal wearable device real-time interaction system based on the light wing vision large model, including

[0007] The multimodal perception module is used to obtain perception data through the perception components, and the data acquisition module collects the acquired perception data;

[0008] A parallel processing module is used to extract features from the perception data collected by the acquisition module and perform task allocation processing on the collected perception data based on a preset functional model to obtain corresponding processing results. The parallel processing module includes a target detection module for real-time target detection and semantic understanding of the environment within the user's field of view. The target detection module detects multiple objects simultaneously and handles overlapping, occluded, and objects of different scales in complex scenes. A visual language model module is used to recognize and communicate the acquired perception data. A semantic segmentation module is used to assign a semantic category to each pixel in the image.

[0009] The parallel processing module includes a visual language model module, which uses a light-wing vision large model to recognize and communicate with the acquired perception data. The light-wing vision large model includes a light-wing vision large model series of four different sizes of models of the module, including light-wing vision-s110M, light-wing vision-n240M, light-wing vision-m290M and light-wing vision-l460M, which are used for image recognition and dialogue. Among them, s110M, n240M, m290M, and l460M are used to represent the scale of the model, respectively. s110M represents a small-scale model with a parameter amount of 110M, n240M represents a normal-scale model with a parameter amount of 240M, m290M represents a medium-scale model with a parameter amount of 290M, and l460M represents a large model with a parameter amount of 460M.

[0010] The multi-scenario application module applies the processing results to the corresponding multi-application scenario application module for real-time interaction;

[0011] The task management module is used for the coordinated operation of various functional modules in multi-scenario applications.

[0012] In some embodiments, the perception components are specifically cameras, microphones, and speakers, and the data acquisition modules respectively collect the perception data obtained by the perception components.

[0013] In some embodiments, the parallel processing module includes a speech recognition module for recognizing the acquired speech perception data, converting the speech perception data into understandable text or instructions, and parsing semantic information after the speech perception data is converted;

[0014] A language translation module is used to translate the semantic information data converted by the speech recognition module;

[0015] The music recognition module is used to obtain the perceived music data and make judgments based on the semantic information converted by the speech recognition module.

[0016] In some embodiments, the language translation module converts text from a language processed by the visual language model module into another language, preserving the original meaning.

[0017] In some embodiments, the music recognition module provides instant feedback of music recognition results based on the acquired semantic information data after the user triggers the music recognition module.

[0018] Another object of the present application is to provide a method for realizing real-time interaction of multimodal wearable devices based on the light wing vision large model, comprising the steps of:

[0019] Step 1: Acquire perception data through the perception component;

[0020] Step 2: The acquisition module collects the acquired perception data and classifies it based on the data features;

[0021] Step 3: Process the classified data based on the Light Wing Vision Large Model, including image recognition and dialogue;

[0022] Step 4: The task management module uses the processed data in multi-scenario applications, and each function runs in coordination.

[0023] An electronic device, comprising:

[0024] Memory, used to store computer programs;

[0025] A processor is used to execute the computer program to implement a real-time interaction method for a multimodal wearable device based on the light wing vision large model.

[0026] A computer-readable storage medium is used to store a computer program, which, when executed by a processor, implements a real-time interaction method for a multimodal wearable device based on a light-wing vision large model.

[0027] Compared with the prior art, the beneficial technical effects of the present invention are:

[0028] 1. The system described in this application is composed of multiple core technology modules, including an object detection module, a speech recognition module, and a cross-modal path aggregation network. The object detection module extracts multi-scale features from images, providing a reliable foundation for object detection and semantic segmentation tasks. The speech translation module extracts high-dimensional semantic embeddings from speech and text and performs cross-modal alignment with visual features. The cross-modal path aggregation network uses context-aware technology to fuse image and language information, significantly improving the accuracy and robustness of multimodal tasks.

[0029] 2. This application significantly reduces the device's hardware performance requirements through a lightweight model architecture, allowing ordinary consumer-grade hardware to also carry multimodal interaction functions. Multimodal fusion capabilities combine image, language, and audio processing to provide a comprehensive smart wearable interaction experience. From language translation to scene perception, the system can adapt to a variety of application scenarios, including travel, learning, and entertainment. With the optimized design of the Light Wing Vision large model, the system achieves an ideal balance between portability and performance, providing users with an efficient and intuitive way of intelligent interaction, while promoting the in-depth application and popularization of smart wearable devices in multiple fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 The system described in the embodiment of this application;

[0031] Figure 2 This is a schematic diagram of the basic architecture of the Light Wing Vision large model described in this application;

[0032] Figure 3 Parameter diagram of the four large light wing vision models shown in this application: light wing vision-s, light wing vision-n, light wing vision-m and light wing vision-l;

[0033] Figure 4 This is a schematic diagram of the operation of this application. DETAILED DESCRIPTION

[0034] The technical solution of the present invention will be further described in detail below with reference to the embodiments and drawings.

[0035] Example

[0036] To fully illustrate the breadth of this application without limiting its scope, the following demonstrates how the multimodal smart wearable device real-time interaction system based on the Light Wing Vision large model can be expanded to other embodiments also protected by this invention without requiring additional creative effort. This system is designed specifically for portable and efficient user interaction needs, highlighting its adaptability and convenience in multiple scenarios.

[0037] like Figure 1 As shown, the present application provides a multimodal intelligent wearable device real-time interaction system based on the light wing vision large model, including a multimodal perception module for acquiring perception data through a perception component, and a data acquisition module for acquiring the acquired perception data;

[0038] A parallel processing module is used to extract features from the perception data collected by the acquisition module and perform task allocation processing on the perception data based on a preset functional model to obtain corresponding processing results;

[0039] A multi-scenario application module applies the processing result to a corresponding multi-application scenario application module;

[0040] The task management module is used for the coordinated operation of various functional modules in multi-scenario applications.

[0041] In this embodiment, the smart wearable device provided by this application integrates multimodal perception modules, such as a high-definition camera, an omnidirectional microphone, and an open speaker. The functional modules together constitute the core perception and interaction basis of the device. The camera serves as the main visual information collector, continuously capturing dynamic images of the user's surroundings, providing real-time and clear visual data. The microphone captures user voice commands and ambient audio, providing input support for voice interaction and music recognition functions. The speaker feeds back information in a natural way, including translation results, music recognition details, and interaction prompts, ensuring a smooth and seamless user experience.

[0042] like Figure 2-4 As shown, the user uses a microphone to input voice, and the audio signal is transmitted to the speech recognition module in real time. The system uses the FunASR model for high-precision speech-to-text conversion, ultimately outputting the user's instructions in text format. The converted text data is then classified based on its features to determine whether the task is a language task or a visual language task.

[0043] For vision-language tasks such as object detection and semantic segmentation, camera data is collected as visual input, and a placeholder replacement mechanism is used to process multimodal input. For the language translation and music recognition modules, a local blank image is selected as visual input. When the system detects a placeholder consisting of 196 consecutive "@" characters (corresponding to the 196 image patches in the CLIP-ViT model), a cross-modal embedding replacement process is initiated. Specifically, the visual encoder first converts the input image into a 196×768 feature matrix, which is then projected into the same embedding space as the text using a learnable projection matrix. This projection process includes LayerNorm normalization and GeLU activation to ensure that the distribution of visual features is aligned with the text embedding. After the embedding layer completes the replacement, the model adds discernible positional identifiers to tokens from different modalities. Visual tokens are assigned positional codes from 0 to 195, while text tokens are numbered starting at 196. This design enables the subsequent RoPE positional encoding to correctly handle relative positional relationships across modalities. In the multi-head attention calculation of the Transformer layer, the model generates three types of attention masks: intra-visual mask (controls the interaction between image patches), intra-text mask (maintains language logic), and cross-modal mask (adjusts the intensity of image-text interaction), and then calls the linear layer to output the results.

[0044] During training, a modality-specific gradient isolation strategy is employed. The gradients of the visual projector are backpropagated only to the image encoding branch, while the gradients of the text embedding layer influence both the language and visual experts. This design ensures the coordinated optimization of cross-modal representations while preventing visual noise from interfering with language modeling. Furthermore, to address local runtime speed issues, four models of different sizes are specially trained for users to choose from: Lightwing Vision-s110M, Lightwing Vision-n240M, Lightwing Vision-m290M, and Lightwing Vision-l420M. LightWing Vision-s110M, LightWing Vision-n240M, LightWing Vision-m290M, and LightWing Vision-l460M are used for image recognition and communication. s110M, n240M, m290M, and l460M represent model size, respectively: s110M represents a small-scale model with 110M parameters, n240M represents a normal-scale model with 240M parameters, m290M represents a medium-scale model with 290M parameters, and l460M represents a large-scale model with 460M parameters. LightWing Vision-s110M, LightWing Vision-n240M, LightWing Vision-m290M, and LightWing Vision-l460M achieve lightweight model size through various model compression techniques, such as pruning and quantization.

[0045] The Light Wing Vision large model uses a jointly trained vision-language encoder to directly extract high-dimensional semantic embeddings of images and text. Furthermore, the Light Wing Vision model has already learned rich cross-modal association knowledge during pre-training, providing a semantic foundation for subsequent path aggregation. The design of context-aware prompt templates is particularly critical in the path guidance of prompt word engineering. By inserting structured prompt words into the input ([number of input images] + [context-aware prompt words, dynamically generating guidance instructions related to image content] + [text description or task objective]), the model is guided to focus on key connections between image and text. The prompt word template, combined with the text description, can be designed to "give the direction of [object name] in the field of view and generate cross-modal association results."

[0046] Furthermore, the device's built-in Light Wing Vision large model processes images captured by the camera and voice data input from the microphone in parallel. For image data, the object detection module quickly identifies key objects in the environment, such as signs, text, and specific items, and generates bounding boxes with semantic labels, helping users instantly understand the surrounding environment. The semantic segmentation module deeply analyzes the image, classifying each pixel into predefined categories such as roads, buildings, and trees, providing users with a more detailed understanding of the surrounding scene.

[0047] In language translation scenarios, the wearable device transmits foreign language content picked up by the microphone in real time to the Light Wing Vision model's language processing module, which uses its semantic understanding and generation capabilities to generate high-quality translation results. These results are played back in real time through open speakers or displayed simultaneously on the user's phone or other supporting device. Voice and text translation functions can be seamlessly switched without additional user interaction, ensuring convenient and efficient interaction.

[0048] In a music recognition scenario, users activate the music recognition function of the system described herein through voice commands. The device instantly analyzes the audio signal captured by the microphone, extracts its features, and then inputs them into the Light Wing Vision large model for comparison. The recognition results, including song title, artist, and related album information, are fed back to the user through speakers or a screen, satisfying their daily music needs.

[0049] It should be noted that this lightweight design is based on the optimized design of the Light Wing Vision large model. Multiple strategies are used to reduce model size and computational complexity while maintaining high accuracy. The lightweight design of the model structure includes parameter compression and low-rank decomposition techniques, which decompose the original high-dimensional parameter matrix into multiple low-rank matrices, thereby reducing the number of parameters by approximately 40%-60%. A specific solution is proposed in the paper "LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation." The core innovation lies in freezing the random projection matrix A while simultaneously performing task-specific sparsification on matrix B. This design not only reduces the trainable parameters to 5% of traditional LoRI, but also, through the mathematical principle of orthogonality, ensures that adapters for different tasks are virtually non-interfering when merged. When the low rank r=d / 4r=d / 4, the reduction ratio is 1-2×(d / 4)d=1-0.5=50%. Furthermore, a parameter sharing mechanism is implemented in the multimodal attention module to reduce redundant computation by sharing cross-modal parameters. In addition, knowledge distillation and dynamic model pruning were also adopted. The former enables lightweight models to maintain an accuracy of over 95% while significantly reducing the number of parameters. The latter dynamically trims model layers according to task requirements, saving 30%-50% of computing resources. This is achieved by adopting the reverse path (Dense-to-Sparse) modeled after "LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation".

[0050] In terms of hardware adaptation and quantization technology, mixed-precision quantization is used to reduce the model volume by 70% (for precision-sensitive layers (activation functions, key weight layers): high-precision quantization (FP16 or INT8) is used to retain more details. For precision-sensitive layers (activation functions, key weight layers): high-precision quantization (FP16 or INT8) is used to retain more details), and edge-side hardware acceleration is used to increase the inference speed to 3 times (through mixed-precision quantization (INT8 / FP16) and hardware simulator optimization, the inference latency on edge devices is reduced by 1.4-1.95 times, and combined with NPU, it can be further increased to 3 times).

[0051] In some embodiments, parallel processing of multimodal data and edge computing optimization are employed to achieve real-time interactions, reducing processing latency to 200-300 milliseconds. On-device model inference and real-time optimization strategies are used to control response times to under 400 milliseconds. Using AGX orin for inference, this approach delivers 275 TOPS (INT8) computing power, eight times that of the previous-generation Xavier. It can simultaneously process multiple high-resolution video streams (such as 8K video) and complex sensor data, splitting multimodal data into subtasks (such as image → CNN feature extraction, speech → MFCC extraction) and assigning them to different GPU SMs for parallel computation. Compared to existing technologies, this approach not only achieves significant improvements in lightweighting and real-time performance, but also ensures high accuracy and low latency.

[0052] The data processing unit, the system's core, aggregates and integrates massive amounts of data from sensor modules like cameras and microphones, performing comprehensive analysis based on the user's current task requirements and contextual information. This unit ensures rapid system switching and intelligent response across different functional scenarios. For example, when a user requests a translation, the system can pause music recognition in real time and automatically resume the original task after the translation is complete, demonstrating advanced task management capabilities and adaptability.

[0053] The control unit coordinates the device's specific operations based on instructions generated by the data processing unit. This includes, but is not limited to, dynamically adjusting speaker volume, switching task modes, and optimizing visual capture range. The device performs tasks accordingly, ensuring smooth and efficient interaction with the user. For example, in a translation scenario, users can use voice to inquire about translation details or context, and the system can instantly adjust the language model's translation logic to provide more comprehensive support.

[0054] Ultimately, the smart wearable device, leveraging its powerful LightWing Vision large-scale model and modular design, provides users with highly convenient intelligent support in a variety of daily scenarios. Whether traveling, working, or enjoying daily entertainment, the device demonstrates exceptional multimodal perception and real-time interaction capabilities, ensuring an intelligent and efficient user experience.

[0055] In summary, the present application provides a real-time interactive system for multimodal wearable devices based on the light-wing vision large model, including a multimodal perception module for acquiring perception data through perception components, and a data acquisition module for acquiring the acquired perception data; a parallel processing module for extracting features from the data collected by the acquisition module, and performing task assignment processing on the data based on a preset functional model to obtain corresponding processing results; a multi-scenario application module for applying the processing results to corresponding multi-application scenario application modules for real-time interaction; and a task management module for the coordinated operation of various functional modules in multi-scenario applications.

[0056] Finally, it is necessary to point out here that the above is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with the present invention within the technical scope disclosed by the present invention should be covered within the scope of protection of the present invention.

Claims

1. A multi-modal wearable device real-time interaction system based on the light wing vision large model, characterized by: include The multimodal perception module is used to obtain perception data through the perception components, and the data acquisition module collects the acquired perception data; A parallel processing module is used to extract features from the perception data collected by the acquisition module and perform task allocation processing on the collected perception data based on a preset functional model to obtain corresponding processing results. The parallel processing module includes a target detection module for real-time target detection and semantic understanding of the environment within the user's field of view. The target detection module detects multiple objects simultaneously and handles overlapping, occluded, and objects of different scales in complex scenes. A visual language model module is used to recognize and communicate with the acquired perception data. Semantic segmentation module, which assigns a semantic category to each pixel in the image; The parallel processing module includes a visual language model module, which uses a light-wing vision large model to recognize and communicate with the acquired perception data. The light-wing vision large model includes a light-wing vision large model series of four different sizes of models of the module, including light-wing vision-s110M, light-wing vision-n240M, light-wing vision-m290M and light-wing vision-l460M, which are used for image recognition and dialogue. Among them, s110M, n240M, m290M, and l460M are used to represent the scale of the model, respectively. s110M represents a small-scale model with a parameter amount of 110M, n240M represents a normal-scale model with a parameter amount of 240M, m290M represents a medium-scale model with a parameter amount of 290M, and l460M represents a large model with a parameter amount of 460M. The multi-scenario application module applies the processing results to the corresponding multi-application scenario application module for real-time interaction; The task management module is used for the coordinated operation of various functional modules in multi-scenario applications.

2. A multi-modal wearable device real-time interaction system based on the light wing vision large model according to claim 1, characterized in that: The perception components are specifically cameras, microphones, and speakers, and the data acquisition modules respectively collect the perception data obtained by the perception components.

3. The multimodal wearable device real-time interaction system based on the light wing vision large model according to claim 1 is characterized in that: The parallel processing module includes a speech recognition module for recognizing acquired speech perception data, converting the speech perception data into understandable text or instructions, and parsing semantic information after the speech perception data is converted; A language translation module is used to translate the semantic information data converted by the speech recognition module; The music recognition module is used to obtain the perceived music data and make judgments based on the semantic information converted by the speech recognition module.

4. A multi-modal wearable device real-time interaction system based on the light wing vision large model according to claim 3, characterized in that: The language translation module converts the text of the language into another language based on the language processed by the visual language model module, retaining the original meaning.

5. The multi-modal wearable device real-time interaction system based on the light wing vision large model according to claim 3 is characterized in that: The music recognition module is triggered by a user and instantly feeds back music recognition results based on the acquired semantic information data.

6. A method for implementing real-time interaction of a multimodal wearable device based on a light wing vision large model as claimed in claim 1, characterized in that: Includes steps Step 1: Acquire perception data through the perception component; Step 2: The acquisition module collects the acquired perception data and classifies it based on the data features; Step 3: Process the classified data based on the Light Wing Vision Large Model, including image recognition and dialogue. Step 4: The task management module uses the processed data in multi-scenario applications, and each function runs in coordination.

7. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement a real-time interaction method for a multimodal wearable device based on a light wing vision large model as described in claim 6.

8. A computer-readable storage medium, characterized in that Used to store a computer program, which, when executed by a processor, implements a real-time interaction method for a multimodal wearable device based on a light-wing vision large model as described in claim 6.

Citation Information

Patent Citations

  • Intelligent glasses control method and device, equipment and storage medium

    CN116300092A