Multi-modal perception lightweight visual enhancement device and method
By integrating multimodal sensors and optimizing the lightweight design of hardware layout, the integration and portability problems of existing visual enhancement devices are solved, the user experience and environmental perception capabilities are improved, and health monitoring and closed-loop control are realized.
Patent Information
- Application Number
- CN202510691810.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-23
AI Technical Summary
Existing vision enhancement devices have shortcomings in device integration, portability and weight control, which affect users' long-term usage experience.
A lightweight visual enhancement device with multimodal perception is designed, including a main component, a display unit, a sensor integration unit, a computing unit, and a power unit. It integrates vision, hearing, motion posture detection, eye tracking, and vital signs monitoring modules, optimizes the hardware layout to reduce weight, and performs multimodal data fusion and display through the computing unit.
It improves user wearing comfort, enhances environmental perception and user experience, and realizes real-time feedback of health status and visual-physiological closed-loop control.
Smart Images

Figure CN120686472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of augmented reality technology, and in particular to a lightweight visual enhancement device and method for multimodal perception. Background Art
[0002] Visual enhancement technology (VET) is a key branch of computer vision and artificial intelligence. Its core foundation is computer vision and image processing, but it also involves deep learning, sensor technology, AR / VR, multimodal perception, and other fields. Using advanced algorithms like deep learning and image processing, VET optimizes images and videos through processes such as denoising, super-resolution reconstruction, and color enhancement, further improving the readability and interactivity of information.
[0003] Multimodal perception is to build a multi-dimensional environmental perception system by integrating different types of sensor data, such as vision, hearing, and body posture, to provide intelligent systems with more comprehensive and accurate environmental information, thereby significantly improving the system's perception capabilities and decision-making efficiency.
[0004] In recent years, the integration of visual enhancement technology and multimodal perception has become a cutting-edge research direction in artificial intelligence and computer vision. This organic combination not only overcomes the limitations of a single modality in information acquisition and processing, but also provides stronger support for complex tasks.
[0005] However, the relevant products on the market currently have serious deficiencies in terms of equipment integration, portability, weight control, etc., which make it difficult to meet long-term usage needs and affect the user experience. Summary of the Invention
[0006] The purpose of the present invention is to provide a lightweight visual enhancement device and method with multimodal perception, improve the weight distribution and environmental perception capabilities of the overall device hardware, increase the user's wearing comfort, and enhance the user experience.
[0007] In order to solve the above technical problems, the embodiments of the present invention provide a technical solution as follows:
[0008] A lightweight visual enhancement device for multimodal perception, comprising:
[0009] Main body assembly, display unit, sensor integration unit, computing unit, power unit and near-infrared light source;
[0010] The main body component is a hollow cavity structure composed of a front shell and a rear shell, including a frame unit and a temple unit, and the sensor integration unit and the near-infrared light source are integrated into the inner cavity of the main body component;
[0011] The display unit includes a micro display screen assembly fixed to the frame unit, a micro projection optical engine, and an optical waveguide lens, wherein the micro display screen assembly is used to provide virtual image source data, the micro projection optical engine projects the virtual image source data to the optical waveguide lens, and the optical waveguide lens fuses the virtual image with the real world to form an augmented reality image;
[0012] The sensor integration unit is used for multi-sensor data acquisition, and includes: a visual perception module, an auditory perception module, a motion posture detection module, an eye tracking module, and a vital sign monitoring module;
[0013] The computing unit is used to process the data collected by the sensor integration unit and perform feature extraction and multimodal information fusion;
[0014] The power unit is used to provide power required by the display unit, the sensor integration unit, and the computing unit;
[0015] The near-infrared light source includes a micro infrared LED array distributed in a ring around the display unit and is used for illuminating the eye area.
[0016] Furthermore, the visual perception module includes an RGB-D camera located on the upper part of the frame unit shell and a high-definition eye-tracking camera facing the user. The RGB-D camera integrates an RGB sensor and a ToF depth sensor for real-time environmental perception and three-dimensional reconstruction of the environment; the eye-tracking camera and the near-infrared light source form a closed-loop control system, and calculates the line of sight direction in real time through the corneal reflection-pupil center vector analysis algorithm.
[0017] Furthermore, the auditory perception module includes a microphone array arranged on the main body component and a micro speaker arranged on the inner side of the temple unit.
[0018] Furthermore, the microphone is an 8-channel MEMS microphone, and the speaker is a bone conduction speaker integrated at the end of the temple unit.
[0019] Furthermore, the motion posture detection module includes an inertial measurement unit and a UWB positioning module, and the inertial measurement unit includes an accelerometer, a gyroscope and a geomagnetic sensor.
[0020] Furthermore, the vital signs monitoring module is arranged inside the temple unit and includes a flexible PPG sensor, which is in contact with the user's skin when the main body component is worn on the human body.
[0021] Furthermore, the sensor integration unit also includes an axial length measurement module, which is arranged at a position corresponding to the inner shell and the human eye, and is used for real-time monitoring of the axial length of the eye.
[0022] In order to solve the technical problem raised in this application, a lightweight visual enhancement method based on multimodal perception of the above-mentioned device is also provided, comprising:
[0023] Step S10: synchronously collecting multi-sensor data through the visual perception module, the auditory perception module, the motion posture detection module, the eye tracking module, and the vital signs monitoring module;
[0024] Step S20: extracting multimodal features from the collected data through the computing unit, and fusing and analyzing the extracted features;
[0025] Step S30: The data processed in S20 is integrated with the real world and displayed through the display unit.
[0026] Furthermore, step S10 also includes synchronously collecting axial length data through an axial length measurement module.
[0027] Furthermore, it also includes dynamically adjusting display parameters according to the user's physiological indicator data collected by the vital signs monitoring module.
[0028] The lightweight visual enhancement device and method with multimodal perception provided by the present invention, compared with the existing technology, optimizes the layout of the display unit, sensor integration unit, power unit, etc. on the main component by integrating sensor units such as the visual perception module, auditory perception module, motion posture detection module, eye tracking module, and vital sign monitoring module into the inner cavity of the main component, optimizes the weight distribution of each hardware, effectively reduces the user's wearing burden, increases the user's wearing comfort, and improves the user experience; the computing unit extracts multimodal features of the collected data under the visual mode, auditory mode, gaze point and other modes, and fuses and analyzes the extracted features to improve the efficiency and comprehensiveness of environmental perception, realize the presentation of augmented reality images, and enhance the user experience; through the setting of the vital sign monitoring module and the axial length detection module, it is possible to detect and monitor the user's physiological indicators, dynamically adjust the display parameters according to the physiological indicator data, realize real-time feedback on health status, realize visual-physiological closed-loop feedback control, and further improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0030] Figure 1 Schematic diagram of the overall structure of a lightweight visual enhancement device with multimodal perception according to an embodiment of the present invention;
[0031] Figure 2 1 is a schematic diagram of the overall structure of the back of the lightweight vision enhancement device with multimodal perception in an embodiment of the present invention;
[0032] Figure 3 1 is a schematic diagram of the structure of a lightweight visual enhancement device with multimodal perception in an embodiment of the present invention without the front shell;
[0033] Figure 4 This is a flowchart of the steps of the lightweight visual enhancement method for multimodal perception in an embodiment of the present invention.
[0034] Explanation of the accompanying drawings: 10. Main body assembly; 11. Front shell; 12. Back shell; 13. Frame unit; 14. Temple unit; 15. Inner cavity; 20. Display unit; 21. Optical waveguide lens; 22. Micro display assembly; 30. Sensor integration unit; 31. RGB-D camera; 32. Eye movement camera; 33. Auditory perception module; 331. Microphone; 332. Speaker; 34. Motion posture detection module; 35. Axis length measurement module; 36. Vital signs monitoring module; 40. Near-infrared light source; 50. Power unit; 51. Charging port. DETAILED DESCRIPTION
[0035] To make the objectives, technical solutions, and advantages of the present invention more apparent, various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will appreciate that many technical details are provided in various embodiments of the present invention to facilitate a better understanding of the present application. However, even without these technical details and the various variations and modifications based on the following embodiments, the technical solutions claimed in the claims of this application can be implemented.
[0036] like Figure 1-3 As shown, one embodiment of the present invention relates to a lightweight vision enhancement device for multimodal perception, including: a main body component 10, a display unit 20, a sensor integration unit 30, a computing unit, a power unit 50 and a near-infrared light source 40.
[0037] The main assembly 10 is a hollow cavity structure composed of a front shell 11 and a rear shell 12, including a frame unit 13 and a temple unit 14. The sensor integration unit 30 and the near-infrared light source 40 are integrated into the inner cavity 15 of the main assembly 10. The temple unit 14 is hingedly connected to the frame unit 13. The curve of the main assembly 10 conforms to ergonomics. The temple unit 14 is worn above the human ear and provides stable support for the main assembly 10. Preferably, the main assembly 10 adopts a lightweight design, combined with ergonomic optimization, to ensure the overall comfort of the device during long-term use. By optimizing the layout of sensors and the weight distribution of the device, the wearing burden of the user is reduced.
[0038] The display unit 20 includes a micro-display assembly 22 fixed to the frame unit 13, a micro-projection optical engine, and an optical waveguide lens 21. When the main assembly 10 is worn on the user, the display unit 20 corresponds to the user's eyes and is located in front of the eyes. The micro-display assembly 22 is used to provide virtual image source data; the micro-projection optical engine amplifies the image source data from the micro-display assembly 22 and projects it into the optical waveguide lens 21. The optical waveguide lens 21 uses the principle of total internal reflection to fuse the virtual image with the real world to form an augmented reality image, which is then transmitted to the user's eyes.
[0039] The sensor integration unit 30 is used for multi-sensor data acquisition and includes a visual perception module, an auditory perception module 33, a motion posture detection module 34, an eye tracking module, an eye axial length measurement module 35, and a vital sign monitoring module 36. The visual perception module includes an RGB-D camera 31 located on the upper portion of the frame unit 13 and a high-definition eye tracking camera 32 positioned facing the user. The RGB-D camera 31 integrates an RGB sensor and a Time-of-Flight (ToF) depth sensor to simultaneously capture color images and depth information, enabling real-time environmental perception and three-dimensional reconstruction of the environment. In one embodiment, the RGB-D camera 31 comprises a composite optical module, including a time-division multiplexing transmitter unit with a switchable narrow-bandpass filter array, which alternately transmits 850nm ToF pulses and 940nm structured light coded patterns. The hybrid sensor array comprises a 2-megapixel RGB sensor interlaced with a 300,000-megapixel ToF pixel area. Using a depth confidence mapping algorithm, the ToF phase data and structured light coded data are weightedly fused to generate a high-confidence depth map.
[0040] The high-definition eye-tracking camera 32 is fixed on the inner shell and is set facing the user for capturing images of the human eye. The eye-tracking camera 32 and the near-infrared light source 40 form a closed-loop control system. The pupil center and the corneal reflection point are identified through the calculation unit by combining the eye image with the reflected light generated by the near-infrared light source 40 on the eye, and the line of sight direction is calculated in real time using the corneal reflection-pupil center vector analysis algorithm. The collected eye movement data is processed and analyzed to achieve an eye tracking effect.
[0041] The auditory perception module 33 includes a microphone 331 and a micro-speaker 332; the microphone 331 array is arranged on the frame unit 13 and the temple unit 14, for capturing ambient sound and user voice, and realizing directional audio capture and noise reduction functions through audio processing technology and voice recognition algorithm; the micro-speaker 332 is installed at the ends of the left and right temple units 14, for providing audio feedback to the user; preferably, the microphone 331 is an 8-channel MEMS microphone 331 arranged along the frame unit 13 and the temple unit 14; the speaker 332 is a bone conduction speaker 332 integrated at the end of the temple unit 14 and in contact with the temporal bone, to realize open audio feedback.
[0042] The motion posture detection module 34 is disposed in the inner cavity 15 of the main component 10 and fixed to the inner wall of the rear shell 12. It includes an inertial measurement unit and a UWB positioning module. The inertial measurement unit is used to measure the three-axis attitude angle and acceleration of the main component 10 when it is worn on the human body. The UWB positioning module is used for high-precision positioning of the main component 10, which can achieve 10-centimeter spatial positioning of the main component. The inertial measurement unit includes an accelerometer, a gyroscope, and a geomagnetic sensor. By measuring acceleration and angular velocity and combining it with the UWB positioning module data, the motion posture detection module 34 realizes multi-source sensor data fusion through an extended Kalman filter. The computing unit calculates the posture and motion trajectory of the main component 10 in real time, thereby realizing motion posture detection.
[0043] The vital signs monitoring module 36 uses a flexible PPG sensor installed inside the left temple. When worn, it contacts the user's skin and emits light of a specific wavelength (such as red light or near-infrared light) to illuminate the skin. It detects changes in the intensity of the reflected light and converts these changes into electrical signals to extract physiological indicators such as the user's heart rate and blood oxygen saturation.
[0044] Preferably, the sensor integration unit 30 also includes an axial length measurement module 35, which is centrally arranged on the side of the rear shell 12 facing the user, and measures the distance from the anterior surface of the cornea to the retinal pigment epithelium by optical methods, thereby realizing real-time detection of the axial length and collecting axial length data.
[0045] The computing unit is used to process the data collected by the sensor integration unit 30 and perform feature extraction and multimodal information fusion. Preferably, the computing unit realizes spatiotemporal alignment and feature fusion of multi-sensor data through a heterogeneous computing architecture;
[0046] The power unit 50 is used to provide the power required by the display unit 20, the sensor integration unit 30, and the computing unit. Preferably, the power unit 50 adopts a distributed power supply design, an integrated flexible battery pack is fixed in the inner cavity 15, and a charging port 51 is provided at the end of the temple unit 14;
[0047] The near-infrared light source 40 includes a micro near-infrared LED array distributed in a ring. The micro near-infrared LEDs are installed around the display unit 20, with the eye area as the lighting target and facing the human body. The near-infrared LEDs uniformly illuminate the eye area to achieve active eye tracking lighting, reduce shadows and ensure that the eye camera 32 captures clear pupil and corneal reflection light spots. The computing unit can calculate the line of sight direction based on the geometric relationship between the pupil and corneal reflection light spots. The eye tracking module is used for gaze point data collection.
[0048] The present invention provides a lightweight visual enhancement device with multimodal perception. Through the lightweight design and ergonomic optimization of the main component 10, the layout of the display unit 20, the sensor integration unit 30, the power unit 50, etc. on the main component 10 optimizes its weight distribution, effectively reduces the user's wearing burden, and increases the user's wearing comfort; through the setting of multiple sensors such as the visual perception module, the auditory perception module 33, the motion posture detection module 34, the eye tracking module, the vital signs monitoring module 36, and the axial length measurement module 35, it can achieve efficient and accurate environmental perception, and monitor the user's axial length, heart rate, blood oxygen saturation and other health data in real time, thereby improving the perception ability of the overall device and user experience.
[0049] like Figure 4 As shown, one embodiment of the present invention relates to a lightweight visual enhancement method for multimodal perception, the method comprising the following steps:
[0050] Step S10: Synchronously collect multi-sensor data through the visual perception module, the auditory perception module 33, the motion posture detection module 34, the eye tracking module, and the vital signs monitoring module 36; wherein, the visual perception module is used to collect real-time environmental information data, the auditory perception module 33 is used to collect environmental sounds and user voice data, the motion posture detection module 34 is used to collect acceleration, angular velocity and position information data, the eye tracking module is used for user gaze point data, and the vital signs monitoring module 36 is used to collect user physiological indicator data, and the physiological indicator data includes heart rate, blood oxygen saturation, etc. Preferably, the collection of sensor data also includes the synchronous collection of axial length data through the axial length measurement module 35.
[0051] Step S20: extracting multimodal features from the collected data through the computing unit, and fusing and analyzing the extracted features;
[0052] Step S30: The data processed in step S20 is integrated with the real world for display through the display unit 20.
[0053] In one embodiment, the method further includes dynamically adjusting display parameters according to the user's physiological indicators to achieve visual-physiological closed-loop feedback control.
[0054] In an exemplary example, a lightweight visual enhancement method for multimodal perception is involved, including using a spatiotemporal calibration matrix to achieve synchronous acquisition of multi-sensor data through a visual perception module, an auditory perception module 33, a motion posture detection module 34, an eye tracking module, and a vital sign monitoring module 36; a computing unit uses an attention mechanism to perform multimodal feature extraction, and the feature extraction includes at least: environmental semantic segmentation feature extraction of data collected by the visual perception module, sound source orientation feature extraction of data collected by the auditory perception module 33, and gaze point heat map feature extraction of data collected by the eye tracking module, cross-modal feature fusion and analysis are performed through a cascade neural network to generate contextual perception information of the augmented reality scene, and the fused data is displayed to the user through the display unit 20; preferably, the display parameters are dynamically adjusted according to the user physiological indicators collected by the vital sign monitoring module 36 to achieve visual-physiological closed-loop feedback control and further improve the user experience.
[0055] The present invention provides a lightweight multimodal perception visual enhancement device and method. By integrating sensor units such as a visual perception module, an auditory perception module, a motion posture detection module, an eye tracking module, and a vital sign monitoring module into the inner cavity of a main component, and by arranging the display unit, sensor integration unit, and power unit on the main component, the weight distribution of each hardware is optimized, effectively reducing the user's wearing burden, increasing wearing comfort, and improving the user experience. A computing unit extracts multimodal features from data collected in modalities such as visual, auditory, and gaze points, and fuses and analyzes the extracted features to improve the efficiency and comprehensiveness of environmental perception, achieve the presentation of augmented reality images, and enhance the user experience. The provision of a vital sign monitoring module and an axial length detection module enables the detection and monitoring of user physiological indicators. Display parameters are dynamically adjusted based on physiological indicator data, providing real-time feedback on health status and implementing visual-physiological closed-loop feedback control, further improving the user experience. By combining multimodal perception with lightweight visual enhancement technology, the present invention enhances the overall device's perception capabilities and user experience, making it suitable for a variety of application scenarios such as intelligent security, autonomous driving, and industrial inspection.
[0056] Those skilled in the art will appreciate that the above-mentioned embodiments are specific examples for implementing the present invention, and that in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present invention.
Claims
1. A lightweight visual enhancement device for multimodal perception, characterized in that: include: A main body assembly (10), a display unit (20), a sensor integration unit (30), a computing unit, a power unit (50), and a near-infrared light source (40); The main body component (10) is a hollow cavity structure composed of a front shell (11) and a rear shell (12), including a frame unit (13) and a temple unit (14); the sensor integration unit (30) and the near-infrared light source (40) are integrated into the inner cavity (15) of the main body component (10); The display unit (20) comprises a micro display screen assembly (22) fixed to the frame unit (13), a micro projection optical machine and an optical waveguide lens (21), wherein the micro display screen assembly (22) is used to provide virtual image source data, the micro projection optical machine projects the virtual image source data to the optical waveguide lens (21), and the optical waveguide lens (21) fuses the virtual image with the real world to form an augmented reality image; The sensor integration unit (30) is used for multi-sensor data collection, and includes: a visual perception module, an auditory perception module (33), a motion posture detection module (34), an eye tracking module, and a vital sign monitoring module (36); The computing unit is used to process the data collected by the sensor integration unit (30) and perform feature extraction and multimodal information fusion; The power unit (50) is used to provide power required by the display unit (20), the sensor integration unit (30), and the computing unit; The near-infrared light source (40) comprises a micro infrared LED array distributed in a ring around the display unit (20) and is used for illuminating the eye area.
2. The lightweight visual enhancement device for multimodal perception according to claim 1, characterized in that: The visual perception module comprises an RGB-D camera (31) located on the upper part of the housing of the frame unit (13) and a high-definition eye-tracking camera (32) facing the user. The RGB-D camera (31) integrates an RGB sensor and a ToF depth sensor for real-time environmental perception and three-dimensional reconstruction of the environment. The eye-tracking camera (32) and a near-infrared light source (40) form a closed-loop control system, and the line of sight direction is calculated in real time through a corneal reflection-pupil center vector analysis algorithm.
3. The lightweight visual enhancement device for multimodal perception according to claim 1, characterized in that: The auditory perception module (33) comprises microphones (331) arranged in an array on the main body component (10) and micro speakers (332) arranged on the inner side of the temple unit (14).
4. The lightweight visual enhancement device for multimodal perception according to claim 3, characterized in that: The microphone (331) is an 8-channel MEMS microphone (331), and the speaker (332) is a bone conduction speaker (332) integrated at the end of the temple unit (14).
5. The lightweight visual enhancement device for multimodal perception according to claim 1, characterized in that: The motion posture detection module (34) includes an inertial measurement unit and a UWB positioning module, and the inertial measurement unit includes an accelerometer, a gyroscope and a geomagnetic sensor.
6. The lightweight visual enhancement device for multimodal perception according to claim 1, characterized in that: The vital sign monitoring module (36) is arranged inside the temple unit (14) and includes a flexible PPG sensor, which is in contact with the user's skin when the main body component (10) is worn on the human body.
7. The lightweight visual enhancement device for multimodal perception according to claim 1, characterized in that: The sensor integration unit (30) further comprises an eye axis length measurement module (35), which is arranged at a position corresponding to the inner shell and the human eye and is used for real-time monitoring of the eye axis length.
8. A lightweight visual enhancement method based on multimodal perception of the device according to any one of claims 1 to 7, characterized in that: include: Step S10: synchronously collecting multi-sensor data through the visual perception module, the auditory perception module (33), the motion posture detection module (34), the eye tracking module, and the vital signs monitoring module (36); Step S20: extracting multimodal features from the collected data through the computing unit, and fusing and analyzing the extracted features; Step S30: The data processed in S20 is integrated with the real world and displayed through the display unit (20).
9. The lightweight visual enhancement method for multimodal perception according to claim 8, characterized in that: Step S10 also includes synchronous collection of axial length data by an axial length measurement module (35).
10. The lightweight visual enhancement method for multimodal perception according to claim 8, characterized in that: It also includes dynamically adjusting display parameters according to the user's physiological index data collected by the vital sign monitoring module (36).
Citation Information
Patent Citations
Augmented reality glasses based on multi-mode imaging and multi-layer perception
CN108919498A
Multi-camera biological recognition imaging system
CN116529786A
Augmented reality glasses, diopter detection method, electronic equipment and storage medium
CN117369142A
Combined birefringent material and reflective waveguide for multiple focal planes in mixed reality head mounted display device
CN117980808A