Multi-source fusion vision enhancement head-mounted display system and device
By designing a multi-source fusion vision-enhanced head-mounted display system and integrating multiple sensors and computing modules, the problems of limited visual display, lack of night vision functions and insufficient interactive experience of existing AR devices are solved, and efficient human-computer interaction and all-weather reconnaissance perception are achieved.
Patent Information
- Application Number
- CN202510080837.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-06
AI Technical Summary
Existing AR glasses equipment has problems such as limited visual display, lack of night vision function, and insufficient interactive experience, and cannot effectively realize multi-source data fusion and night vision functions, which limits the use of complex outdoor environments.
A multi-source fusion vision-enhanced head-mounted display system is designed, including a head-mounted AR device module and a main control computing unit module. It adopts multi-sensing fusion, spatial computing and optimized human-computer interaction methods, integrating a tracking system, a depth system, an audio system, an infrared low-light enhancement system and an eye movement system to achieve efficient human-computer interaction through multi-source data fusion.
It improves human-computer interaction efficiency, realizes natural and harmonious human-computer interaction, improves the accuracy and robustness of gesture recognition, enhances gesture feature expression capabilities, improves the interactivity and ease of use of head-mounted AR devices, and supports all-weather reconnaissance perception and immersive deep virtual and real interaction.
Smart Images

Figure CN120103967A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of AR equipment, and in particular relates to a multi-source fusion vision enhancement head-mounted display system and device. Background Art
[0002] With the development of augmented reality (AR) technology, head-mounted display devices are gradually applied to tactical, industrial and consumer fields. Existing AR glasses have problems such as limited visual display, lack of night vision function, and insufficient interactive experience. They usually lack support for multi-source data fusion, cannot effectively realize the real-time fusion display of low-light and infrared images, and often do not have night vision function, which limits their use in complex outdoor environments. At the same time, the interaction mode of existing devices is single, and they fail to make full use of emerging human-computer interaction technologies such as eye tracking and gesture recognition, and cannot be accurately controlled by gestures or eye movements. To this end, a multi-source fusion vision-enhanced head-mounted display system is proposed to improve the user experience of the device through multi-sensor fusion, spatial computing and optimized human-computer interaction. Summary of the invention
[0003] The purpose of the present invention is to provide a multi-source fusion vision enhanced head-mounted display system and device, which improves the human-computer interaction efficiency, realizes natural and harmonious human-computer interaction between people and head-mounted devices, improves the accuracy and robustness of gesture recognition, and enhances the ability to express gesture features.
[0004] In order to achieve the purpose of the present invention, on the one hand, the present invention provides a multi-source fusion visual enhancement head-mounted display system, comprising the following modules:
[0005] Head-mounted AR device module, used to provide clear augmented reality images under different lighting conditions and provide positioning and voice services;
[0006] The main control computing unit module is used to process display data and sensor data, and update the user interface and interactive response in real time, and is connected to the head-mounted AR device module via USB or wirelessly.
[0007] The head-mounted AR device module includes a tracking system, a depth system, an audio system, a display system, an infrared low-light enhancement system, and an eye movement system;
[0008] The tracking system calculates the pupil center based on image processing technology, obtains the optical axis of the eyeball through the line connecting the corneal curvature center and the pupil center, and determines the real sight direction - the visual axis - using the angle between the optical axis and the visual axis, so as to achieve the effect that the display image changes accordingly with the change of the position of the eye's gaze point;
[0009] The depth system uses the time-of-flight TOF module to obtain depth information, complete large-space mapping, and obtain a real 3D map of the scene in front of the device; the depth system uses a 3D-TOF image sensor to measure distance and size, track movement, and convert the shape of an object into a three-dimensional model;
[0010] The eye movement system is used to provide pupil distance recognition and gaze point tracking, support accurate user interaction, and enable users to control the operation of interface elements and virtual objects by gaze;
[0011] The audio system is used to realize the voice interaction function; the person and the multi-source fusion visual enhancement head-mounted display system transmit information through natural voice, which is compared with the interaction mode of gesture recognition and eye tracking; the audio and semantic information are obtained according to the voice recognition and noise reduction algorithm to complete the voice interaction;
[0012] The infrared low-light enhancement system is used for low-light night vision, infrared night vision, infrared and low-light dual-light fusion imaging, and displays images in dark and complex environments;
[0013] The optical display module of the display system is composed of an optical waveguide and a microdisplay, and is used for large-field-of-view near-eye display, realizing a transparent virtual-real fusion display, providing a high-resolution, low-latency visual experience, and ensuring the user's visual stability in dynamic scenes.
[0014] The main control computing unit module includes a VSLAM module, a display processing module, and a multi-source data fusion module; the main control computing unit module adopts the RK3588S chip, integrates the NPU to accelerate data processing and image calculation, and supports multi-threaded operation; the VSLAM module is used to realize six-degree-of-freedom head posture tracking, captures the user's spatial position information in real time through a binocular camera, and generates real-time spatial positioning data; the display processing module presents the fused augmented reality image to the user through a waveguide optical system; the multi-source data fusion module is used to simultaneously process data from multiple sensors, perform image registration and fusion, including the superposition of low-light and infrared data, and the fusion of gesture recognition data and environmental data.
[0015] On the other hand, the present invention provides a multi-source fusion visual enhancement head-mounted display device, which consists of a head-mounted AR device and a main control computing unit in a split architecture; the head-mounted AR device includes an infrared camera located in the center of the head-mounted AR device, with two symmetrical low-light cameras and a fisheye camera on both sides, a depth camera at the bottom, and two eye-controlled cameras respectively placed under the lenses of the optical display module; the outside of the head-mounted AR device is compatible with the helmet through an adjustable strap, and the audio system includes a speaker and a microphone, the speakers are on both sides of the head-mounted AR device, and the microphone is in the center of the head-mounted AR device; the main control computing unit is connected to the head-mounted AR device through a Type-C interface.
[0016] Compared with the prior art, the significant progress of the present invention lies in: the present invention realizes all-weather reconnaissance perception, large field of view, see-through augmented reality display and immersive deep virtual-reality interaction through integrated hardware integration, spatial computing, augmented reality near-eye display, multi-source data fusion, human-computer interaction and other technologies, thereby enhancing the user's perception of the external environment and the ability to quickly process information, improving the efficiency of human-computer interaction, achieving natural and harmonious human-computer interaction between people and head-mounted devices, improving the accuracy and robustness of gesture recognition, enhancing the ability to express gesture features, and improving the interactivity and ease of use of head-mounted AR devices.
[0017] In order to more clearly illustrate the functional characteristics and structural parameters of the present invention, further description is given below in conjunction with the accompanying drawings and specific implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0019] Figure 1 It is the structural composition of the multi-source fusion vision enhancement head-mounted display system of an embodiment of the present invention;
[0020] Figure 2 is a schematic diagram of a multi-source fusion vision enhancement head-mounted display system according to an embodiment of the present invention;
[0021] Figure 3 is a schematic diagram of a data fusion process of a low-light camera and an infrared camera according to an embodiment of the present invention;
[0022] Figure 4 is a system structure diagram of a main control computing unit according to an embodiment of the present invention;
[0023] Figure 5 is a schematic diagram of the working principle of the depth system of an embodiment of the present invention;
[0024] Figure 6Schematic diagram of the working principle of the eye movement system according to an embodiment of the present invention.
[0025] The reference numerals in the figure are: head-mounted AR device (100); low-light camera (110); infrared camera (120); fisheye camera (130); depth camera (140); eye-controlled camera (150); audio system (160); optical display module (170); VSLAM module (210); display processing module (220); and multi-source data fusion module (230). DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0027] The multi-source fusion visual enhancement head-mounted display system of the present invention is combined with Figure 1 , including the following modules:
[0028] Head-mounted AR device module, used to provide clear augmented reality images under different lighting conditions and provide positioning and voice services;
[0029] The main control computing unit module is used to process display data and sensor data, and update the user interface and interactive response in real time, and is connected to the head-mounted AR device module via USB or wirelessly.
[0030] The head-mounted AR device module includes a tracking system, a depth system, an audio system, a display system, an infrared low-light enhancement system, and an eye movement system;
[0031] The tracking system calculates the pupil center based on image processing technology, obtains the optical axis of the eyeball through the line connecting the corneal curvature center and the pupil center, and determines the real sight direction - the visual axis - using the angle between the optical axis and the visual axis, so as to achieve the effect that the display image changes accordingly with the change of the position of the eye's gaze point;
[0032] The depth system uses a time-of-flight (TOF) module to obtain depth information, complete large-space mapping, and obtain a true 3D map of the scene in front of the device, thereby achieving precise fusion of virtual and reality; the depth system uses a 3D-TOF image sensor to measure distance and size, track movement, and convert the shape of an object into a three-dimensional model.
[0033] The eye movement system is used to provide pupil distance recognition and gaze point tracking, support accurate user interaction, and enable users to control the operation of interface elements and virtual objects by gaze;
[0034] The audio system 160 is used to realize the voice interaction function; the person and the multi-source fusion visual enhancement head-mounted display system transmit information through natural voice. Compared with the interaction methods of gesture recognition and eye tracking, voice interaction can free both hands and eyes, has less restrictions, and is more convenient, greatly improving the interaction efficiency; the voice interaction is completed by obtaining audio and semantic information based on the voice recognition and noise reduction algorithm;
[0035] The infrared low-light enhancement system is used for low-light night vision, infrared night vision, infrared and low-light dual-light fusion imaging, and displays images in dark and complex environments;
[0036] The optical display module 170 of the display system is composed of an optical waveguide and a microdisplay, and is used for large-field-of-view near-eye display, realizing a transparent display of virtual-real fusion, providing a high-resolution, low-latency visual experience, and ensuring the user's visual stability in dynamic scenes.
[0037] Combination Figure 3 The infrared low-light enhancement system includes a binocular low-light camera 110 and an infrared camera 120, which are respectively installed at the front end of the AR device to obtain low-light and infrared images; the acquired infrared image extracts the target area through threshold segmentation, and extracts edge information from the segmentation result through edge extraction; the edge is morphologically processed to repair the edge details and enhance the target contour, and then the infrared image is pseudo-colored to improve the visualization effect of the image; the low-light image acquired by the low-light camera is first denoised to remove environmental noise and enhance the image quality; the infrared image processed with pseudo-color and the low-light image processed with noise reduction are image registered and fused to finally generate a fused image containing infrared and low-light information; the true color image fusion technology that conforms to the visual characteristics of the human eye can simultaneously retain the image features of the infrared image and the low-light image and eliminate redundant information, and the infrared image is first preprocessed and the low-light image is denoised before the two images are registered and fused; the fused image can be clearly displayed at night or in a low-light environment, and is equipped with an image enhancement algorithm.
[0038] The resolution of the binocular low-light camera 110 is 1280×960, and the resolution of the infrared camera 120 is 640×512, and the frame rate is 30 Hz, which ensures the clarity and detail of night vision images and meets the needs of various application scenarios.
[0039] The depth system includes a depth camera 140, which separates image frames containing gesture information from video streams captured by multiple front cameras, detects gestures in the video stream and extracts them according to the interactive model of gesture input; in the gesture analysis stage, feature detection and parameter estimation are performed on the separated gestures, and a skin color-based positioning technology is selected. A lookup table is established using skin training data for feature detection. At the same time, feature parameters such as the area, boundary, contour or fingertip of the gesture are initially estimated and updated over time, supporting the recognition of at least 10 static or dynamic gesture actions, including confirmation, cancellation, scaling, rotation, etc.; through the hand movements captured by the depth camera 140, the system generates three-dimensional gesture coordinates and performs spatial operations; the gesture recognition algorithm uses a convolutional neural network to extract and classify gesture features, has self-learning capabilities, and the recognition response time does not exceed 100 milliseconds.
[0040] Combination Figure 5 The depth system includes a depth camera 140, which is used to obtain image frames containing gesture information. The image information is used as an information source input for gesture recognition. Through hand detection, feature extraction, feature encoding and classification recognition, efficient static gesture recognition can be achieved; combined with Fisher vector encoding and SVM classifier, the accuracy and robustness of gesture recognition can be improved. The acquired image first needs to be feature extracted, and then the possible hand area in the image is detected through depth information to perform hand candidate area detection, and the background and noise irrelevant to the gesture are removed, the hand features are retained, and the arm information is removed; the segmented hand image is feature described As described above, local geometry and texture information are extracted, and the extracted local descriptors are converted into high-dimensional feature vectors through Fisher vector encoding to enhance the expression ability of gesture features; the encoded gesture features are classified by support vector machine SVM, and the classifier outputs the corresponding gesture category label to realize gesture recognition; according to the classification results, a gesture model is established and updated to generate an explanation or instruction for the current gesture to provide support for subsequent human-computer interaction; and support recognition of static or dynamic gesture actions, including confirmation, cancellation, scaling, and rotation; through the hand movements collected by the depth camera 140, the system generates three-dimensional gesture coordinates and performs spatial operations.
[0041] Combination Figure 6The eye movement system includes an eye-controlled camera 150 and an eye-tracking algorithm; the eye-controlled camera 150 is connected to the system through a driver, and the data acquisition function is started to capture the user's eye image in real time, monitor the eye movement, pupil position, center of the pupil and the size and position characteristics of the Pu'erqin spot; the acquired original image needs to be preprocessed, such as denoising, contrast enhancement, and extraction of key features of the eye, pupil and Pu'erqin spot; according to the image brightness and ambient light conditions, the camera exposure parameters are dynamically adjusted to ensure the stability of the captured image quality and perform exposure control; the processed image and related data are formatted and packaged and transmitted to the main control computing unit module in real time through a high-speed USB 3.0 interface. After receiving the data, the main control computing unit module uses the eye tracking algorithm to analyze the eye movement trajectory and pupil gaze point, and according to the gaze point analysis result, the user's gaze point is accurately mapped to the display interface of the head-mounted AR device module to realize the eye movement function.
[0042] The main control computing unit module includes a VSLAM (Visual Simultaneous Localization and Mapping) module 210, a display processing module 220, and a multi-source data fusion module 230;
[0043] The main control computing unit module has a built-in NPU for accelerating the visual SLAM algorithm and the three-dimensional reconstruction function of the depth camera; the main control computing unit module adopts the RK3588S chip, and integrates the NPU to accelerate data processing and image calculation, supports multi-threaded operation, improves the processing capability of the device in complex computing scenarios, and processes sensor data in real time and outputs high-quality display effects.
[0044] The VSLAM module 210 is used to implement six-degree-of-freedom head posture tracking, capture the user's spatial position information in real time through a binocular camera, and generate real-time spatial positioning data;
[0045] The display processing module 220 presents the fused augmented reality image to the user through a waveguide optical system;
[0046] The multi-source data fusion module 230 is used to simultaneously process data from multiple sensors and perform image registration and fusion, including superposition of low-light and infrared data, and fusion of gesture recognition data and environmental data.
[0047] Combination Figure 4The main control computing unit module is connected to the head-mounted AR device module via Type-C wired or Wi-Fi6 and Bluetooth 5.0 wireless connection to transmit sensor data in real time; the head-mounted AR device module transmits the data to the service process of the main control computing unit module via USB data, runs functions and algorithms, display rendering services and algorithms, and graphics API adaptation in the service process, and the service process provides services for the application process and connects to the API interface of the open source XR platform OpenXR, thereby supporting the application operation and development of the three-dimensional engine Unity or Unreal.
[0048] The audio system 160 supports spatial voice interaction function, and is used to provide a highly immersive audio experience through a built-in dual-channel audio system and noise reduction algorithm, and to control controls on the interface or trigger specific operation instructions through voice.
[0049] The present invention also includes a device for a multi-source fusion visual enhancement head-mounted display system, combined with Figure 2 , which consists of a head-mounted AR device 100 and a main control computing unit 200 in a split structure; the head-mounted AR device 100 includes an infrared camera (120) located in the center of the head-mounted AR device 100, two symmetrical low-light cameras 110 and a fisheye camera 130 on both sides, a depth camera 140 at the bottom, and two eye-controlled cameras 150 respectively placed under the lens of an optical display module 170; the outside of the head-mounted AR device 100 is compatible with a helmet through an adjustable strap; the audio system 160 includes a speaker and a microphone, the speakers are on both sides of the head-mounted AR device 100, and the microphone is in the center of the head-mounted AR device 100; the main control computing unit 200 is connected to the head-mounted AR device 100 through a Type-C interface.
[0050] The head-mounted AR device 100 adopts a brand-new hardware and appearance structure design, with built-in tracking system, display system, audio system, infrared low-light enhancement system, eye tracking system and other subsystem modules, and is also equipped with sensor cameras including spatial positioning camera, eye control camera, low-light camera, depth camera, infrared camera, and integrated inertial measurement unit, distance sensor, audio module, light sensor, etc.; the main control computing unit 200 is built based on RK3588S and includes application process and service process. The head-mounted AR device 100 is responsible for driving various hardware modules at the front end, and transmits sensor data to the main control computing unit via USB. The main control computing unit 200, as the core part of algorithm processing and content production, is responsible for producing application content, cooperating with the "service process included in the AR SDK" to optimize the content and apply various functional algorithms, and finally presents the content through the display channel.
[0051] Example
[0052] Example 1
[0053] The tracking system of the multi-source fusion vision-enhanced head-mounted display device adopts a two-eye inside-out VSLAM algorithm and works based on the real-time mapping and positioning principle of binocular vision. It uses the indirect method of sparse point cloud plus keyframes, and relies on the original efficient spatial description operator and algorithm to achieve high-speed VSLAM calculation at the head end, accurately output the camera's own 6DOF posture information, and effectively achieve high-precision spatial positioning and tracking, giving AR devices key AR functions such as spatial motion tracking and virtual object anchoring. At the same time, combined with the loop detection, mapping and optimization functions provided by the back-end algorithm, it can further improve the re-positioning accuracy and robustness, ensuring the stable operation of the system in complex environments.
[0054] Example 2
[0055] The depth system of the multi-source fusion visual enhancement head-mounted display device uses an indirect TOF algorithm. Specifically, the modulated light pulse is emitted by the infrared transmitter. After the light pulse is reflected by an object, the receiver receives the reflected light pulse and calculates the distance to the object based on the round-trip time of the light pulse. Among them, the image signal processor (ISP) chip performs depth calculation and processing operations on the original image data to obtain the depth value of each pixel, providing accurate data support for subsequent 3D reconstruction, depth perception and other applications.
[0056] Example 3
[0057] The infrared low-light system of the multi-source fusion visual enhancement head-mounted display device includes two low-light cameras and one infrared camera. The system performs pixel fusion processing on the infrared image and the low-light camera image on the AR device side, effectively improving the richness and availability of image information. Through the carefully designed video perspective principle, the visual perception process of the human eye is simulated, and the real-world scene is captured with the help of the built-in binocular real-scene camera to achieve real-world reconstruction and fusion with the virtual world image in the application content.
[0058] Example 4
[0059] The audio system of the multi-source fusion visual enhancement head-mounted display device adopts a dual-channel audio external amplifier design, and through the built-in advanced noise reduction algorithm and audio equalization algorithm, it brings users a highly immersive audio experience. The input and output data streams of its audio system are completed through the standard UAC (USB Audio Class) of the operating system.
[0060] Example 5
[0061] The display system of the multi-source fusion visual enhancement head-mounted display device has reconstructed the display rendering pipeline and provided core display optimization measures for mixed reality. The display front end is responsible for the initialization of the display and the adjustment and setting of parameters such as brightness to ensure the stability and comfort of the display effect. The display system reconstructs the operating system rendering pipeline, implements the self-developed display synthesizer, and performs a series of necessary display optimization operations on the rendering content, such as distortion and color correction, positional time warp (PTW), asynchronous time warp (ATW), asynchronous space warp (ASW), etc. At the same time, it enables the direct display and single buffer functions of the graphics card, reduces the display link latency, and makes the MTP (Motion to Photon) and display smoothness of the AR glasses reach the optimal state.
[0062] Example 6
[0063] The eye movement system of the multi-source fusion visual enhancement head-mounted display device is designed based on the principle of pupil corneal reflection method, with a design frame rate of up to 60fps. The system can not only provide real-time gaze point data, but also has the function of automatic pupil distance recognition. The eye movement front end is responsible for image processing and exposure setting after data sampling, and then packages the data and transmits it to the host via USB. The eye movement algorithm uses the built-in eye movement gaze point algorithm to accurately calculate the gaze point data and realize automatic identification of the pupil distance. Usually, the current user's pupil distance is output after the calibration module completes the calibration.
[0064] Example 7
[0065] The 3D gesture tracking system of the multi-source fusion visual enhanced head-mounted display device adopts an advanced 3D gesture recognition solution based on vision and deep learning technology. The gesture recognition algorithm constructs a 2D hand shape representation based on a unique Fisher vector encoding method, extracts multiple geometric features from the segmented human hand binary image to form a local descriptor, and then obtains the overall 2D hand shape representation through Fisher vector encoding, and uses a linear SVM classifier for efficient classification. The 3D coordinate estimation of the gesture uses the JGR-P20 model, which takes a single depth image as input and outputs the uvZ coordinates of each joint point through a complex network architecture and algorithm modules.
[0066] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0067] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-source fusion visual enhancement head-mounted display system, characterized in that: Includes the following modules: Head-mounted AR device module, used to provide clear augmented reality images under different lighting conditions and provide positioning and voice services; The main control computing unit module is used to process display data and sensor data, and update the user interface and interactive response in real time, and is connected to the head-mounted AR device module via USB or wirelessly.
2. A multi-source fusion visual enhancement head mounted display system according to claim 1, characterized in that: The head-mounted AR device module includes a tracking system, a depth system, an audio system, a display system, an infrared low-light enhancement system, and an eye movement system; The tracking system calculates the pupil center based on image processing technology, obtains the optical axis of the eyeball through the line connecting the corneal curvature center and the pupil center, and determines the real sight direction - the visual axis - using the angle between the optical axis and the visual axis, so as to achieve the effect that the display image changes accordingly with the change of the position of the eye's gaze point; The depth system uses the time-of-flight TOF module to obtain depth information, complete large-space mapping, and obtain a real 3D map of the scene in front of the device; the depth system uses a 3D-TOF image sensor to measure distance and size, track movement, and convert the shape of an object into a three-dimensional model; The eye movement system is used to provide pupil distance recognition and gaze point tracking, support accurate user interaction, and enable users to control the operation of interface elements and virtual objects by gaze; The audio system (160) is used to realize the voice interaction function; the person and the multi-source fusion visual enhancement head mounted display system transmit information through natural voice, which is compared with the interaction mode of gesture recognition and eye tracking; the audio and semantic information are obtained according to the voice recognition and noise reduction algorithm to complete the voice interaction; The infrared low-light enhancement system is used for low-light night vision, infrared night vision, infrared and low-light dual-light fusion imaging, and displays images in dark and complex environments; The optical display module (170) of the display system is composed of an optical waveguide and a microdisplay, and is used for large-field-of-view near-eye display, realizing a transparent virtual-real fusion display, providing a high-resolution, low-latency visual experience, and ensuring the visual stability of users in dynamic scenes.
3. The multi-source fusion visual enhancement head mounted display system according to claim 2, characterized in that: The infrared low-light enhancement system comprises a binocular low-light camera (110) and an infrared camera (120) which are respectively installed at the front end of an AR device for acquiring low-light and infrared images; the acquired infrared image extracts a target area through threshold segmentation, and extracts edge information from the segmentation result through edge extraction; morphological processing is performed on the edge to repair edge details and enhance the target contour, and then pseudo-color processing is performed on the infrared image to improve the visualization effect of the image; the low-light image acquired by the low-light camera is first subjected to noise reduction to remove environmental noise and enhance image quality; the infrared image subjected to pseudo-color processing and the low-light image subjected to noise reduction are image registered and fused, and finally a fused image containing infrared and low-light information is generated; the true color image fusion technology that conforms to the visual characteristics of the human eye can simultaneously retain the image features of the infrared image and the low-light image and eliminate redundant information, and the infrared image is first preprocessed and the low-light image is subjected to noise reduction before the two images are registered and fused; The fused image can be displayed clearly at night or in low-light environments and is equipped with an image enhancement algorithm.
4. The multi-source fusion visual enhancement head mounted display system according to claim 2, characterized in that: The resolution of the binocular low-light camera (110) is 1280×960, the resolution of the infrared camera (120) is 640×512, and the frame rate is 30 Hz.
5. The multi-source fusion visual enhancement head mounted display system according to claim 2, characterized in that: The depth system comprises a depth camera (140) for acquiring image frames containing gesture information, the image information being input as an information source for gesture recognition, and realizing efficient static gesture recognition through hand detection, feature extraction, feature encoding and classification recognition; combining Fisher vector encoding and SVM classifier, the acquired image first needs to be feature extracted, and then the possible hand area in the image is detected through depth information to detect the hand candidate area, and the background and noise irrelevant to the gesture are removed, the hand features are retained, and the arm information is removed; the segmented hand image is feature described, the local geometry and texture information is extracted, and the extracted local descriptor is converted into a high-dimensional feature vector through Fisher vector encoding; the encoded gesture features are classified by using a support vector machine (SVM), and the classifier outputs the corresponding gesture category label to realize gesture recognition; according to the classification result, a gesture model is established and updated to generate an explanation or instruction for the current gesture; It also supports the recognition of static or dynamic gestures, including confirmation, cancellation, zooming, and rotation; through the hand movements captured by the depth camera (140), the system generates three-dimensional gesture coordinates and performs spatial operations.
6. The multi-source fusion visual enhancement head mounted display system according to claim 2, characterized in that: The eye movement system comprises an eye control camera (150) and an eye movement tracking algorithm; the eye control camera (150) is connected to the system via a driver program, and a data acquisition function is activated to capture the user's eye image in real time, and monitor eye movement, pupil position, pupil center, and the size and position characteristics of the Pu'erqin spots; The acquired original image needs to be preprocessed to extract the key features of the eyeball, pupil and Pu'erqin spots; the camera exposure parameters are dynamically adjusted according to the image brightness and ambient light conditions to ensure the stability of the captured image quality and perform exposure control; the processed image and related data are formatted and packaged and transmitted to the main control computing unit module in real time through the high-speed USB3.0 interface. After receiving the data, the main control computing unit module uses an eye tracking algorithm to analyze the eye movement trajectory and pupil gaze point, and according to the gaze point analysis results, the user's gaze point is accurately mapped to the display interface of the head-mounted AR device module to realize the eye movement function.
7. The multi-source fusion visual enhancement head mounted display system according to claim 1, characterized in that: The main control computing unit module comprises a VSLAM module (210), a display processing module (220), and a multi-source data fusion module (230); the main control computing unit module adopts an RK3588S chip, integrates an NPU for accelerating data processing and image calculation, and supports multi-threaded operation; the VSLAM module (210) is used to achieve six-degree-of-freedom head posture tracking, captures the user's spatial position information in real time through a binocular camera, and generates real-time spatial positioning data; the display processing module (220) presents the fused augmented reality image to the user through a waveguide optical system; The multi-source data fusion module (230) is used to simultaneously process data from multiple sensors and perform image registration and fusion, including superposition of low-light and infrared data, and fusion of gesture recognition data and environmental data.
8. The multi-source fusion visual enhancement head mounted display system according to claim 1, characterized in that: The main control computing unit module is connected to the head-mounted AR device module via Type-C wired or Wi-Fi6 and Bluetooth 5.0 wireless connection to transmit sensor data in real time; the head-mounted AR device module transmits the data to the service process of the main control computing unit module via USB data, runs functions and algorithms, display rendering services and algorithms, and graphics API adaptation in the service process, and the service process provides services for the application process and connects to the API interface of the open source XR platform OpenXR, thereby supporting the application operation and development of the three-dimensional engine Unity or Unreal.
9. The multi-source fusion visual enhancement head mounted display system according to claim 2, characterized in that: The audio system (160) supports a spatial voice interaction function, and is used to provide a highly immersive audio experience through a built-in dual-channel audio system and a noise reduction algorithm, and to control controls on an interface or trigger specific operation instructions through voice.
10. A device for a multi-source fusion visual enhancement head mounted display system according to any one of claims 1 to 9, characterized in that: The invention comprises a head-mounted AR device (100) and a main control computing unit (200) in a split-type structure; the head-mounted AR device (100) comprises an infrared camera (120) located in the center of the head-mounted AR device (100), two symmetrical low-light cameras (110) and a fisheye camera (130) on both sides, a depth camera (140) at the bottom, and two eye-controlled cameras (150) respectively placed under the lenses of an optical display module (170); the outside of the head-mounted AR device (100) is compatible with a helmet via an adjustable strap; the audio system (160) comprises a speaker and a microphone, the speakers are located on both sides of the head-mounted AR device (100), and the microphone is located in the center of the head-mounted AR device (100); the main control computing unit (200) is connected to the head-mounted AR device (100) via a Type-C interface.