Intelligent glasses environment sensing method and system based on machine vision
Through the multi-sensor tightly coupled factor graph optimization framework and dynamic environment perception algorithm, smart glasses achieve high-precision, low-power dynamic environment perception and natural interaction, solving the problems of insufficient perception accuracy and fragmented interactive experience in existing technologies. It is suitable for all-weather industrial inspections and drone collaborative operations.
Patent Information
- Application Number
- CN202511201421.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing smart glasses have problems in environmental perception, such as insufficient perception accuracy, contradiction between real-time and power consumption, lack of dynamic obstacle processing, and fragmented interactive experience. In particular, it is difficult to achieve high-precision perception, low-power computing and natural interaction in dynamic environments.
It adopts a multi-sensor tightly coupled factor graph optimization framework, combines binocular cameras, IMU and millimeter-wave radar, and uses motion clustering algorithm and social-LSTM trajectory prediction model to build a dynamic three-dimensional environment map, and realizes natural interaction through AR display and spatial audio algorithm.
In a dynamic environment, the horizontal positioning error is ≤1.2cm, the height error is ≤2cm, the depth estimation error is ≤5cm, the system power consumption is <1.2W, the battery life is extended to 10 hours, the user's visual fatigue is reduced by 50%, and it supports a wide range of scene applications.
Smart Images

Figure CN120708114A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and wearable devices, and in particular to a method and system for environmental perception of smart glasses based on machine vision. Background Art
[0002] With the rapid development of augmented reality (AR) technology, smart glasses have shown great potential in the field of environmental perception and interaction. However, existing technologies still have the following key issues: Insufficient perception accuracy: Existing technologies mostly rely on monocular cameras, which cannot obtain depth information. This leads to large errors in obstacle distance estimation (typical error >15%). In dynamic environments, the lack of multi-sensor fusion results in low recognition rates for moving targets. The contradiction between real-time performance and power consumption: Traditional SLAM algorithms require high-performance GPU support, making them difficult to deploy on mobile devices. If simplified models are used, the integrity of environmental modeling is sacrificed. Lack of dynamic obstacle handling: Existing technologies are mostly designed for static environments and do not consider trajectory prediction for moving objects such as pedestrians and vehicles, resulting in a high risk of obstacle avoidance path planning failure. Fragmented interactive experience: AR display and voice prompts lack synergy, multimodal feedback suffers from information overload, and user context understanding is inefficient. Based on the above problems, there is an urgent need for a smart glasses solution that can take into account high-precision perception, low-power computing, dynamic environment adaptability and natural interaction. Summary of the Invention
[0003] The purpose of the present invention is to provide a machine vision-based smart glasses environmental perception method and system. Based on a multi-sensor tightly coupled factor graph optimization framework, including binoculars + IMU + millimeter wave radar, it can achieve horizontal positioning error ≤ 1.2cm and height error ≤ 2cm in a dynamic environment. Through the motion clustering algorithm and social-LSTM trajectory prediction model, it can accurately distinguish between static and dynamic targets, aiming to solve the problems in the existing technology.
[0004] The present invention is implemented as follows: a method for intelligent glasses environment perception based on machine vision, applied to VR wearable devices, specifically comprising the following steps: S101: Collecting environmental stereo image data in real time through a binocular camera, synchronously acquiring multimodal sensor data, and storing them in a cache queue in a timestamp-aligned manner; S102: performing dedistortion and optical flow compensation preprocessing on the stereo image data, and generating pose increments from the multimodal sensor data through a pre-integration algorithm, and performing spatial alignment and fusion with the obstacle point cloud data from the millimeter-wave radar; S103: A tightly coupled visual-inertial SLAM algorithm is used to construct a dynamic 3D environment map based on the fused multi-source data. The static background of the 3D environment map is modeled using a voxel grid, and dynamic obstacles are marked as independent moving targets through motion clustering. S104: A lightweight semantic segmentation network is used to perform pixel-by-pixel classification of the 3D environment map to identify the type, geometry, and motion state of obstacles. Based on the semantic segmentation results and the user's real-time motion speed, a spatiotemporal joint optimization algorithm is used to generate multi-level obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels. S105: The obstacle avoidance path, heat map of the dangerous area and semantic labels of the obstacles are dynamically superimposed on the user's field of view through AR display. At the same time, a directional voice warning signal is generated through the spatial audio algorithm to complete the environmental perception of the smart glasses.
[0005] Furthermore, in S101, the real-time collection of environmental stereo image data by a binocular camera includes: Receiving a connection request sent by a user terminal through a preset network, wherein the connection request is used to request to establish a connection with the VR wearable device; Detecting whether the current account of the user terminal is the target account; If the current account of the user terminal is the target account, a connection is made with the user terminal according to the connection request. After the connection is completed, a request is sent by the user terminal to collect environmental stereo image data in real time through a binocular camera.
[0006] Furthermore, multimodal sensor data is synchronously acquired, and the multimodal sensor data includes: IMU inertial measurement unit data, which acquires the acceleration and angular velocity of the smart glasses in real time within the interval between binocular camera frames. Motion blur is compensated through IMU data interpolation, and the device's posture change is calculated. Combined with accelerometer data, the direction of the gravity vector is determined to assist in the vertical alignment of the 3D map. Barometer data calculates relative altitude through air pressure changes, detects the user's current state, and provides floor-level coarse positioning in the SLAM system, complementing the visual positioning results; Millimeter-wave radar data provides high-precision distance measurement in scenarios where binocular vision fails, identifies the speed of moving obstacles through the Doppler effect, and detects non-line-of-sight obstacles.
[0007] Furthermore, the preset network includes one or a combination of 3G network, 4G network, 5G network, and WIFI network.
[0008] Furthermore, in S103, a tightly coupled visual-inertial SLAM algorithm is used to construct a dynamic three-dimensional environment map based on the fused multi-source data, including: The tightly coupled visual-inertial SLAM algorithm adopts a factor graph optimization framework to jointly construct an objective function by combining binocular feature point reprojection error, IMU pre-integration error and radar point cloud matching error; The optimal pose of the objective function is solved iteratively through the LM algorithm, and the sliding window management mechanism is used to limit the computational complexity; Complete multi-source error joint optimization, sliding window complexity control and adaptive weight adjustment to achieve a balance between centimeter-level positioning accuracy, real-time performance and robustness in dynamic environments.
[0009] Furthermore, in S104, a lightweight semantic segmentation network is used to classify the three-dimensional environment map pixel by pixel to identify the category, geometric size and motion state of obstacles, including: The lightweight semantic segmentation network uses MobileNetV3 as the encoder, and the decoder is composed of cascaded dilated convolution modules; The cascaded dilated convolution module embeds a binocular disparity attention module in the skip connection. The module generates attention weights by calculating the cosine similarity of the left and right eye feature maps to suppress occlusion area errors in depth estimation. The segmentation network fuses the left and right eye feature maps through a binocular disparity attention module and introduces radar point cloud data as deep supervision.
[0010] Furthermore, based on the semantic segmentation results and the user's real-time movement speed, a spatiotemporal joint optimization algorithm is used to generate multi-level obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels, including: A multi-level obstacle avoidance path is generated through a spatiotemporal joint optimization algorithm, in which a trajectory prediction model is used to calculate the collision risk probability for moving obstacles; The trajectory prediction model adopts a social-LSTM network to extract the historical motion trajectory and scene context features of dynamic obstacles; Predict the motion path within the next 3 seconds and generate a confidence score for the dynamic obstacle avoidance path based on the safety distance threshold.
[0011] Furthermore, in S105, the obstacle avoidance path, the heat map of the dangerous area, and the semantic label of the obstacle are dynamically superimposed on the user's field of view through AR display, including: Obtain obstacle avoidance paths, heat maps of dangerous areas, and obstacle semantic labels, and record coordinate positioning of these paths, heat maps of dangerous areas, and obstacle semantic labels on a three-dimensional environment map; Obtain the outer contours of the obstacle avoidance path, the heat map of the dangerous area, and the semantic labels of the obstacles, and complete the copying of the outer contours of the obstacle avoidance path, the heat map of the dangerous area, and the semantic labels of the obstacles; The stacking in the three-dimensional environment map is completed according to the specific outline of the copied data. After the stacking is completed, the coordinate points are identified and shifted to complete the dynamic superposition display.
[0012] Compared with the existing technology, the machine vision-based smart glasses environment perception method and system provided by the present invention have the following beneficial effects: 1. Based on a multi-sensor tightly coupled factor graph optimization framework, including binoculars, IMUs, and millimeter-wave radar, it achieves horizontal positioning error ≤1.2cm and height error ≤2cm in dynamic environments. Through a motion clustering algorithm and a social-LSTM trajectory prediction model, it can accurately distinguish between static and dynamic targets and predict the average displacement error of the trajectory for the next three seconds to be ≤0.3m. The millimeter-wave radar and binocular parallax attention mechanism work together to achieve a depth estimation error of ≤5cm in harsh conditions such as low light, rain, and fog. 2. The FPGA+NPU heterogeneous architecture and sliding window optimization mechanism achieve end-to-end processing latency of ≤35ms, meeting the real-time requirements of mobile scenarios. The semantic segmentation network based on MobileNetV3 is compressed to 7.8MB, supports INT8 quantized inference, and achieves a processing speed of 28FPS on the NPU. Through DVFS technology and adaptive resource scheduling strategies, the system average power consumption is <1.2W, and the battery life is extended to 10 hours, making it suitable for all-weather industrial inspection applications. 3. Holographic grating waveguide technology achieves 85% transmittance and minimal error in viewing distance consistency, seamlessly integrating virtual information with real scenes and reducing user visual fatigue by 50%. The spatial audio algorithm is linked to the haptic vibration module, ensuring low latency in warning information response time. Furthermore, it supports millimeter-wave radar array upgrades, extending the detection range to 50 meters, making it suitable for large-scale scenarios such as drone collaborative operations. 4. Through multi-sensor conflict detection and adaptive weight adjustment, the system's positioning accuracy remains low even when 30% of visual features are blocked or interfered with. The millimeter-wave radar has a false detection rate of less than 5% in rain, fog, and dusty environments. It also supports ROS / Android system interfaces and can be quickly integrated into industrial robots, AGVs, and other equipment.
[0013] A machine vision-based smart glasses environment perception system, which executes the above-mentioned smart glasses environment perception method, includes: Multimodal sensing module: includes binocular cameras, IMU, millimeter-wave radar, and barometer. It is used to acquire multimodal sensor data and store it in a cache queue in a timestamp-aligned manner. The hardware synchronization error of each sensor is ≤1ms. Edge computing unit: Integrates a SLAM coprocessor, NPU accelerator, and dynamic memory allocation controller, supporting multi-threaded parallel processing of sensor data streams. It is used to generate pose increments from multimodal sensor data through a pre-integration algorithm, spatially align and fuse it with obstacle point cloud data from millimeter-wave radar, and use a tightly coupled visual-inertial SLAM algorithm to build a dynamic 3D environment map based on the fused multi-source data. Semantic Understanding Module: This module stores pre-trained lightweight semantic segmentation models and an obstacle knowledge base, supports online incremental learning, and is used to perform pixel-by-pixel classification of 3D environment maps, identify obstacle categories, geometric dimensions, and motion states, and generate multi-level obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels using a spatiotemporal joint optimization algorithm based on the semantic segmentation results and the user's real-time motion speed. Interactive output module: This module includes an AR waveguide display, bone conduction headphones, and a tactile feedback device. It supports multi-channel human-computer interaction and is used to dynamically overlay obstacle avoidance paths, heat maps of dangerous areas, and semantic labels of obstacles onto the user's field of view through AR displays. It also generates directional voice warning signals through spatial audio algorithms. Energy management module: used to supply energy to the system, using dynamic voltage and frequency scaling technology to adjust processor power consumption according to computing load.
[0014] Furthermore, the edge computing unit includes: A SLAM coprocessor, configured with dedicated visual-inertial SLAM logic circuits, is used to perform dedistortion and optical flow compensation preprocessing on stereo image data, and to construct a dynamic 3D environment map based on the fused multi-source data. The static background of the 3D environment map is modeled using a voxel grid, and dynamic obstacles are marked as independent moving targets through motion clustering. The NPU accelerator supports INT8 quantized inference, generates pose increments from multimodal sensor data through a pre-integration algorithm, and spatially aligns and fuses it with the obstacle point cloud data from the millimeter-wave radar. A dynamic memory allocation controller manages the memory usage of binocular images, radar point clouds, and IMU data through priority queues. It is used to generate pose increments through a pre-integration algorithm and spatially align and fuse them with the obstacle point cloud data of the millimeter-wave radar. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a schematic flow chart of a method for environmental perception of smart glasses based on machine vision proposed in the present invention; Figure 2 This is a schematic block diagram of the process of real-time acquisition of environmental stereoscopic image data by a binocular camera in a machine vision-based smart glasses environment perception method proposed by the present invention; Figure 3This is a schematic diagram of the structure of a machine vision-based smart glasses environment perception system proposed by the present invention; Figure 4 This is a structural diagram of the edge computing unit in the machine vision-based smart glasses environment perception system proposed by the present invention. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0017] The implementation of the present invention is described in detail below with reference to specific embodiments.
[0018] The same or similar numbers in the drawings of this embodiment correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "up", "down", "left", "right", etc. indicate directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0019] Reference Figure 1-2 As shown in FIG, a method for intelligent glasses environment perception based on machine vision is applied to VR wearable devices, specifically comprising the following steps: S101: Collecting environmental stereo image data in real time through a binocular camera, synchronously acquiring multimodal sensor data, and storing them in a cache queue in a timestamp-aligned manner; Among them, the real-time collection of environmental stereo image data through the binocular camera includes: Receive a connection request sent by a user terminal through a preset network, where the connection request is used to request to establish a connection with the VR wearable device; Check whether the current account of the user terminal is the target account; If the current account of the user terminal is the target account, a connection is made with the user terminal according to the connection request. After the connection is completed, a request is sent by the user terminal to collect stereoscopic image data of the environment in real time through the binocular camera. S102: Dedistortion and optical flow compensation are pre-processed on the stereo image data. At the same time, the multimodal sensor data is used to generate pose increments through a pre-integration algorithm, and spatially aligned and fused with the obstacle point cloud data from the millimeter-wave radar. S103: A tightly coupled visual-inertial SLAM algorithm is used to construct a dynamic 3D environment map based on the fused multi-source data. The static background of the 3D environment map is modeled using a voxel grid, and dynamic obstacles are marked as independent moving targets through motion clustering. Among them, a tightly coupled visual-inertial SLAM algorithm is used to build a dynamic three-dimensional environment map based on the fused multi-source data, including: The tightly coupled visual-inertial SLAM algorithm adopts a factor graph optimization framework to jointly construct the objective function by combining the binocular feature point reprojection error, IMU pre-integration error and radar point cloud matching error; The optimal pose of the objective function is solved iteratively through the Levenberg-Marquardt algorithm, and a sliding window management mechanism is used to limit computational complexity. The FPGA+NPU heterogeneous architecture and sliding window optimization mechanism achieve end-to-end processing latency of ≤35ms, meeting the real-time requirements of mobile scenarios. The semantic segmentation network based on MobileNetV3 is compressed to 7.8MB, supports INT8 quantized inference, and achieves a processing speed of 28FPS on the NPU. Through DVFS technology and adaptive resource scheduling strategies, the system's average power consumption is <1.2W, and battery life is extended to 10 hours, making it suitable for all-weather industrial inspection applications. Complete multi-source error joint optimization, sliding window complexity control and adaptive weight adjustment to achieve a balance between centimeter-level positioning accuracy, real-time performance and robustness in dynamic environments; S104: A lightweight semantic segmentation network is used to perform pixel-by-pixel classification of the 3D environment map to identify the type, geometry, and motion state of obstacles. Based on the semantic segmentation results and the user's real-time motion speed, a spatiotemporal joint optimization algorithm is used to generate multi-level obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels. Among them, a lightweight semantic segmentation network is used to classify the 3D environment map pixel by pixel to identify the category, geometric size and motion state of obstacles, including: The lightweight semantic segmentation network uses MobileNetV3 as the encoder, and the decoder consists of cascaded dilated convolution modules; The cascaded dilated convolution module embeds a binocular disparity attention module in the skip connection. The module generates attention weights by calculating the cosine similarity of the left and right eye feature maps to suppress occlusion area errors in depth estimation. The segmentation network fuses the left and right eye feature maps through the binocular disparity attention module and introduces radar point cloud data as deep supervision; S105: The obstacle avoidance path, heat map of the dangerous area, and semantic labels of obstacles are dynamically superimposed on the user's field of view through AR display. At the same time, a directional voice warning signal is generated through a spatial audio algorithm to complete the smart glasses' environmental perception. Among them, obstacle avoidance paths, heat maps of dangerous areas, and obstacle semantic labels are dynamically superimposed on the user's field of view through AR display, including: Obtain obstacle avoidance paths, heat maps of dangerous areas, and obstacle semantic labels, and record coordinate positioning of these paths, heat maps of dangerous areas, and obstacle semantic labels on a three-dimensional environment map; Obtain the outer contours of the obstacle avoidance path, the heat map of the dangerous area, and the semantic labels of the obstacles, and complete the copying of the outer contours of the obstacle avoidance path, the heat map of the dangerous area, and the semantic labels of the obstacles; The stacking in the three-dimensional environment map is completed according to the specific outline of the copied data. After the stacking is completed, the coordinate point identification and shifting are performed to complete the dynamic superposition of the display. The holographic grating waveguide technology achieves 85% transmittance and small error in line of sight consistency. Virtual information and real scenes are seamlessly integrated, and user visual fatigue is reduced by 50%. The spatial audio algorithm is linked with the tactile vibration module, and the alarm information response time delay is low. It also supports millimeter wave radar array upgrades, and the detection distance is extended to 50 meters, which is suitable for large-scale scenarios such as drone collaborative operations.
[0020] In S101 of this embodiment, multimodal sensor data is synchronously acquired, and the multimodal sensor data includes: IMU inertial measurement unit data, which acquires the acceleration and angular velocity of the smart glasses in real time within the interval between binocular camera frames. Motion blur is compensated through IMU data interpolation, and the device's posture change is calculated. Combined with accelerometer data, the direction of the gravity vector is determined to assist in the vertical alignment of the 3D map. Barometer data calculates relative altitude through air pressure changes, detects the user's current state, and provides floor-level coarse positioning in the SLAM system, complementing the visual positioning results; Millimeter-wave radar data provides high-precision distance measurement in scenarios where binocular vision fails, identifies the speed of moving obstacles through the Doppler effect, and detects non-line-of-sight obstacles.
[0021] In this embodiment, the preset network includes one of a 3G network, a 4G network, a 5G network, and a WIFI network, or a combination thereof.
[0022] In this embodiment, based on the semantic segmentation results and the user's real-time movement speed, a spatiotemporal joint optimization algorithm is used to generate a multi-level obstacle avoidance path, a heat map of the dangerous area, and obstacle semantic labels, including: A multi-level obstacle avoidance path is generated through a spatiotemporal joint optimization algorithm, in which a trajectory prediction model is used to calculate the collision risk probability for moving obstacles; The trajectory prediction model uses a social-LSTM network to extract the historical motion trajectory and scene context features of dynamic obstacles; Predict the motion path within the next 3 seconds and generate a confidence score for the dynamic obstacle avoidance path based on the safety distance threshold.
[0023] This technical solution is based on a multi-sensor tightly coupled factor graph optimization framework, including binoculars, IMUs, and millimeter-wave radar. It achieves horizontal positioning error ≤1.2cm and height error ≤2cm in dynamic environments. Through a motion clustering algorithm and a social-LSTM trajectory prediction model, it can accurately distinguish between static and dynamic targets and predict the average displacement error of the trajectory for the next three seconds to be ≤0.3m. The millimeter-wave radar and binocular parallax attention mechanism work together to achieve a depth estimation error of ≤5cm in adverse conditions such as low light, rain, and fog. Reference Figure 3-4 As shown, a smart glasses environment perception system based on machine vision executes the above-mentioned smart glasses environment perception method, and the environment perception system includes: a multimodal sensing module: including a binocular camera, an IMU, a millimeter-wave radar, and a barometer, for acquiring multimodal sensor data and storing it in a cache queue in a timestamp-aligned manner, and the hardware synchronization error of each sensor is ≤1ms; an edge computing unit: integrating a SLAM coprocessor, an NPU accelerator, and a dynamic memory allocation controller, supporting multi-threaded parallel processing of sensor data streams, for generating pose increments from multimodal sensor data through a pre-integration algorithm, spatially aligning and fusing the obstacle point cloud data of the millimeter-wave radar, and using a tightly coupled visual-inertial SLAM algorithm to build a dynamic three-dimensional environment map based on the fused multi-source data; a semantic understanding module: storing a pre-trained lightweight semantic segmentation model and an obstacle knowledge base, supporting online incremental learning, for pixel-by-pixel classification of the three-dimensional environment map, identifying the category, geometric size, and motion state of the obstacle, and performing real-time object recognition based on the semantic segmentation results and the user's real-time operation. The system uses a spatiotemporal joint optimization algorithm to generate multi-level obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels. The interactive output module includes an AR waveguide display, bone conduction headphones, and a tactile feedback device, supporting multi-channel human-computer interaction. It is used to dynamically overlay obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels onto the user's field of view through the AR display, while generating directional voice warning signals through a spatial audio algorithm. The energy management module is used to supply energy to the system. It uses dynamic voltage and frequency scaling technology to adjust processor power consumption according to the computational load. Based on a multi-sensor tightly coupled factor graph optimization framework, including binoculars + IMU + millimeter-wave radar, it achieves a horizontal positioning error of ≤1.2cm and a height error of ≤2cm in dynamic environments. The motion clustering algorithm and the social-LSTM trajectory prediction model can accurately distinguish between static and dynamic targets, and predict the average displacement error of the trajectory in the next 3 seconds is ≤0.3m. The millimeter-wave radar and the binocular parallax attention mechanism work together to achieve a depth estimation error of ≤5cm in harsh conditions such as low light, rain, and fog.
[0024] In this embodiment, the edge computing unit includes: a SLAM coprocessor configured with a dedicated visual-inertial SLAM logic circuit for performing dedistortion and optical flow compensation preprocessing on stereo image data, and for constructing a dynamic three-dimensional environment map based on the fused multi-source data, wherein the static background of the three-dimensional environment map is modeled using a voxel grid, and dynamic obstacles are marked as independent moving targets through motion clustering; an NPU accelerator supporting INT8 quantized inference, generating pose increments from multimodal sensor data through a pre-integration algorithm, and spatially aligning and fusing the obstacle point cloud data from the millimeter-wave radar; A dynamic memory allocation controller manages the memory usage of binocular images, radar point clouds, and IMU data through priority queues. This is used to generate pose increments using a pre-integration algorithm and spatially align and fuse them with obstacle point cloud data from the millimeter-wave radar. The system utilizes an FPGA + NPU heterogeneous architecture and a sliding window optimization mechanism to achieve end-to-end processing latency of ≤35ms, meeting the real-time requirements of mobile scenarios. The semantic segmentation network based on MobileNetV3 is compressed to 7.8MB, supports INT8 quantized inference, and achieves a processing speed of 28FPS on the NPU. Through DVFS technology and an adaptive resource scheduling strategy, the system's average power consumption is <1.2W, extending battery life to 10 hours, making it suitable for all-weather industrial inspection applications. Holographic grating waveguide technology achieves 85% transmittance and minimal line-of-sight consistency error, seamlessly integrating virtual information with real scenes and reducing user visual fatigue by 50%. The spatial audio algorithm is linked to the haptic vibration module, ensuring low latency in alarm response times. Furthermore, support for millimeter-wave radar array upgrades extends the detection range to 50 meters, making it suitable for large-scale scenarios such as drone collaborative operations. This technical solution uses multi-sensor conflict detection and adaptive weight adjustment. Even when 30% of visual features are blocked or interfered with, the system positioning accuracy still maintains a small error. The millimeter-wave radar has a false detection rate of less than 5% in rainy, foggy, and dusty environments. It also supports ROS / Android system interfaces and can be quickly integrated into industrial robots, AGVs and other equipment.
[0025] In this embodiment, the entire operation process can be controlled by a computer to provide signal feedback to implement the steps in sequence. These are all conventional knowledge of current automated control and will not be described in detail in this embodiment.
[0026] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for smart glasses environment perception based on machine vision, characterized in that: Applied to VR wearable devices, specifically including the following steps: S101: Collecting environmental stereo image data in real time through a binocular camera, synchronously acquiring multimodal sensor data, and storing them in a cache queue in a timestamp-aligned manner; S102: performing dedistortion and optical flow compensation preprocessing on the stereo image data, and generating pose increments from the multimodal sensor data through a pre-integration algorithm, and performing spatial alignment and fusion with the obstacle point cloud data from the millimeter-wave radar; S103: A tightly coupled visual-inertial SLAM algorithm is used to construct a dynamic 3D environment map based on the fused multi-source data. The static background of the 3D environment map is modeled using a voxel grid, and dynamic obstacles are marked as independent moving targets through motion clustering. S104: A lightweight semantic segmentation network is used to perform pixel-by-pixel classification of the 3D environment map to identify the type, geometry, and motion state of obstacles. Based on the semantic segmentation results and the user's real-time motion speed, a spatiotemporal joint optimization algorithm is used to generate multi-level obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels. S105: The obstacle avoidance path, heat map of the dangerous area and semantic labels of the obstacles are dynamically superimposed on the user's field of view through AR display. At the same time, a directional voice warning signal is generated through the spatial audio algorithm to complete the environmental perception of the smart glasses.
2. The method for intelligent glasses environment perception based on machine vision according to claim 1, characterized in that: In S101, the binocular camera collects the environmental stereo image data in real time, including: Receiving a connection request sent by a user terminal through a preset network, wherein the connection request is used to request to establish a connection with the VR wearable device; Detecting whether the current account of the user terminal is the target account; If the current account of the user terminal is the target account, a connection is made with the user terminal according to the connection request. After the connection is completed, a request is sent by the user terminal to collect environmental stereo image data in real time through a binocular camera.
3. The method for intelligent glasses environment perception based on machine vision according to claim 2, characterized in that: Synchronously acquire multimodal sensor data, the multimodal sensor data including: IMU inertial measurement unit data, which acquires the acceleration and angular velocity of the smart glasses in real time within the interval between binocular camera frames. Motion blur is compensated through IMU data interpolation, and the device's posture change is calculated. Combined with accelerometer data, the direction of the gravity vector is determined to assist in the vertical alignment of the 3D map. Barometer data calculates relative altitude through air pressure changes, detects the user's current state, and provides floor-level coarse positioning in the SLAM system, complementing the visual positioning results; Millimeter-wave radar data provides high-precision distance measurement in scenarios where binocular vision fails, identifies the speed of moving obstacles through the Doppler effect, and detects non-line-of-sight obstacles.
4. The method for intelligent glasses environment perception based on machine vision according to claim 3, characterized in that: The preset network includes one of a 3G network, a 4G network, a 5G network, a WIFI network, or a combination thereof.
5. The method for intelligent glasses environment perception based on machine vision according to claim 4, characterized in that: In S103, a tightly coupled visual-inertial SLAM algorithm is used to construct a dynamic three-dimensional environment map based on the fused multi-source data, including: The tightly coupled visual-inertial SLAM algorithm adopts a factor graph optimization framework to jointly construct an objective function by combining binocular feature point reprojection error, IMU pre-integration error and radar point cloud matching error; The optimal pose of the objective function is solved iteratively through the LM algorithm, and the sliding window management mechanism is used to limit the computational complexity; Complete multi-source error joint optimization, sliding window complexity control and adaptive weight adjustment to achieve a balance between centimeter-level positioning accuracy, real-time performance and robustness in dynamic environments.
6. The method for intelligent glasses environment perception based on machine vision according to claim 5, characterized in that: In S104, a lightweight semantic segmentation network is used to classify the 3D environment map pixel by pixel to identify the category, geometric size, and motion state of obstacles, including: The lightweight semantic segmentation network uses MobileNetV3 as the encoder, and the decoder is composed of cascaded dilated convolution modules; The cascaded dilated convolution module embeds a binocular disparity attention module in the skip connection. The module generates attention weights by calculating the cosine similarity of the left and right eye feature maps to suppress occlusion area errors in depth estimation. The segmentation network fuses the left and right eye feature maps through a binocular disparity attention module and introduces radar point cloud data as deep supervision.
7. The method for intelligent glasses environment perception based on machine vision according to claim 6, characterized in that: Based on the semantic segmentation results and the user's real-time movement speed, a spatiotemporal joint optimization algorithm is used to generate multi-level obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels, including: A multi-level obstacle avoidance path is generated through a spatiotemporal joint optimization algorithm, in which a trajectory prediction model is used to calculate the collision risk probability for moving obstacles; The trajectory prediction model adopts a social-LSTM network to extract the historical motion trajectory and scene context features of dynamic obstacles; Predict the motion path within the next 3 seconds and generate a confidence score for the dynamic obstacle avoidance path based on the safety distance threshold.
8. The method for intelligent glasses environment perception based on machine vision according to claim 7, characterized in that: In S105, the obstacle avoidance path, the heat map of the dangerous area, and the semantic labels of the obstacles are dynamically superimposed on the user's field of view through AR display, including: Obtain obstacle avoidance paths, heat maps of dangerous areas, and obstacle semantic labels, and record coordinate positioning of these paths, heat maps of dangerous areas, and obstacle semantic labels on a three-dimensional environment map; Obtain the outer contours of the obstacle avoidance path, the heat map of the dangerous area, and the semantic labels of the obstacles, and complete the copying of the outer contours of the obstacle avoidance path, the heat map of the dangerous area, and the semantic labels of the obstacles; The stacking in the three-dimensional environment map is completed according to the specific outline of the copied data. After the stacking is completed, the coordinate points are identified and shifted to complete the dynamic superposition display.
9. A machine vision-based smart glasses environment perception system, characterized in that: The smart glasses environment perception method according to any one of claims 1 to 8 is implemented, wherein the environment perception system comprises: Multimodal sensing module: includes binocular cameras, IMU, millimeter-wave radar, and barometer. It is used to acquire multimodal sensor data and store it in a cache queue in a timestamp-aligned manner. The hardware synchronization error of each sensor is ≤1ms. Edge computing unit: Integrates a SLAM coprocessor, NPU accelerator, and dynamic memory allocation controller, supporting multi-threaded parallel processing of sensor data streams. It is used to generate pose increments from multimodal sensor data through a pre-integration algorithm, spatially align and fuse it with obstacle point cloud data from millimeter-wave radar, and use a tightly coupled visual-inertial SLAM algorithm to build a dynamic 3D environment map based on the fused multi-source data. Semantic Understanding Module: This module stores pre-trained lightweight semantic segmentation models and an obstacle knowledge base, supports online incremental learning, and is used to perform pixel-by-pixel classification of 3D environment maps, identify obstacle categories, geometric dimensions, and motion states, and generate multi-level obstacle avoidance paths, dangerous area heat maps, and obstacle semantic labels using a spatiotemporal joint optimization algorithm based on the semantic segmentation results and the user's real-time motion speed. Interactive output module: This module includes an AR waveguide display, bone conduction headphones, and a tactile feedback device. It supports multi-channel human-computer interaction and is used to dynamically overlay obstacle avoidance paths, heat maps of dangerous areas, and semantic labels of obstacles onto the user's field of view through AR displays. It also generates directional voice warning signals through spatial audio algorithms. Energy management module: used to supply energy to the system, using dynamic voltage and frequency scaling technology to adjust processor power consumption according to computing load.
10. The machine vision-based smart glasses environment perception system according to claim 9, characterized in that: The edge computing unit includes: A SLAM coprocessor, configured with dedicated visual-inertial SLAM logic circuits, is used to perform dedistortion and optical flow compensation preprocessing on stereo image data, and to construct a dynamic 3D environment map based on the fused multi-source data. The static background of the 3D environment map is modeled using a voxel grid, and dynamic obstacles are marked as independent moving targets through motion clustering. The NPU accelerator supports INT8 quantized inference, generates pose increments from multimodal sensor data through a pre-integration algorithm, and spatially aligns and fuses it with the obstacle point cloud data from the millimeter-wave radar. A dynamic memory allocation controller manages the memory usage of binocular images, radar point clouds, and IMU data through priority queues. It is used to generate pose increments through a pre-integration algorithm and spatially align and fuse them with the obstacle point cloud data of the millimeter-wave radar.