A guide method and system based on monocular depth estimation

CN122768049APending Publication Date: 2026-09-18XIAMEN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611273497.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提出一种一种基于单目深度估计的导盲方法与系统,以解决现有导盲方案中深度感知硬件成本高、地形适应性差、障碍物检测不全面及头部威胁缺失等问题

Benefits of technology

[0077] (1) Low hardware cost and lightweight device: This invention only requires a monocular camera and an IMU sensor, without the need for additional depth sensing hardware such as binocular cameras, structured light modules or ToF sensors. It can be integrated into a wearable device in the form of ordinary glasses, which is lightweight and has low power consumption, making it suitable for visually impaired users to wear for a long time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122768049A_ABST
    Figure CN122768049A_ABST
Patent Text Reader

Abstract

This invention proposes a method and system for guiding the visually impaired based on monocular depth estimation, belonging to the fields of computer vision and smart wearable technology. The invention uses a monocular camera to acquire images, which, after quality assessment, are input into a lightweight depth estimation model for depth estimation. Post-processing then outputs an inverted normalized depth map. A dynamic column-level ground reference is constructed using bottom reference strips, and the region of interest is divided into left, center, and right sections and analyzed in a grid pattern to identify various obstacle shapes, including blocky, elongated, grid-like, and wall-like obstacles. Simultaneously, the IMU's pitch angle is adaptively adjusted to adjust the top region of the image for head-mounted obstacles and suspended objects detection, and voice commands are generated based on distance classification and priority arbitration. The system is integrated into a glasses-like wearable device. This invention features lightweight design, terrain adaptation, posture perception, and a low false alarm rate, effectively assisting visually impaired individuals in safe travel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information technology, specifically to the fields of computer vision, smart wearables and guide technology, and discloses a guide method and system based on monocular depth estimation. Background Technology

[0002] Visual impairment severely impacts the daily safety and quality of life of patients. Visually impaired individuals face significant challenges when traveling independently, including the inability to perceive the distance and shape of obstacles, judge the movement of dynamic targets, and identify overhead obstacles, severely limiting their ability to travel independently. Therefore, developing intelligent assistive systems capable of real-time environmental perception and providing precise navigation guidance is of significant social importance.

[0003] Currently available guide technologies for the visually impaired have the following limitations: solutions based on ultrasonic or infrared sensors can only obtain a single distance value, cannot distinguish the shape of obstacles, and lack the ability to estimate the dynamic target collision time; solutions based on wireless sensor networks require the deployment of nodes in multiple locations on the user's body, resulting in complex wearing conditions, high communication latency, and difficulty in meeting real-time obstacle avoidance requirements; solutions based on traditional computer vision rely on two-dimensional image processing, lack depth information, cannot accurately estimate obstacle distances, and are prone to generating false alarms; solutions based on binocular or structured light depth sensors can directly obtain three-dimensional information, but the modules are large, consume a lot of power, and their performance degrades under strong outdoor light, making them unsuitable for integration into lightweight glasses-like devices.

[0004] In summary, existing navigation technologies for the visually impaired generally suffer from inaccurate distance estimation and insufficient obstacle shape recognition. Therefore, there is an urgent need for a lightweight depth estimation method based on a monocular camera, capable of comprehensively perceiving the distance, shape, and spatial position of obstacles without adding additional depth sensors, and combining this with a multi-priority decision fusion mechanism to provide safe and reliable navigation guidance for visually impaired users. Summary of the Invention

[0005] The purpose of this invention is to propose a guide method and system based on monocular depth estimation to solve the problems of high cost of depth sensing hardware, poor terrain adaptability, incomplete obstacle detection and lack of head threat detection in existing guide solutions.

[0006] To achieve the above objectives, the technical solution of the present invention is as follows:

[0007] A guide system for the blind based on monocular depth estimation includes an image acquisition module, a posture perception module, a main control processor, a voice output module, and a power supply subsystem integrated into a head-worn device.

[0008] The image acquisition module is used to acquire RGB image streams of the user's travel path in real time;

[0009] The attitude perception module includes an inertial measurement unit, which is used to acquire three-axis acceleration data and three-axis angular velocity data of the user's head in real time;

[0010] The main control processor is configured as follows:

[0011] The user's head posture information is obtained based on the three-axis acceleration data and three-axis angular velocity data. The user's head is then judged to be in an abnormal motion state based on the posture information. If so, the voice output module is paused. Otherwise, the quality of the current frame RGB image is evaluated.

[0012] If the quality assessment result of the current frame RGB image does not meet the preset quality requirements, the obstacle detection result of the previous valid frame will be used; otherwise, the depth of the current frame RGB image will be estimated by a monocular depth estimation model, and the original depth map output by the model will be post-processed to obtain a normalized depth map.

[0013] The user's walking motion is identified based on triaxial acceleration data to estimate single-step displacement. The scale factor is then calculated by combining the corner displacement information of the ground area and the normalized depth value of the corner, so as to convert the normalized depth value in the normalized depth map into physical distance.

[0014] Based on the pose information, determine the region of interest and the upper detection region of the normalized depth map; within the region of interest, identify ground obstacles based on the normalized depth value; within the upper detection region, identify obstacles above the head based on physical distance.

[0015] Sliding window decision fusion is performed on obstacle detection results from multiple consecutive frames to generate voice navigation commands;

[0016] The voice output module is used to broadcast voice navigation instructions to the user;

[0017] The power subsystem is used to supply power to the guide system.

[0018] Preferably, the user's head posture information is obtained based on triaxial acceleration data and triaxial angular velocity data, and the user's head is judged to be in an abnormal movement state based on the posture information. If so, the voice output module's broadcast is paused, as follows:

[0019] The triaxial acceleration data output by the inertial measurement unit is low-pass filtered, and the pitch angle is calculated based on the filtered triaxial acceleration data. and roll angle :

[0020]

[0021]

[0022]

[0023] in, This is the filtered gravity vector at the current moment. This is the filtered gravity vector from the previous time step. The data represents the triaxial acceleration output by the inertial measurement unit at the current moment. These are the low-pass filter coefficients for the gravity vector; , , These are the components of the filtered gravity vector along the three axes of the inertial measurement unit at the current moment.

[0024] A dead zone threshold is set for the yaw axis angular velocity output by the inertial measurement unit. If the absolute value of the current yaw axis angular velocity is less than or equal to the dead zone threshold, the yaw axis angular velocity increment at the current moment is ignored, and the yaw angle at the previous moment is kept as the yaw angle at the current moment. If the absolute value of the current yaw axis angular velocity is greater than the dead zone threshold, the product of the current yaw axis angular velocity and the sampling interval is accumulated and added to the yaw angle at the previous moment. A leakage coefficient is applied to the accumulated yaw angle to attenuate it, and the yaw angle at the current moment is obtained.

[0025] The pitch, roll, and yaw angles are smoothed using exponential moving averages, and user status is determined accordingly.

[0026] When the smoothed pitch angle is greater than the first preset threshold, it is determined to be a nose-down state; when the smoothed pitch angle is less than the second preset threshold, it is determined to be a nose-up state; when the rate of change of the smoothed yaw angle is greater than the third preset threshold, it is determined to be a rapid head-turn state.

[0027] The voice broadcast will pause when the user is looking down, looking up, or turning their head quickly.

[0028] Preferably, the quality assessment of the current frame's RGB image is performed as follows:

[0029] After converting the RGB image of the current frame to grayscale, the variance of the Laplacian operator response is calculated. When the variance is lower than a preset threshold, the current frame is determined to be a blurry frame. The depth estimation of the current frame is skipped and the obstacle detection result of the previous valid frame is used.

[0030] If a preset number of consecutive frames are all judged as blurry frames, the broadcast frequency will be reduced.

[0031] Preferably, the post-processing of the original depth map output by the model to obtain a normalized depth map specifically includes:

[0032] Extreme value suppression is performed on the original depth map: the lower quantile and higher quantile of the depth values ​​of the entire original depth map are calculated, pixel values ​​less than the lower quantile are clamped to the lower quantile, and pixel values ​​greater than the higher quantile are clamped to the higher quantile.

[0033] Contrast stretching is performed on the depth map after extreme value suppression to map the depth values ​​to a predetermined range;

[0034] The mapped depth map is inverted so that the depth values ​​of nearby objects are less than the depth values ​​of distant objects, resulting in a normalized depth map.

[0035] Untrusted regions are detected in the normalized depth image based on the proportion of abnormal connected regions with abrupt changes in depth value, and the corresponding grid blocks of untrusted regions are excluded in the subsequent grid block analysis.

[0036] Preferably, the step of identifying the user's walking motion based on triaxial acceleration data to estimate single-step displacement, and solving for the scale factor by combining the corner displacement information of the ground area and the normalized depth value of the corner, specifically includes:

[0037] Based on triaxial acceleration data Perform zero-velocity detection to identify the user's gait cycle: calculate acceleration amplitude. , , These are the triaxial accelerations output by the inertial measurement unit, i.e. When the acceleration amplitude When the acceleration amplitude falls back to the preset amplitude range after reaching the peak and the rate of change of the acceleration amplitude is less than the preset rate of change threshold, it is marked as a zero velocity point. The interval between adjacent zero velocity points is one gait cycle.

[0038] Horizontal acceleration during gait period Perform trapezoidal integration to obtain the horizontal displacement components. ,in Triaxial acceleration data The acceleration component in the horizontal direction after deducting the gravitational component;

[0039] Corner points are extracted from the bottom ground region of the previous RGB image. The horizontal parallax of each corner point is obtained in the current RGB image through optical flow tracing. The corresponding normalized depth value is obtained from the normalized depth map of the current frame based on the image coordinates of each corner point.

[0040] According to the horizontal displacement component A linear equation is established based on the horizontal parallax and the normalized depth value, and the scale factor is solved using the least squares method. :

[0041]

[0042] in: This refers to the focal length of the camera in the horizontal direction. For the first The normalized depth values ​​corresponding to each corner point For the first Horizontal parallax of each corner point This represents the total number of corner points;

[0043] The obtained scale factor is smoothed by exponential moving average to obtain a smoothed scale factor. When the preset pause condition is met, the update of the smoothed scale factor is paused and the historical value is used. When the rate of change of the smoothed scale factor exceeds the preset rate of change threshold, the scene is switched, the historical cumulative value of the exponential moving average smoothing is reset, and the scale factor obtained in the current frame is used as the new historical benchmark.

[0044] Preferably, identifying ground obstacles based on normalized depth values ​​within the region of interest specifically includes:

[0045] Based on the smoothed pitch angle, the exclusion ratio of the invalid region at the top of the normalized depth map is calculated to eliminate non-ground areas and obtain the region of interest; the exclusion ratio is negatively correlated with the smoothed pitch angle, and the exclusion ratio is limited between a preset upper limit value and a preset lower limit value.

[0046] The region of interest is divided into left, middle and right sections according to a preset ratio, with the width of the middle section being greater than that of the left and right sections;

[0047] Based on the smoothed roll angle, calculate the horizontal offset of each row in the region of interest. Then, translate the partition boundary points on the vertical dividing lines of the left, middle and right zones according to the horizontal offset of their respective rows. Connect the translated partition boundary points to form the corrected partition boundary lines, so that the partition boundaries are aligned with the real horizontal plane.

[0048] The left, middle and right areas are gridded, and the average depth of each grid block is calculated based on the normalized depth value. If the average depth of the grid block is less than a preset multiple of the average ground reference depth of the corresponding column, the grid block is marked as an obstacle block. Multiple consecutive obstacle blocks are determined to be effective obstacles. When the total proportion of effective obstacle blocks in the current area exceeds a preset proportion, the area ahead is determined to be blocked.

[0049] For each grid block in the central area, depth features are extracted based on normalized depth values. These depth features include dynamic range, edge pixel ratio, and low depth ratio. Based on the depth features, the obstacle type for each grid block is determined using preset classification rules. These obstacle types include wall obstacles, elongated obstacles, grid-like obstacles, head obstacles, suspended obstacles, and block obstacles. Voice prompts are provided for each obstacle type.

[0050] Preferably, the main control processor is further configured as follows:

[0051] Extract pixel strips at the bottom of the normalized depth map with a preset ratio as the ground reference area;

[0052] Determine the minimum depth value for all pixels in each column of the ground reference region;

[0053] Calculate the standard deviation of depth values ​​for all pixels in each column of the ground reference region, and calculate the global standard deviation of depth values ​​for all pixels in the ground reference region.

[0054] The confidence weight of each column is determined based on the standard deviation of the depth value of each column and the standard deviation of the global depth value: if the standard deviation of the depth value of the current column is less than the product of the standard deviation of the global depth value and the preset coefficient, then the confidence weight of the current column is set to the first weight value; otherwise, it is set to the second weight value. The first weight value is less than the second weight value.

[0055] The minimum depth values ​​of each column are weighted by median filtering based on their confidence weights to obtain the ground reference depth for each column:

[0056]

[0057] in, For the first List the ground reference depth, To preset window radius, For column index variables within the window, For the first Minimum depth value of the column, For the first Confidence weights for columns.

[0058] Preferably, identifying obstacles above the head based on physical distance within the upper detection area specifically includes:

[0059] The height of the upper detection area at the top of the normalized depth map is dynamically adjusted based on the smoothed pitch angle; the height of the upper detection area is negatively correlated with the smoothed pitch angle.

[0060] The percentage of pixels in the upper detection area whose physical distance is less than a preset distance threshold is counted.

[0061] When the pixel ratio is greater than a preset ratio threshold, it is determined that there is an obstacle above the head, and the physical distance of the corresponding ground area directly below the upper detection area is obtained; if the physical distance of the corresponding ground area directly below the upper detection area is greater than a preset ground distance threshold, it is determined that there is a suspended obstacle.

[0062] Preferably, the step of performing sliding window decision fusion on obstacle detection results from multiple consecutive frames to generate voice navigation commands specifically includes:

[0063] Maintain a sliding window with a preset frame length to store obstacle detection results and obstruction indicators for each area across multiple consecutive frames;

[0064] If there are wall obstacle labels or suspended obstacle labels in the sliding window and they appear a first preset number of times, then output the corresponding voice prompt.

[0065] The number of frames identified as blocked in the left, middle, and right areas of the sliding window is counted separately. If the number of blocked frames in either the left or right area reaches a second preset number and the number of blocked frames in the other area is less than the second preset number, a voice direction command pointing to the other area is output. If the number of blocked frames in the middle area reaches a third preset number, a stop command is output. Otherwise, a straight-through command is generated.

[0066] Set a cooldown period for non-emergency voice commands, and do not repeat the same command during the cooldown period; emergency voice commands are not subject to the cooldown period.

[0067] A guide method for the blind based on monocular depth estimation, applied to any of the above-mentioned systems, is characterized by comprising the following steps:

[0068] The image acquisition module acquires RGB image streams of the user's travel path in real time.

[0069] The inertial measurement unit acquires real-time triaxial acceleration and triaxial angular velocity data of the user's head.

[0070] The user's head posture information is obtained based on the three-axis acceleration data and three-axis angular velocity data. The user's head is then judged to be in an abnormal motion state based on the posture information. If so, the voice output module is paused. Otherwise, the quality of the current frame RGB image is evaluated.

[0071] If the quality assessment result of the current frame RGB image does not meet the preset quality requirements, the obstacle detection result of the previous valid frame will be used; otherwise, the depth of the current frame RGB image will be estimated by a monocular depth estimation model, and the original depth map output by the model will be post-processed to obtain a normalized depth map.

[0072] The user's walking motion is identified based on triaxial acceleration data to estimate single-step displacement. The scale factor is then calculated by combining the corner displacement information of the ground area and the normalized depth value of the corner, so as to convert the normalized depth value in the normalized depth map into physical distance.

[0073] Based on the pose information, determine the region of interest and the upper detection region of the normalized depth map; within the region of interest, identify ground obstacles based on the normalized depth value; within the upper detection region, identify obstacles above the head based on physical distance.

[0074] Sliding window decision fusion is performed on obstacle detection results from multiple consecutive frames to generate voice navigation commands;

[0075] The voice output module reads the voice navigation instructions to the user.

[0076] Compared with the prior art, the present invention has the following beneficial effects:

[0077] (1) Low hardware cost and lightweight device: This invention only requires a monocular camera and an IMU sensor, without the need for additional depth sensing hardware such as binocular cameras, structured light modules or ToF sensors. It can be integrated into a wearable device in the form of ordinary glasses, which is lightweight and has low power consumption, making it suitable for visually impaired users to wear for a long time.

[0078] (2) Comprehensive obstacle perception dimensions: This invention uses four complementary mechanisms, namely block grid analysis, Sobel gradient detection, wall detection and overhead threat detection, to identify ground obstacles of various shapes such as block, elongated, grid, and wall, as well as obstacles in special locations at head height and suspended height, covering the main types of obstacles faced by visually impaired users when traveling.

[0079] (3) Accurate distance estimation and precise direction suggestion: This invention determines obstacles based on a relative comparison mechanism of ground reference depth (threshold 0.85) rather than absolute depth value, which can adapt to changes in ground distance in different scenarios; through independent depth analysis of the left, middle and right regions, it provides users with accurate suggestions for left turn, right turn or straight-line direction, rather than just telling them "there is an obstacle ahead".

[0080] (4) User posture adaptive and low false alarm rate: The present invention monitors the change rate of the user's head pitch angle and yaw angle in real time through IMU attitude calculation. When the user lowers his head, raises his head or turns his head quickly, the broadcast will be automatically paused to avoid invalid alarms when the camera deviates from the normal viewing angle. At the same time, the threat detection area above is dynamically adjusted according to the pitch angle to improve the accuracy of overhead obstacle detection.

[0081] (5) Absolute scale online self-calibration: No pre-calibration or additional sensors are required. Metric depth can be recovered by using IMU and ground optical flow during walking. This enables the system to automatically adapt to the safe distance threshold in different scenarios and provide quantitative navigation instructions, significantly improving practicality and safety.

[0082] (6) IMU-Vision Tightly Coupled Domain Mapping: Breaking through the traditional fixed ROI clipping, the IMU is used to correct camera tilt and field of view changes in real time, ensuring that the physical meaning of left, middle and right partitions on uphill, downhill and side slopes is consistent, and avoiding misleading steering suggestions.

[0083] (7) Sliding window decision smoothing: eliminates broadcast flicker caused by inter-frame jitter, greatly reduces false alarm rate while ensuring response speed, and has natural robustness to short-term line of sight deviation or depth estimation noise.

[0084] (8) This invention does not rely on any target detection annotation data, is unsupervised, and is plug-and-play. Attached Figure Description

[0085] Figure 1 This is a flowchart of the overall processing of the present invention;

[0086] Figure 2 This is a photograph of the guide glasses of the present invention.

[0087] Figure 3 This is a schematic diagram of the ROI division in the depth map of the present invention, showing the top 20% exclusion area, the effective ROI area (divided into 16 rows × 9 columns of grid and horizontal partitions of the left, middle and right regions), and the bottom 15% ground reference strip;

[0088] Figure 4 This is a schematic diagram of the block obstacle detection method of the present invention;

[0089] Figure 5 This is a schematic diagram of the elongated obstacle detection method of the present invention;

[0090] Figure 6 This is a schematic diagram of the wall surface detection method of the present invention;

[0091] Figure 7 This is a schematic diagram of the head obstacle detection area of ​​the present invention. Detailed Implementation

[0092] The following is in conjunction with the appendix Figure 1-7 The technical solution of the present invention will be described in detail below.

[0093] The system described in this invention is integrated into a wearable device in the form of glasses. Its hardware components are as follows: The main control processor uses an ESP32, responsible for unified scheduling of camera acquisition, IMU data reading, depth estimation inference, and voice broadcasting; the image acquisition module uses an OV3660 camera, configured via an SCCB interface and connected to the main control processor, outputting an RGB image stream with a resolution of no less than 640×480; the attitude perception module uses an ICM42688 six-axis IMU, communicating with the ESP32 via an SPI interface to provide data from a three-axis accelerometer and a three-axis gyroscope; the voice output module uses a MAX98357A digital audio amplifier, connected to the main control processor via an I2S interface to receive audio data and drive the speakers on the glasses; the power subsystem uses a TP4056 lithium battery charge / discharge management chip in conjunction with an LDO voltage regulator circuit to provide stable 3.3V and 5V power supplies, and is equipped with ESD protection, reverse connection protection, and overcurrent protection circuits to improve the reliability of the wearable device. Figure 1 As shown, the complete system processing chain starts with image acquisition, goes through IMU pose calculation, image quality assessment, monocular depth estimation, ground reference depth calculation, ROI cropping, multi-shape obstacle detection, head obstacle detection, and finally outputs voice broadcast.

[0094] The method of the present invention includes the following steps:

[0095] Step 1, Image Acquisition. A monocular camera mounted on the glasses acquires a real-time RGB image stream of the travel path, which serves as input for subsequent depth estimation and obstacle detection.

[0096] Step 2, IMU Attitude Calculation. The inertial measurement unit (IMU) installed on the glasses acquires the user's three-axis head attitude angles in real time, providing a basis for subsequent broadcast suppression and dynamic adjustment of the overhead threat detection area. The IMU attitude calculation employs a fusion algorithm combining accelerometer and gyroscope data, including gravity vector low-pass filtering, pitch and roll angle calculation, yaw angle integration, and EMA smoothing of angle values.

[0097] ;

[0098] in, This is the filtered gravity vector at the current moment. This is the filtered gravity vector from the previous time step. The data represents the triaxial acceleration output by the inertial measurement unit at the current moment. These are the low-pass filter coefficients for the gravity vector;

[0099] Pitch angle With roll angle They were calculated as follows:

[0100]

[0101]

[0102] in, , , These are the components of the filtered gravity vector along the three axes of the inertial measurement unit at the current moment.

[0103] Yaw angle Using dead zones ( The integration of the leakage coefficient (0.2) and the integral method involves the following two steps:

[0104]

[0105]

[0106] in, The yaw axis angular velocity, The sampling interval;

[0107] All angle values ​​are smoothed using an exponential moving average (EMA, coefficient 0.15):

[0108] in, Refers to the smoothed pitch angle, roll angle, or yaw angle;

[0109] When the user is looking down (smoothed tilt angle) ), head tilt (smoothed pitch angle) ) or rapid turn (smoothed yaw rate of change) When in the ) state, the system pauses all broadcast output to avoid invalid alarms caused by the camera deviating from the normal viewing angle or motion blur.

[0110] Step 3, Image Quality Assessment. The system performs a quality assessment of the input image before performing depth estimation. After converting the input image to grayscale, the variance of the Laplacian operator response is calculated. To quantify image sharpness:

[0111]

[0112] in, This is the grayscale image obtained after converting the input image to grayscale.

[0113] like If the current frame is below a preset threshold, it is determined that there is motion blur or focus blur. The depth estimation of the current frame is skipped, and the obstacle detection result of the previous valid frame is used. If multiple consecutive frames are blurry frames, the broadcast frequency is reduced to avoid misleading prompts.

[0114] Step 4, Monocular Depth Estimation. The image that has passed the quality assessment is input into the lightweight depth estimation model DepthAnything V2-Small (ViT-S encoder, feature dimension 64, output channels [48,96,192,384]) for inference, and the original depth map is output.

[0115] Step 5, Depth Contrast Adaptation Based on Extreme Value Suppression. The original depth map is first subjected to extreme value suppression: calculate the 1% and 99th quantiles p1 and p99 of the total depth values, clamp pixels smaller than p1 to p1, and pixels larger than p99 to p99; then linearly stretch to the [0,1] interval, and finally invert (smallest values ​​are closer) to obtain the normalized depth map. This process avoids interference from extreme depth values ​​(isolated extremely deep / shallow points) caused by lens smudges or strong reflective spots, which could lead to errors in subsequent gradient and statistical calculations.

[0116] In addition, if there are spatial discontinuities in the depth map that exceed 5% of the total pixels (caused by lens smudges or water vapor), the system will mark the area as an untrusted area and exclude the corresponding grid block in the subsequent block grid analysis, making obstacle avoidance decisions only based on trustworthy areas.

[0117] Step 6, Least Squares-EMA Depth Online Calibration. Before obstacle detection, the system attempts to recover the absolute scale using walking motion.

[0118] Step 6.1, IMU zero-velocity detection. Calculate the acceleration amplitude. :

[0119]

[0120] in, These are the triaxial accelerations output by the inertial measurement unit, i.e. ;

[0121] When the amplitude reaches its peak and then falls back to around 9.8 m / s², and the rate of change satisfies When the zero velocity point is reached, it is marked as the zero velocity point. There is one gait cycle between adjacent zero velocity points.

[0122] Step 6.2, Single-step displacement estimation. Estimating the horizontal acceleration during the gait period. Perform trapezoidal integration to obtain the horizontal displacement components. :

[0123]

[0124] in, Triaxial acceleration data The acceleration component in the horizontal direction after deducting the gravitational component; , These are the start and end times of integration, respectively. For integration variables; , These are the times of the k-th and (k+1)-th sampling points, respectively. , These are the horizontal accelerations at the k-th and (k+1)-th sampling points, respectively.

[0125] Step 6.3, Ground Corner Extraction and Optical Flow Tracking. Within the bottom 15% ground region of the previous frame, 15-20 points are extracted using the FAST corner detector. These points are then tracked using Lucas-Kanade sparse optical flow in the current frame to obtain the horizontal disparity of each point. .

[0126] Step 6.4, establish the linear equation. Based on the pinhole model, for each valid point, we have:

[0127]

[0128] in For the first Normalized depth values ​​for each point, Let be the focal length of the camera in the horizontal direction. Solve for the scale factor using the least squares method. :

[0129]

[0130] in, To extract the total number of points using the FAST corner detector;

[0131] Step 6.5, Smoothing and Degradation. For consecutive valid frames... Perform EMA smoothing (coefficient) )

[0132]

[0133] in, , These are the smoothing scale factors for the current frame and the previous frame, respectively. The scale factor obtained by solving for the current frame;

[0134] The system determines that reliable scale calibration conditions are not available at the current moment if any of the following conditions are met:

[0135] Condition A (Stillness): The user is in a stationary state, as determined by the IMU zero-speed detection in step 6.1, and the duration of stillness exceeds 3 seconds;

[0136] Condition B (Solution Anomaly): The least squares solution residuals for 5 consecutive frames exceed a preset residual threshold (the residual is the mean square error between the actual and estimated values ​​of the least squares fit in step 6.4; for the i-th corner point, the actual value is...). The estimated value is ).

[0137] When any of the above conditions are met, the system pauses updating the smoothing scaling factor and directly uses the most recent valid historical value of the smoothing scaling factor:

[0138] when When the change exceeds 30%, a scene switch is determined, and the EMA is reset.

[0139]

[0140] get Afterwards, all subsequent depth comparison thresholds are converted to physical distances, such as safety distances. The corresponding normalization threshold is It is used to convert normalized depth maps into physical distances (meters), enabling obstacle threshold scene adaptation, quantitative voice broadcasting (such as "There is an obstacle 2 meters ahead"), and collision time estimation; it adopts least squares method + exponential moving average smoothing to reduce computational overhead and supports static degradation and automatic reset during scene switching.

[0141] Step 7, adaptive anchoring of ground depth column by column. Extraction The bottom 15% pixel strip, for each column Calculate the minimum depth value for this column. Simultaneously, the depth contrast characterization of this column within the striped region is calculated:

[0142]

[0143] in, For the first The total number of pixels listed in the bottom 15% strip. For pixel index, For pixels The normalized depth value, For the first The average normalized depth value listed in the bottom 15% band;

[0144] Calculate local standard deviation :

[0145]

[0146] Calculate the global ground strip standard deviation :

[0147]

[0148] in, This is the average normalized depth value of all pixels within the entire bottom 15% strip area;

[0149] Determine the confidence level :

[0150]

[0151] Final Column ground reference depth The result is obtained from weighted median filtering (with a kernel size of 5):

[0152]

[0153] The influence of low-texture columns is mitigated, preserving structural continuity; among them, For column index variables within the window, For the first Minimum depth value of the column, For the first Confidence weights for columns.

[0154] Furthermore, the local contrast weight is calculated independently for each column, and the confidence of anchoring in sparse texture areas is automatically reduced to avoid errors introduced by weak textured ground, making the ground depth anchoring robust to changes in lighting and surface material.

[0155] Step 8, IMU-driven region of interest (ROI) cropping. For example... Figure 3 As shown, this is to eliminate interference from non-ground areas such as the sky and ceiling on obstacle detection:

[0156] Step 8.1, based on the smoothed current pitch angle Calculate the top exclusion ratio :

[0157]

[0158] Ensure that when the pitch angle is positive (looking down), more sky area is excluded, and when it is negative (looking up), more information above is retained, so that the effective detection area is always aligned with the user's actual line of sight while walking.

[0159] Step 8.2, based on the smoothed roll angle Affine transformation correction is applied to the normalized depth map.

[0160]

[0161] in, Image height, The first line after excluding the top;

[0162] And the vertical dividing lines of the left, middle and right zones are as follows Perform a global translation to align the partition boundaries with the projection of the real horizontal plane onto the image plane. In practice, this can be achieved by performing an affine transformation on the depth map or by adjusting only the region segmentation lines.

[0163] Step 9: Unsupervised obstacle classification based on multi-dimensional fusion of deep features. The ROI after domain mapping in Step 8 is divided into three regions: left, middle, and right, in a 1:2:1 ratio. The left, middle, and right regions are then gridded, and the average depth of each grid block is calculated based on the normalized depth value. If the average depth of a grid block is less than a preset multiple of the average ground reference depth of the corresponding column, the grid block is marked as an obstacle block. Multiple consecutive obstacle blocks are determined to be valid obstacles. When the total proportion of valid obstacle blocks in the current region exceeds a preset ratio, the area ahead is determined to be obstructed.

[0164] The central section is divided into a 16x9 grid array; feature vectors are extracted for each grid block, including dynamic range. Edge pixel ratio and low depth ratio For each grid block ,set up This represents the total number of pixels within the block. For pixels The normalized depth value is then used to define the following four features:

[0165]

[0166]

[0167] Gradient detection is performed by applying a horizontal Sobel operator to the normalized depth map, with the convolution kernel being...

[0168]

[0169] For each pixel p in the entire image, its horizontal gradient magnitude Defined as:

[0170]

[0171] in, These are the row index and the column index, respectively. coordinates The convolution weight coefficients corresponding to the position, In pixels Centered on the 3x3 area (up, down, left, right), offset The normalized depth value of the location pixel;

[0172] Set edge threshold Statistical analysis of the horizontal gradient magnitude across the entire graph. Greater than The number of pixels is calculated, and its proportion to the total number of pixels in the entire image is denoted as the edge pixel proportion. :

[0173]

[0174] The classification rules are as follows:

[0175] like If identified as a wall / large flat surface, output "Front Wall";

[0176] like If the obstacle is identified as a long and thin barrier (railing, thin post), the output will be "Beware of thin railings";

[0177] like If the block is located in the upper 30% of the image, it is identified as a head / suspended obstacle. Based on the ground depth, it is determined whether it is suspended and the output is "Caution: Top of head".

[0178] Others were identified as blocky obstacles;

[0179] The percentage of effective obstacle blocks in the left, center, and right zones is statistically analyzed to determine the obstruction situation in each zone and to output turning suggestions.

[0180] Step 10, Head Obstacle Detection. The detection area height is dynamically adjusted based on the IMU pitch angle obtained in Step 2. :

[0181]

[0182] The detection area is made to adapt to the user's head pose. The physical distance corresponding to the depth value within this area is calculated, and the scale factor obtained in step 6 is used. Convert the normalized depth threshold to a physical threshold: Normalized Depth The corresponding physical distance is The system calculates the physical distance within the upper detection area. pixel ratio ,when If a threat is detected from above, a priority alert is triggered.

[0183] Furthermore, if the upper region The physical distance to the corresponding column of ground area directly below (i.e., the ground is far away) If the normalized depth value of the ground area directly below the upper detection area is used, it is identified as a suspended obstacle (such as a balcony extension or an open window), and a differentiated warning is output: "Caution: There is a suspended object overhead; you can duck to pass." When IMU data is unavailable, a downgrade judgment is made by comparing the average depth of the upper and lower halves of the depth map: if the average depth of the upper half is significantly less than that of the lower half, it is determined that there may be a threat above, and a warning is output.

[0184] Step 11, sliding time window decision fusion. Maintain a circular buffer of length 5 to store the classification results of each frame and the blocking flags of each region. After processing each frame, the data within the window is fused using majority voting and priority weighting:

[0185] Step 11.1, Prioritize Emergency Decisions: If the "Wall" or "Hanging Obstacle" label exists in the window and appears ≥2 times, immediately output the corresponding voice message without cooldown.

[0186] Step 11.2, Smoothing the steering decision: Count the number of "obstructed" frames in the left, middle and right zones respectively. If the number of obstructed frames in a certain zone is ≥3 and the number of obstructed frames in another zone is <3, then output a steering command to the unobstructed zone; if the number of obstructed frames in the middle zone is ≥4, then output "Stop"; if there is no obvious decision, output "Go straight".

[0187] Step 11.3, Cooling Management: The minimum broadcast interval for non-emergency instructions (such as "Obstruction ahead, please turn left" or "Safe ahead, you can go straight") is 2 seconds, while emergency instructions (such as "Wall ahead, please stop immediately!" or "Caution: There is a horizontal bar overhead, please duck down to pass!") are not restricted.

[0188] This method ensures that single-frame noise does not cause command jitter while maintaining a fast response to real obstacles.

[0189] This mechanism effectively eliminates false alarms and repeated broadcasts caused by single-frame jitter.

[0190] Step 12, Voice Broadcast Output. The decision results from each of the above detection modules are merged according to priority and converted into natural language voice commands, which are then broadcast to the user via the speaker on the glasses through the MAX98357A power amplifier. A cooling-off time limit is applied to the same decision type to avoid information overload caused by repeated broadcasts.

[0191] The priorities, from highest to lowest, are: Overhead Threat Alarm, Emergency Stop, Directional Guidance, and Safe Straight Ahead; Overhead Threat Alarm is not subject to a cooldown period; Directional Guidance outputs left or right turn suggestions based on the average depth of the left and right zones.

[0192] Example 1: Outdoor sidewalk walking scenario

[0193] As the user walks on the outdoor sidewalk, the system has completed online scale self-calibration; the current scale factor is... =0.30 (i.e., a normalized depth value of 1.0 corresponds to a physical distance of 0.30 m, and this value has been converged through step 6). There is a parked bicycle (block-shaped obstacle) about 2.5 meters ahead, occupying the central and right passageways.

[0194] The OV3660 camera captures the current frame, and the ICM42688 calculates... =1°, the system executes the complete detection process. The Depth Anything V2-Small inference output is post-processed to obtain a normalized depth map, which is then transformed in step 6 to obtain the physical distance:

[0195]

[0196] The physical distance corresponding to the middle area. The physical distance corresponding to the area on the right. : The physical distance corresponding to the left area.

[0197] The obstacles in the central and right zones are both within 2 meters, classifying them as "emergency" level; the left zone is clear. Grid analysis shows that obstacles account for 35% of the area in the central and right zones, and the system determines that the right side is blocked while the left side is clear. Based on the distance classification and turning suggestions in step 9, the system announces: "There is an obstacle approximately 1.7 meters ahead, please turn left."

[0198] The user veered to the left and continued moving forward; the next frame detected... After conversion, the physical distance is greater than 5 meters (safe), the proportion of obstacle blocks drops to 3%, and the system announces: "Safe ahead, you can go straight," guiding the user to successfully bypass the obstacle.

[0199] : Normalized depth of the intermediate region : Normalized depth of the right region : Normalized depth of the left region.

[0200] Example 2: Abnormal User Posture Scenario

[0201] When a user looks down at their phone while walking in an indoor corridor, the ICM42688 calculates... If the angle is greater than 15° (=22°), the system immediately pauses all broadcast output to prevent false alarms when the camera is facing the ground. This is after the user looks up. The temperature dropped to 8°, the system resumed normal detection procedures, detected that the corridor ahead was clear, and announced: "It is safe ahead, you can proceed straight ahead."

[0202] Example 3: Overhead horizontal bar scenario

[0203] As the user walks to the entrance of the underground parking lot, a height restriction bar (a suspended obstacle) is in front of them. The system has already calibrated the scale factor. =0.35 (In this scenario, a normalized depth of 1.0 corresponds to 0.35m). Step 10: Calculate the detection area based on the pitch angle:

[0204]

[0205] This refers to the upper 27% of the image. The system calculates the physical distance corresponding to the depth values ​​within this region. :

[0206]

[0207] calculate Pixel ratio ( The system detected a close-range threat from above. Simultaneously, the normalized depth of the ground area directly below... ), Therefore, it is determined to be a head obstacle (such as a horizontal bar directly blocking the road), and the system announces: "Caution: There is a horizontal bar overhead. Please duck down to pass."

[0208] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A guide system for the blind based on monocular depth estimation, characterized in that, It includes an image acquisition module, a posture perception module, a main control processor, a voice output module, and a power supply subsystem integrated into the head-mounted wearable device; The image acquisition module is used to acquire RGB image streams of the user's travel path in real time; The attitude perception module includes an inertial measurement unit, which is used to acquire three-axis acceleration data and three-axis angular velocity data of the user's head in real time; The main control processor is configured as follows: The user's head posture information is obtained based on the three-axis acceleration data and three-axis angular velocity data. The user's head is then judged to be in an abnormal motion state based on the posture information. If so, the voice output module is paused. Otherwise, the quality of the current frame RGB image is evaluated. If the quality assessment result of the current frame's RGB image does not meet the preset quality requirements, the obstacle detection result of the previous valid frame will be used. Otherwise, the depth of the current frame's RGB image is estimated using a monocular depth estimation model, and the original depth map output by the model is post-processed to obtain a normalized depth map. The user's walking motion is identified based on triaxial acceleration data to estimate single-step displacement. The scale factor is then calculated by combining the corner displacement information of the ground area and the normalized depth value of the corner, so as to convert the normalized depth value in the normalized depth map into physical distance. Based on the pose information, determine the region of interest and the upper detection region of the normalized depth map; within the region of interest, identify ground obstacles based on the normalized depth value; within the upper detection region, identify obstacles above the head based on physical distance. Sliding window decision fusion is performed on obstacle detection results from multiple consecutive frames to generate voice navigation commands; The voice output module is used to broadcast voice navigation instructions to the user; The power subsystem is used to supply power to the guide system.

2. The guide system for the blind based on monocular depth estimation according to claim 1, characterized in that, The system obtains the user's head posture information based on triaxial acceleration and triaxial angular velocity data, and determines whether the user's head is in an abnormal movement state based on the posture information. If so, the voice output module's broadcast is paused, as detailed below: The triaxial acceleration data output by the inertial measurement unit is low-pass filtered, and the pitch angle is calculated based on the filtered triaxial acceleration data. and roll angle : in, This is the filtered gravity vector at the current moment. This is the filtered gravity vector from the previous time step. The data represents the triaxial acceleration output by the inertial measurement unit at the current moment. These are the low-pass filter coefficients for the gravity vector; , , These are the components of the filtered gravity vector along the three axes of the inertial measurement unit at the current moment. A dead zone threshold is set for the yaw axis angular velocity output by the inertial measurement unit. If the absolute value of the current yaw axis angular velocity is less than or equal to the dead zone threshold, the yaw axis angular velocity increment at the current moment is ignored, and the yaw angle at the previous moment is kept as the yaw angle at the current moment. If the absolute value of the current yaw axis angular velocity is greater than the dead zone threshold, the product of the current yaw axis angular velocity and the sampling interval is accumulated and added to the yaw angle at the previous moment. A leakage coefficient is applied to the accumulated yaw angle to attenuate it, and the yaw angle at the current moment is obtained. The pitch, roll, and yaw angles are smoothed using exponential moving averages, and user status is determined accordingly. When the smoothed pitch angle is greater than the first preset threshold, it is determined to be a nose-down state; when the smoothed pitch angle is less than the second preset threshold, it is determined to be a nose-up state; when the rate of change of the smoothed yaw angle is greater than the third preset threshold, it is determined to be a rapid head-turn state. The voice broadcast will pause when the user is looking down, looking up, or turning their head quickly.

3. The guide system for the blind based on monocular depth estimation according to claim 1, characterized in that, The quality assessment of the current frame's RGB image is specifically as follows: After converting the RGB image of the current frame to grayscale, the variance of the Laplacian operator response is calculated. When the variance is lower than a preset threshold, the current frame is determined to be a blurry frame. The depth estimation of the current frame is skipped and the obstacle detection result of the previous valid frame is used. If a preset number of consecutive frames are all judged as blurry frames, the broadcast frequency will be reduced.

4. A guide system for the blind based on monocular depth estimation according to claim 1, characterized in that, The post-processing of the original depth map output by the model to obtain a normalized depth map specifically includes: Extreme value suppression is performed on the original depth map: the lower quantile and higher quantile of the depth values ​​of the entire original depth map are calculated, pixel values ​​less than the lower quantile are clamped to the lower quantile, and pixel values ​​greater than the higher quantile are clamped to the higher quantile. Contrast stretching is performed on the depth map after extreme value suppression to map the depth values ​​to a predetermined range; The mapped depth map is inverted so that the depth values ​​of nearby objects are less than the depth values ​​of distant objects, resulting in a normalized depth map. Untrusted regions are detected in the normalized depth image based on the proportion of abnormal connected regions with abrupt changes in depth value, and the corresponding grid blocks of untrusted regions are excluded in the subsequent grid block analysis.

5. A guide system for the blind based on monocular depth estimation according to claim 2, characterized in that, The process of identifying the user's walking motion based on triaxial acceleration data to estimate single-step displacement, and combining this with corner displacement information of the ground area and normalized depth values ​​of the corners to solve for the scale factor, specifically includes: Based on triaxial acceleration data Perform zero-velocity detection to identify the user's gait cycle: calculate acceleration amplitude. , , These are the triaxial accelerations output by the inertial measurement unit, i.e. When the acceleration amplitude When the acceleration amplitude falls back to the preset amplitude range after reaching the peak and the rate of change of the acceleration amplitude is less than the preset rate of change threshold, it is marked as a zero velocity point. The interval between adjacent zero velocity points is one gait cycle. Horizontal acceleration during gait period Perform trapezoidal integration to obtain the horizontal displacement components. ,in Triaxial acceleration data The acceleration component in the horizontal direction after deducting the gravitational component; Corner points are extracted from the bottom ground region of the previous RGB image. The horizontal parallax of each corner point is obtained in the current RGB image through optical flow tracing. The corresponding normalized depth value is obtained from the normalized depth map of the current frame based on the image coordinates of each corner point. According to the horizontal displacement component A linear equation is established based on the horizontal parallax and the normalized depth value, and the scale factor is solved using the least squares method. : in: This refers to the focal length of the camera in the horizontal direction. For the first The normalized depth values ​​corresponding to each corner point For the first Horizontal parallax of each corner point This represents the total number of corner points; The obtained scale factor is smoothed by exponential moving average to obtain a smoothed scale factor. When the preset pause condition is met, the update of the smoothed scale factor is paused and the historical value is used. When the rate of change of the smoothed scale factor exceeds the preset rate of change threshold, the scene is switched, the historical cumulative value of the exponential moving average smoothing is reset, and the scale factor obtained in the current frame is used as the new historical benchmark.

6. A guide system for the blind based on monocular depth estimation according to claim 2, characterized in that, The process of identifying ground obstacles within the region of interest based on normalized depth values ​​specifically includes: Based on the smoothed pitch angle, the exclusion ratio of the invalid region at the top of the normalized depth map is calculated to eliminate non-ground areas and obtain the region of interest; the exclusion ratio is negatively correlated with the smoothed pitch angle, and the exclusion ratio is limited between a preset upper limit value and a preset lower limit value. The region of interest is divided into left, middle and right sections according to a preset ratio, with the width of the middle section being greater than that of the left and right sections; Based on the smoothed roll angle, calculate the horizontal offset of each row in the region of interest. Then, translate the partition boundary points on the vertical dividing lines of the left, middle and right zones according to the horizontal offset of their respective rows. Connect the translated partition boundary points to form the corrected partition boundary lines, so that the partition boundaries are aligned with the real horizontal plane. The left, middle and right areas are gridded, and the average depth of each grid block is calculated based on the normalized depth value. If the average depth of the grid block is less than a preset multiple of the average ground reference depth of the corresponding column, the grid block is marked as an obstacle block. Multiple consecutive obstacle blocks are determined to be effective obstacles. When the total proportion of effective obstacle blocks in the current area exceeds a preset proportion, the area ahead is determined to be blocked. For each grid block in the central area, depth features are extracted based on normalized depth values. These depth features include dynamic range, edge pixel ratio, and low depth ratio. Based on the depth features, the obstacle type for each grid block is determined using preset classification rules. These obstacle types include wall obstacles, elongated obstacles, grid-like obstacles, head obstacles, suspended obstacles, and block obstacles. Voice prompts are provided for each obstacle type.

7. A guide system for the blind based on monocular depth estimation according to claim 6, characterized in that, The main control processor is also configured to: Extract pixel strips at the bottom of the normalized depth map with a preset ratio as the ground reference area; Determine the minimum depth value for all pixels in each column of the ground reference region; Calculate the standard deviation of depth values ​​for all pixels in each column of the ground reference region, and calculate the global standard deviation of depth values ​​for all pixels in the ground reference region. The confidence weight of each column is determined based on the standard deviation of the depth value of each column and the standard deviation of the global depth value: if the standard deviation of the depth value of the current column is less than the product of the standard deviation of the global depth value and the preset coefficient, then the confidence weight of the current column is set to the first weight value; otherwise, it is set to the second weight value. The first weight value is less than the second weight value. The minimum depth values ​​of each column are weighted by median filtering based on their confidence weights to obtain the ground reference depth for each column: in, For the first List the ground reference depth, To preset window radius, For column index variables within the window, For the first Minimum depth value of the column, For the first Confidence weights for columns.

8. A guide system for the blind based on monocular depth estimation according to claim 1, characterized in that, The process of identifying obstacles above the head based on physical distance within the upper detection area specifically includes: The height of the upper detection area at the top of the normalized depth map is dynamically adjusted based on the smoothed pitch angle; the height of the upper detection area is negatively correlated with the smoothed pitch angle. The percentage of pixels in the upper detection area whose physical distance is less than a preset distance threshold is counted. When the pixel ratio is greater than a preset ratio threshold, it is determined that there is an obstacle above the head, and the physical distance of the corresponding ground area directly below the upper detection area is obtained; if the physical distance of the corresponding ground area directly below the upper detection area is greater than a preset ground distance threshold, it is determined that there is a suspended obstacle.

9. A guide system for the blind based on monocular depth estimation according to claim 1, characterized in that, The step of performing sliding window decision fusion on obstacle detection results from multiple consecutive frames to generate voice navigation commands specifically includes: Maintain a sliding window with a preset frame length to store obstacle detection results and obstruction indicators for each area across multiple consecutive frames; If there are wall obstacle labels or suspended obstacle labels in the sliding window and they appear a first preset number of times, then output the corresponding voice prompt. The number of frames identified as blocked in the left, middle, and right areas of the sliding window is counted separately. If the number of blocked frames in either the left or right area reaches a second preset number and the number of blocked frames in the other area is less than the second preset number, a voice direction command pointing to the other area is output. If the number of blocked frames in the middle area reaches a third preset number, a stop command is output. Otherwise, a straight-through command is generated. Set a cooldown period for non-emergency voice commands, and do not repeat the same command during the cooldown period; emergency voice commands are not subject to the cooldown period.

10. A guide method for the blind based on monocular depth estimation, applied to the system described in any one of claims 1-9, characterized in that, Includes the following steps: The image acquisition module acquires RGB image streams of the user's travel path in real time. The inertial measurement unit acquires real-time triaxial acceleration and triaxial angular velocity data of the user's head. The user's head posture information is obtained based on the three-axis acceleration data and three-axis angular velocity data. The user's head is then judged to be in an abnormal motion state based on the posture information. If so, the voice output module is paused. Otherwise, the quality of the current frame RGB image is evaluated. If the quality assessment result of the current frame's RGB image does not meet the preset quality requirements, the obstacle detection result of the previous valid frame will be used. Otherwise, the depth of the current frame's RGB image is estimated using a monocular depth estimation model, and the original depth map output by the model is post-processed to obtain a normalized depth map. The user's walking motion is identified based on triaxial acceleration data to estimate single-step displacement. The scale factor is then calculated by combining the corner displacement information of the ground area and the normalized depth value of the corner, so as to convert the normalized depth value in the normalized depth map into physical distance. Based on the pose information, determine the region of interest and the upper detection region of the normalized depth map; within the region of interest, identify ground obstacles based on the normalized depth value; within the upper detection region, identify obstacles above the head based on physical distance. Sliding window decision fusion is performed on obstacle detection results from multiple consecutive frames to generate voice navigation commands; The voice output module reads the voice navigation instructions to the user.