A blind obstacle avoidance system based on deep learning and dynamic window algorithm

CN122835387APending Publication Date: 2026-09-29JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610890834.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

它们仅能提示前方有障碍物;缺乏动态的路径规划,无法为用户提供向左绕行或向右绕行的具体导航指令,导致用户在复杂环境中依然感到迷茫

Benefits of technology

[0065]本发明通过深度学习模型的轻量化处理、二维栅格地图替代三维点云、局部规划算法替代全局建图的多重优化,能够在低功耗嵌入式平台上实现低延迟的实时避障响应,满足盲人行走场景对系统轻便性与实时性的双重需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122835387A_ABST
    Figure CN122835387A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of computer vision intelligent obstacle avoidance navigation, and particularly relates to a blind person obstacle avoidance system based on deep learning and a dynamic window algorithm; the system comprises a hardware control perception module, which is used for collecting environmental visual information output by a binocular depth camera; a visual obstacle identification module, which is used for identifying and classifying obstacles according to environmental visual information based on a lightweight deep learning model; an occupancy grid map generation module, which is used for converting depth information in a depth image into a two-dimensional occupancy grid map and determining the grid state; a dynamic window obstacle avoidance planning module, which is used for planning an obstacle avoidance path in real time on the two-dimensional occupancy grid map and selecting an optimal path; and a multi-level voice feedback module, which is used for broadcasting obstacle avoidance prompts and walking direction instructions to a user according to the danger level of obstacles and the optimal path; the application can realize low-delay real-time obstacle avoidance response on a low-power embedded platform, and meet the dual requirements of portability and real-time performance of the system in a blind person walking scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision intelligent obstacle avoidance and navigation technology, and in particular relates to an obstacle avoidance system for the blind based on deep learning and dynamic window algorithm. Background Technology

[0002] Currently, mobility aids for visually impaired individuals are mainly divided into three categories: traditional physical aids (such as guide canes), satellite-based voice navigation systems (such as mobile phone maps), and environmental perception-based electronic obstacle avoidance systems. Traditional guide canes can only detect a very small area below the feet and cannot detect obstacles above the waist (such as protruding air conditioner units or tree branches) or moving pedestrians or vehicles, posing significant safety hazards. GPS-based voice navigation systems can only provide macroscopic road directions and cannot detect sudden obstacles in the microscopic environment (such as temporary construction barriers or parked non-motorized vehicles).

[0003] With the development of computer vision and sensor technology, some vision-based obstacle avoidance systems have emerged in recent years. A search revealed a prior art obstacle avoidance device for the blind based on a monocular camera, which detects objects ahead through image recognition and issues an alarm. However, this type of technology has the following technical drawbacks:

[0004] Limited perception dimensions and lack of geometric information: Existing solutions mostly use monocular cameras or ultrasonic sensors. Although monocular cameras can identify object types (such as pedestrians and cars) through deep learning, they have difficulty obtaining precise distance and contour information between objects and users, which easily leads to false alarms; although ultrasonic sensors can measure distances, they cannot identify object types and cannot tell users whether the road ahead is traversable grass or a dangerous pit.

[0005] The conflict between computational load and real-time performance: To achieve high-precision obstacle avoidance, some existing technologies attempt to combine LiDAR or stereo vision with complex SLAM algorithms. These solutions typically require high-performance industrial PCs or cloud computing support, resulting in large system size, high power consumption, and high cost. More importantly, when running on embedded platforms, if the algorithm is not lightweighted, the computational latency for visual recognition and path planning is high (typically greater than 200 milliseconds). In highly dynamic scenarios such as blind people walking, this latency causes obstacle avoidance commands to lag, failing to meet real-time requirements.

[0006] Obstacle avoidance strategies and interaction methods are often mechanical: Most existing electronic obstacle avoidance systems use simple logic such as stopping upon encountering an obstacle or sounding an alarm. They can only indicate that there is an obstacle ahead; they lack dynamic path planning and cannot provide users with specific navigation instructions to detour to the left or right, causing users to feel confused even in complex environments.

[0007] In summary, how to achieve efficient integration of binocular vision and deep learning while ensuring lightweight design and low power consumption, generate smooth obstacle avoidance paths using two-dimensional grid maps combined with dynamic window algorithm (DWA), and provide blind users with navigation information that includes both object types and precise spatial orientation through multi-level semantic voice is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0008] In view of this, the present invention aims to provide a blind obstacle avoidance system based on deep learning and dynamic window algorithm. The present invention achieves low-latency real-time obstacle avoidance response on a low-power embedded platform through multiple optimizations, such as lightweight processing of deep learning model, replacement of three-dimensional point cloud with two-dimensional grid map and replacement of global mapping with local planning algorithm, thus meeting the dual requirements of system portability and real-time performance for blind people in walking scenarios.

[0009] To achieve the above objectives, the technical solution of the present invention is implemented as follows:

[0010] A blind obstacle avoidance system based on deep learning and dynamic programming algorithms includes:

[0011] The hardware control perception module is used to acquire environmental visual information output by the binocular depth camera, the environmental visual information including color RGB images and depth images;

[0012] The visual obstacle recognition module uses a lightweight deep learning model to identify and classify obstacles based on environmental visual information, outputs obstacle category labels and bounding box information, and obtains obstacle distance by registering color RGB images and depth images.

[0013] The occupied raster map generation module is used to convert depth information in a depth image into a two-dimensional occupied raster map and determine the raster state.

[0014] The dynamic window obstacle avoidance planning module is used to plan obstacle avoidance paths in real time on the two-dimensional occupied grid map based on the dynamic window algorithm, and select the optimal path through the evaluation function.

[0015] The multi-level voice feedback module is used to broadcast obstacle avoidance prompts and walking direction instructions to the user based on the obstacle hazard level and the optimal path.

[0016] The visual obstacle recognition module uses the YOLOv5-Lite network as the target detection model, and its backbone network is a lightweight structure based on ShuffleNetV2.

[0017] The training process of the target detection model is as follows:

[0018] A. Construct an obstacle dataset, which includes multiple color RGB images captured by an Orbbec Gemini Pro camera, and obstacle annotation information corresponding to each color RGB image;

[0019] The obstacle annotation information includes: the bounding box of the obstacle in the color RGB image, the bounding box coordinates, and the obstacle category label. The bounding box coordinates are composed of the coordinates of the upper left and lower right vertices of the bounding box in the pixel coordinate system of the color RGB image. B. The YOLOv5-Lite network is trained using the obstacle dataset. That is, the color RGB images in the obstacle dataset are used as the input of the YOLOv5-Lite network, and the obstacle annotation information corresponding to each color RGB image is used as the expected output of the YOLOv5-Lite network to train the YOLOv5-Lite network. Training is stopped when the training rounds reach a preset number of rounds, and a trained object detection model is obtained.

[0020] The preferred size of the training input image for the YOLOv5-Lite network is 320×320;

[0021] After model training is completed, the obtained PyTorch format weight file is converted to ONNX format, and inference is performed on an embedded single-board computer platform using the ONNXRuntime inference engine to achieve model lightweighting.

[0022] The visual obstacle recognition module acquires obstacle distance information specifically including:

[0023] Calculate the pixel coordinates of the geometric center of the obstacle bounding box in the color RGB image; perform pixel-level registration and alignment between the color RGB image and the depth image;

[0024] Based on the pixel coordinates of the geometric center, the depth value at the corresponding position in the indexed depth image is used as the obstacle distance.

[0025] The specific content of the occupation grid map generation module in constructing the two-dimensional occupation grid map includes:

[0026] S1. Generation of 3D point cloud: Generate a 3D point cloud based on the depth information in the depth image;

[0027] S2. Two-dimensional top-view projection: Project the obstacle point cloud in the three-dimensional point cloud onto the horizontal plane, discard the height coordinate information of each point in the obstacle point cloud, and obtain the two-dimensional top-view coordinates corresponding to each three-dimensional point.

[0028] S3. Raster Mapping: Based on the two-dimensional top-view coordinates, the corresponding two-dimensional projection points are assigned to the corresponding raster cells, thereby generating a two-dimensional occupied raster map;

[0029] S4. Parameter settings for 2D occupied grid map:

[0030] Map extent: raster resolution 0.1m × 0.1m;

[0031] S5. Grid stabilization mechanism: Count the number of two-dimensional projection points falling into each grid cell. If the number of points is greater than or equal to a preset threshold, the grid cell is determined to be occupied; otherwise, it is in an idle state.

[0032] The determination of the obstacle point cloud: points in the three-dimensional point cloud whose height coordinates exceed a preset height threshold form the obstacle point cloud, and only the two-dimensional top-view projection and raster mapping are performed on the obstacle point cloud.

[0033] The update process of the two-dimensional occupied grid map is as follows:

[0034] When the binocular depth camera moves, the following dynamic update operation is performed:

[0035] The current coverage area of ​​the sliding window is determined based on the current position of the camera;

[0036] Remove the historical map data of the current coverage area. The historical map data includes the number of points stored in the corresponding grid and the occupancy status information.

[0037] The newly acquired 3D point cloud within the current coverage area is projected onto the horizontal plane. Based on the 2D top-view coordinates of each point after projection, the number of points falling into each grid is counted, and this number is recorded as the current number of points in the corresponding grid.

[0038] Based on the current number of points in each grid, the occupancy status is redefined; thereby, the content of the two-dimensional occupied grid map is dynamically refreshed.

[0039] The dynamic window obstacle avoidance planning module transforms the local path planning problem into an optimization problem in the velocity space, as detailed below:

[0040] S1. A preset velocity window, wherein the velocity window includes a selectable set of linear velocities and a selectable set of angular velocities;

[0041] S2. For a selected velocity combination, starting from the origin of the coordinate system on the two-dimensional occupied grid map, iteratively calculate K steps according to the kinematic model of the predicted trajectory to obtain a discrete trajectory point sequence, which is denoted as a predicted trajectory.

[0042] S3. Iterate through the velocity combinations within the window to generate multiple predicted trajectories;

[0043] S4. Filter valid trajectories for calculating the comprehensive score: For each predicted trajectory, iterate through all points on the trajectory and check whether the grid where each point is located is occupied; if any point on the trajectory falls into an occupied grid, the trajectory is marked as a collision trajectory and is not included in the calculation of the comprehensive score.

[0044] S5. Calculate the overall score for each valid trajectory. :

[0045] ;

[0046] in Score the distance to the finish line:

[0047] ;

[0048] The coordinates of the endpoint of the valid trajectory on a two-dimensional occupied grid map; The coordinates of the destination point on the two-dimensional occupied grid map;

[0049] Obstacle distance rating ;

[0050] in Obstacle grid distance:

[0051] ;

[0052] Where n is the number of occupies in a two-dimensional grid map to predict the trajectory of the th element. Centered on a point, with a radius of... The number of grid cells occupied within the circular area. , To predict the trajectory of the first The raster index of the raster cell containing each point; , These represent the first and second digits within the circular region, respectively. A raster index that occupies a raster cell;

[0053] Steering angle penalty split ;

[0054] in The turning angle of the effective trajectory endpoint relative to the starting point;

[0055] S6. Optimal Path Selection

[0056] Calculate the comprehensive score of each valid trajectory and select the trajectory with the highest score as the obstacle avoidance path at the current moment.

[0057] The multi-level voice feedback module is specifically configured as follows:

[0058] The multi-level voice feedback module calculates the direction angle from the starting point to the ending point based on the received optimal path, and generates and broadcasts the corresponding direction guidance command based on the preset angle interval mapping relationship.

[0059] The calculation of the direction angle from the starting point to the ending point specifically includes:

[0060] Obtain the starting coordinates of the optimal path coordinates of the endpoint Construct direction vector ( Based on the direction vector, calculate the direction angle θ from the starting point to the ending point, where θ = arctan( );

[0061] The method of generating corresponding directional guidance instructions based on the preset angle interval mapping relationship specifically includes: generating a straight-ahead instruction when the direction angle θ satisfies 80°≤θ≤110°; generating a right-turn instruction when the direction angle θ satisfies 0°≤θ<80°; generating a left-turn instruction when the direction angle θ satisfies 110°<θ≤180°; otherwise, it is determined to be an unreachable area.

[0062] The multi-level voice feedback module receives the obstacle recognition results from the visual obstacle recognition module and broadcasts them via voice; the obstacle recognition results include the obstacle's category label and the obstacle's distance.

[0063] The hardware control and perception module includes an embedded single-board computer platform, a power supply module, and a binocular depth camera. The power supply module provides power to the embedded single-board computer platform and the binocular depth camera. The binocular depth camera acquires environmental visual information and embeds it into the embedded single-board computer platform, and then configures the system on the embedded single-board computer platform.

[0064] Compared with the prior art, the present invention can achieve the following beneficial effects:

[0065] This invention achieves low-latency, real-time obstacle avoidance response on a low-power embedded platform through multiple optimizations, including lightweight processing of deep learning models, replacement of three-dimensional point clouds with two-dimensional grid maps, and replacement of global mapping with local planning algorithms. This meets the dual requirements of portability and real-time performance for blind people in walking scenarios.

[0066] The grid stabilization mechanism introduced in this invention effectively overcomes the problem of frequent grid state jumps caused by single-point noise in traditional grid maps, providing a stable and reliable environmental representation for path planning, improving the continuity and reliability of obstacle avoidance. Through dynamic window algorithm sampling and multi-dimensional evaluation in velocity space, it can generate smooth and safe obstacle avoidance trajectories and effectively suppress unnecessary frequent turns, making the obstacle avoidance path more in line with human walking habits.

[0067] This invention provides graded voice broadcasts based on the hazard level of obstacles and converts planned paths into intuitive walking direction instructions. Combined with periodic prompts and a priority arbitration mechanism, it avoids information overload, ensures the priority delivery of key information, and improves the naturalness and usability of human-computer interaction. Attached Figure Description

[0068] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0069] Figure 1 This is a block diagram illustrating the overall working principle of the present invention.

[0070] Figure 2 This is a flowchart illustrating the workflow of the visual impairment recognition module of the present invention.

[0071] Figure 3 This is a flowchart of the dynamic window obstacle avoidance planning module of the present invention.

[0072] Figure 4 This is a rendering of the visual impairment recognition module of the present invention.

[0073] Figure 5 The image shows the effect of the grid map generation module and the dynamic window obstacle avoidance planning module of the present invention. Detailed Implementation

[0074] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0075] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0076] Please see Figure 1 A blind obstacle avoidance system based on deep learning and dynamic programming algorithms includes:

[0077] The hardware control perception module is used to acquire environmental visual information output by the binocular depth camera, the environmental visual information including color RGB images and depth images;

[0078] The visual obstacle recognition module uses a lightweight deep learning model to identify and classify obstacles based on environmental visual information, and outputs the category labels and bounding box information of the obstacles. At the same time, this module also obtains the distance information of obstacles by registering color RGB images and depth images.

[0079] The occupying grid map generation module is used to acquire a depth image, generate a point cloud based on the depth information in the depth image, generate a two-dimensional occupying grid map based on the point cloud, and introduce a grid stabilization mechanism to determine the grid state.

[0080] The dynamic window obstacle avoidance planning module is used to plan obstacle avoidance paths in real time on the two-dimensional occupied grid map based on the dynamic window algorithm, and select the optimal path through the evaluation function.

[0081] The multi-level voice feedback module is used to broadcast obstacle avoidance prompts and walking direction instructions to the user through a priority arbitration mechanism based on the obstacle hazard level and the optimal path.

[0082] The modules work together to form a closed-loop system of perception-planning-feedback. The system structure of this invention is as follows: Figure 1 As shown.

[0083] The hardware control and sensing module includes: an embedded single-board computer platform, a binocular depth camera, and a power supply module.

[0084] The embedded single-board computer platform uses a Raspberry Pi 5 as the core computing unit. This platform boasts high computing power and a good power efficiency, meeting the real-time computing requirements of a lightweight obstacle avoidance system. This embodiment of the invention utilizes Raspberry Pi OS (based on a Linux kernel system) as the software development environment. Necessary drivers and software libraries are installed on this system, including the Orbbec Gemini Pro camera SDK, a Python runtime environment, and the ONNX Runtime inference engine.

[0085] The binocular depth camera used is the Gemini Pro model from Orbbec, which connects to a Raspberry Pi 5 via a Type-C interface, enabling the input of acquired environmental visual information to an embedded single-board computer platform. This camera can simultaneously output color RGB images, depth images, and infrared images, with a maximum depth measurement range of approximately 0.2m to 3m, meeting the perception distance requirements for obstacle avoidance scenarios for the blind.

[0086] The power supply module uses the MicroSnow UPS Module 3S as the power supply system. This module supports lithium battery power supply and charging management, and can provide a stable 5V / 5A power output for the Raspberry Pi 5 and the binocular depth camera, ensuring continuous operation of the system in mobile scenarios.

[0087] The hardware control and sensing module calls the Orbbec SDK, which is compatible with the Orbbec GeminiPro camera, in the embedded single-board computer platform to obtain the output color RGB image, depth image, and infrared image. It then performs time synchronization on the color RGB image, depth image, and infrared image to ensure that the images are synchronized in time. Finally, it unifies the color space of the time-synchronized images, that is, it aligns the color characteristics of all images to the same standard color space to obtain the processed standardized visual information data.

[0088] The technical solution of the visual impairment recognition module specifically includes:

[0089] (I) Model Selection and Training

[0090] The target detection model employs the YOLOv5-Lite network, a lightweight convolutional neural network-based target detection network. The input to this model is a color RGB image to be detected, and the output is a color RGB image with obstacle detection results. The detection results include the bounding boxes and coordinates of each obstacle in the color RGB image, as well as the category label for each obstacle. The bounding box coordinates are composed of the coordinates of the top-left and bottom-right vertices of the bounding box in the pixel coordinate system of the color RGB image.

[0091] The YOLOv5-Lite network source is:

[0092] https: / / github.com / ppogg / YOLOv5-Lite is a lightweight and improved version of YOLOv5. It reduces the number of model parameters and computation through channel pruning and depthwise separable convolution techniques, making it suitable for deployment on embedded platforms.

[0093] The YOLOv5-Lite network includes an input terminal, a backbone network, a neck network, and a head network, wherein:

[0094] The input terminal is used to receive the color RGB image to be detected and input the received color RGB image into the Backbone network;

[0095] The backbone network adopts a lightweight structure based on the ShuffleNetV2 network to extract feature maps layer by layer from the color RGB image. The feature maps include multiple levels from low-level features to high-level semantic features.

[0096] The Neck network, located between the Backbone network and the Head network, receives multiple feature maps of different levels output by the Backbone network and performs bidirectional feature fusion on the received feature maps from top to bottom and bottom to top, so as to output three feature maps corresponding to targets of different scales to the Head network, thereby enhancing the network's ability to perceive targets of different scales.

[0097] The Head network receives the fused feature map output by the Neck network and predicts the target category, bounding box, and bounding box coordinates on the feature maps at three different scales, thereby generating the final detection result.

[0098] The training process of YOLOv5-Lite is as follows:

[0099] 1. Construct an obstacle dataset, which contains 1043 color RGB images captured by an Orbbec Gemini Pro camera, and obstacle annotation information corresponding to each color RGB image;

[0100] The obstacle labeling information includes: the bounding box and bounding box coordinates of the obstacle in the color RGB image, and the obstacle category label. The bounding box coordinates are composed of the coordinate values ​​of the upper left and lower right vertices of the bounding box in the pixel coordinate system of the color RGB image.

[0101] In one embodiment of the present invention, the obstacle dataset used for training is an open-source dataset that already contains images and their corresponding bounding boxes, bounding box coordinates, and object category annotation information.

[0102] In one embodiment of the present invention, the obstacle dataset used for training is partly a self-constructed dataset and partly an open-source dataset. The self-constructed dataset uses the labelimg annotation tool to delineate bounding boxes for obstacles in each image and define the corresponding object categories, while retaining the bounding box coordinates and category annotation results.

[0103] During training, the color RGB images are scaled to a pre-selected input size, which is chosen from 480×480, 480×640, or 320×320; 2. The YOLOv5-Lite network is trained using the obstacle dataset, that is, the color RGB images in the obstacle dataset are used as the input of the YOLOv5-Lite network, and the obstacle annotation information corresponding to each color RGB image is used as the expected output of the YOLOv5-Lite network to train the YOLOv5-Lite network. Training stops when the preset number of training rounds is reached, and a trained object detection model is obtained. In this embodiment, the number of rounds is 300, and training stops after 300 rounds.

[0104] (ii) Lightweighting of the model

[0105] To enable real-time inference on an embedded single-board computer platform, this module performs lightweight processing on the trained model:

[0106] After model training is completed, a weight file (.pt file) in PyTorch format is obtained. To achieve efficient deployment on an embedded single-board computer, the PyTorch weight file is converted into an ONNX (Open Neural Network Exchange) model file. The ONNX Runtime inference engine is used to load the ONNX model file and perform inference on a Raspberry Pi 5.

[0107] ONNX is an open deep learning model exchange format that supports model migration between various frameworks. It reduces model inference latency and computational resource consumption while maintaining detection accuracy, thereby enabling real-time obstacle recognition on the embedded single-board computer. The ONNX Runtime is optimized for ARM architecture, significantly improving inference speed compared to native PyTorch inference.

[0108] (III) Optimal Input Size

[0109] In this embodiment, the preferred input image size during training is 320×320. Actual testing shows that a 320×320 image size provides the fastest inference speed and consumes the least computational resources, meeting the real-time requirements of the obstacle avoidance system. While 480×480 and 480×640 images offer higher detection accuracy, they increase inference latency by approximately 40%-60%, potentially causing delays in obstacle avoidance commands in real-world walking scenarios. Therefore, this system preferentially uses 320×320 as the model input size.

[0110] (iv) Obstacle localization and ranging

[0111] 1. Let the coordinates of the top-left and bottom-right vertices of the obstacle bounding box in the pixel coordinate system of the color RGB image be respectively... and The coordinates of the geometric center point of the obstacle in the pixel coordinate system are... Calculate using the following formula:

[0112]

[0113] A pixel coordinate system is established with the top left corner of the color RGB image as the origin, the positive x-axis is to the right along the horizontal direction of the image, and the positive y-axis is downward along the vertical direction of the image.

[0114] 2. Perform pixel-level registration and alignment between the color RGB image acquired by the binocular depth camera and the depth image, so that each pixel in the color RGB image can find a corresponding depth value in the depth image that represents the distance to that point;

[0115] 3. Based on the coordinates of the geometric center point of the obstacle In the depth image, pixels at the same location are indexed, and the depth value corresponding to that pixel is read as the depth value corresponding to the geometric center point of the obstacle. The unit is meters; this depth value represents the distance from the geometric center to the camera along the camera's optical axis, denoted as obstacle distance.

[0116] In one embodiment, if the depth value corresponding to the geometric center of the obstacle is invalid, it is considered to be infinite and denoted as NA.

[0117] The occupation grid map generation module constructs a two-dimensional occupation grid map, specifically including:

[0118] 1. Definition and projection of three-dimensional coordinates

[0119] A depth image is acquired, and a 3D point cloud is generated based on the depth information in the depth image, as follows:

[0120] Map each valid pixel in the depth image to the camera coordinate system to generate the corresponding three-dimensional spatial coordinates of that pixel. ,in = (Depth value);

[0121] The camera coordinate system has the camera optical center as its origin. The axis is horizontal to the right. The z-axis points forward along the camera's optical axis (depth direction), and the z-axis points vertically upward (height direction).

[0122] The three-dimensional point cloud is formed by mapping all the valid pixels in the depth image to obtain all the three-dimensional spatial points.

[0123] The three-dimensional coordinates of each point in the three-dimensional point cloud are denoted as follows: The obstacle point cloud in the 3D point cloud is subjected to dimensionality reduction processing, that is, the obstacle point cloud is projected onto the horizontal plane along the z-axis of the camera coordinate system, and the height coordinate information of each point in the obstacle point cloud is discarded to obtain the 2D projection points corresponding to each 3D point. The 2D top view coordinates are denoted as follows. .

[0124] Based on the two-dimensional top-down coordinates, the corresponding two-dimensional projection points are assigned to the corresponding grids to generate a two-dimensional occupied grid map; wherein, each grid of the two-dimensional occupied grid map corresponds to a region on the horizontal plane (XY plane), and the attribute value of the grid is determined according to the number of two-dimensional points falling into the region.

[0125] Obstacle point cloud is defined as a 3D point cloud whose height coordinates exceed a preset height threshold. The point cloud is composed of points; in this embodiment, the height threshold is set to 10cm. That is, points with a height exceeding 10cm are identified as obstacle point clouds, and point clouds with a height below this threshold (such as ground, pebbles, etc.) are considered passable areas.

[0126] 2. Map parameter settings:

[0127] Map range:

[0128] Based on the actual needs of blind people for obstacle avoidance, this module is constructed with the optical center of the binocular depth camera as the origin and the direction along the camera's optical axis pointing towards the scene as the reference point. The positive direction of the optical axis is defined by the horizontal direction perpendicular to the optical axis. The camera space coordinate system of the axis is a reference to a two-dimensional occupied grid map;

[0129] The map range is set as follows:

[0130] The axial direction, i.e. the depth direction (front): 0.1m to 4.0m. The area within 0.1m is the emergency obstacle avoidance zone, and the area beyond 4.0m is the far field region. Obstacles beyond this range have little impact on the current obstacle avoidance decision and are not modeled.

[0131] The axial direction, i.e. the horizontal direction, is centered on the camera's optical axis and extends 3m to the left and right, for a total width of 6.0m. This range covers the width of the safe passage for blind people to walk on.

[0132] Raster resolution:

[0133] The grid resolution is 0.1m, which means the grid size is set to 0.1m × 0.1m. This resolution can effectively represent the outline details of obstacles (such as stone blocks, road posts, etc.) while ensuring that the map matrix size is moderate (the map matrix size is 35×60, corresponding to 35 rows in the depth direction and 60 columns in the horizontal direction), and the computational cost is controllable.

[0134] Raster mapping:

[0135] For any two-dimensional projection point Its corresponding raster index The calculation is as follows:

[0136]

[0137] in:

[0138] , where L is the raster resolution, Lx = 6.0m, and L is the total horizontal width of the two-dimensional raster map. For depth-direction raster index (0 ≤ i ≤ 35), For horizontal raster index (0 ≤ j ≤ 60);

[0139] 3. Grid stabilization mechanism

[0140] To address the issue of raster state jumps caused by single-point clouds in traditional raster maps, this module introduces a raster stabilization mechanism.

[0141] For each grid cell in a two-dimensional occupied grid map Count the number of two-dimensional projection points projected onto the grid based on the grid index. (Counting points involves cumulatively counting points with the same index within each grid cell); in this embodiment, the point count threshold is set to... The rules for determining the occupancy status of grid cells are as follows:

[0142] like If the grid is occupied, it is determined that the grid is occupied and recorded as an occupied grid, indicating that there is an obstacle there and it is impassable.

[0143] like If the grid is detected, it is determined to be in an empty state and recorded as an empty grid, indicating that the area has been detected and is unobstructed, and can be passed through.

[0144] This mechanism determines the raster state by accumulating statistical results from multiple point clouds, avoiding frequent changes in raster state caused by single-point noise (such as depth camera measurement errors), and effectively improving the stability of the map.

[0145] 4. Two-dimensional grid map update

[0146] The occupied grid map generation module uses a sliding window mechanism to dynamically update the two-dimensional occupied grid map, and the map coverage area is updated in real time as the stereo depth camera moves.

[0147] Specifically, the two-dimensional occupancy grid map consists of multiple grids, each grid corresponding to the camera coordinate system. A fixed-size area on a plane (i.e., a horizontal plane); the map data maintained by each grid includes: the number of two-dimensional projection points projected onto the grid cell, and the occupancy status of the grid determined based on the number;

[0148] When the binocular depth camera moves, the following dynamic update operation is performed:

[0149] The current coverage area of ​​the sliding window (i.e., the range of the color RGB image currently captured by the camera) is determined based on the current position of the camera.

[0150] Remove the historical map data of the current coverage area, which includes the number of points stored and their occupancy status information in all grid cells of the current coverage area;

[0151] For newly acquired 3D point clouds that fall within the current coverage area (i.e. the range of the color RGB image currently captured by the camera), project them onto the horizontal plane along the z-axis of the camera coordinate system, and assign them to the corresponding grid according to the two-dimensional top-view coordinates of each point after projection. Count the number of points falling into each grid and update the number of points in that grid.

[0152] Based on the updated point count, redetermine the occupancy status of each grid cell;

[0153] This enables dynamic refreshing of the content of the two-dimensional grid map.

[0154] The map always centers on the camera's current field of view, maintaining only local environmental information within a certain range around it. There is no need to build or maintain a globally consistent map, which effectively reduces the consumption of computing resources and ensures the real-time performance of the system on the embedded platform.

[0155] The dynamic window obstacle avoidance planning module transforms the local path planning problem into an optimization problem in the velocity space, and generates candidate trajectories based on a preset fixed velocity window;

[0156] The velocity window is jointly defined by the selectable set of linear velocities and the selectable set of angular velocities. Subsequent trajectory prediction steps traverse each velocity combination within this window to generate candidate trajectories; specifically:

[0157] (a) The speed window is set as follows:

[0158] The set of possible linear velocities is set to V, where (unit: );

[0159] The selectable set of angular velocities is set as follows: ,in The range is arrive ,interval This interval setting can smoothly cover the turning direction while controlling the number of samples to ensure computational efficiency.

[0160] (ii) Trajectory prediction parameters

[0161] The trajectory prediction parameters are set as follows:

[0162] Number of prediction steps in a single path prediction: ;

[0163] Duration per step: ;

[0164] Total prediction time is the time required for one path prediction: ;

[0165] The kinematic model formula for predicting the trajectory is as follows:

[0166]

[0167] in It is the x-axis coordinate (i.e., the x-axis coordinate in the camera space rectangular coordinate system) of the trajectory point on the two-dimensional occupied grid map at step t, with units of . ; +1 It predicts the x-axis coordinates of the trajectory point up to step t+1; It is the y-axis coordinate of the trajectory point on the two-dimensional grid map at step t, in units of ; +1 It predicts the y-axis coordinate of the trajectory point up to step t+1; It is the angle of the velocity direction predicted up to step t, that is, the angle between the line connecting the predicted trajectory point at step t and the origin of the coordinate system on the 2D grid map, and the x-axis, with units of 1000°. ; +1 It is the angle of the velocity direction predicted at step t+1, that is, the angle between the line connecting the predicted trajectory point at step t+1 and the origin of the coordinates on the two-dimensional grid map and the x-axis; It is the speed selected from the available set of online speeds, in units of ; The angular velocity is selected from the available set of angular velocities, in units of . ;

[0168] For the selected speed combination Starting from the current pose, i.e., the origin of the coordinate system on the 2D occupied grid map (the origin of the camera space Cartesian coordinate system), iterate for K steps according to the above formula to obtain the discrete trajectory point sequence. ; where the predicted angle of the current pose yes This angle indicates that the initial direction was straight ahead.

[0169] That is, after selecting the linear velocity and angular velocity, the motion is simulated step by step according to the above conditions. A total of K steps are simulated and predicted. The path taken in K steps is the predicted trajectory under the given conditions.

[0170] Iterate through all velocity combinations within the velocity window to generate multiple predicted trajectories.

[0171] (III) Destination setting

[0172] The target point is set 1.75m in front of the camera; that is, its coordinates on the 2D grid map are... , The point (in the direction of the camera's optical axis); this distance corresponds to a position about 3-4 steps in front of the blind person, and is a key decision area for obstacle avoidance planning.

[0173] (iv) Evaluation function design

[0174] The evaluation function for this module is designed as follows:

[0175]

[0176] in A comprehensive score for the predicted trajectory;

[0177] The definitions are as follows:

[0178] 1. Finish line distance score ;

[0179] in To predict the coordinates of the trajectory's endpoint; this score takes a negative value, so that the closer the trajectory is to the destination point, the higher the score.

[0180] 2. Obstacle Distance Scoring

[0181] For each point on the predicted trajectory Calculate the distance between obstacle grids within a radius of 1m centered on that point;

[0182] Specifically, for the first [item] on the predicted trajectory A point, on a two-dimensional grid map, with that point as the center and a radius of... Within the circular area, search for all occupied grid cells, calculate the Euclidean distance from each occupied grid cell to the point, and take the average value as the obstacle grid distance to that point. :

[0183]

[0184] Where n is the number of grid cells occupied within the circular area. For grid cells that only partially fall within the circular area, they are still counted as a complete grid cell in n. , To predict the trajectory of the first The raster index of the raster cell containing each point; , These represent the first and second digits within the circular region, respectively. A raster index that occupies a raster cell;

[0185] Then, the obstacle grid distances of all points on the predicted trajectory are summed, and the average of the steps is taken to obtain the obstacle distance score. :

[0186]

[0187] The higher the score, the further away the predicted trajectory is from the obstacle, and the higher the safety.

[0188] 3. Steering angle penalty is shared equally.

[0189] The absolute value of the total steering angle of the predicted trajectory is the equal part of the steering angle penalty.

[0190]

[0191] in The cumulative turning angle (in radians) of the predicted trajectory endpoint relative to the starting point.

[0192] The larger the steering angle, the greater the penalty, thus reducing the overall score of the trajectory and suppressing unnecessary frequent steering.

[0193] (v) Collision determination

[0194] For each predicted trajectory, iterate through all points on the trajectory. Check whether the grid where each point is located is occupied; if any point on the trajectory falls into an occupied grid, the trajectory is marked as a collision trajectory and is directly excluded from the calculation of the comprehensive score.

[0195] (vi) Optimal path selection

[0196] Iterate through all feasible speed combinations within the speed window (3 options for linear velocity × 17 options for angular velocity = 51 combinations) to generate multiple predicted trajectories; each speed combination corresponds to one trajectory, calculate the comprehensive score of each predicted trajectory, and select the predicted trajectory with the highest score as the obstacle avoidance path at the current moment.

[0197] The technical solution of the multi-level voice feedback module specifically includes:

[0198] (a) Hazard level classification

[0199] Based on the obstacle category output by the visual impairment recognition module, obstacles are classified into three danger levels:

[0200] Hazardous obstacles: These include vehicle-type obstacles, specifically cars, trucks, motorcycles, buses, and tricycles. These obstacles move at high speeds, are large in size, and pose a high risk of collision, requiring the highest level of warning.

[0201] More dangerous category: This includes obstacles with slow movement speed, specifically bicycles, pedestrians, and pets; these obstacles have a certain degree of unpredictability (they may move suddenly), but their movement speed is relatively slow, requiring a moderate level of attention.

[0202] Non-hazardous obstacles: These include other static obstacles, specifically tactile paving, trash cans, spherical barriers, posts, fire hydrants, parking signs, warning posts, and reflective cones. These obstacles are static objects, and the tactile paving is a passable area, requiring only simple prompts or no prompts.

[0203] (ii) Setting the feedback cycle

[0204] This module uses a periodic timed triggering method, and the feedback period can be adjusted according to actual needs.

[0205] Obstacle avoidance planning feedback: The multi-level voice feedback module receives the optimal path from the dynamic window obstacle avoidance planning module and converts the optimal path into a walking direction command, wherein the path starting point is based on the two-dimensional grid map. and the finish line The straight line vector generates semantic instructions including detouring to the left, detouring to the right, or going straight; specifically:

[0206] The multi-level voice feedback module calculates the direction angle from the starting point to the ending point based on the received optimal path, and generates corresponding directional guidance instructions based on the preset angle interval mapping relationship.

[0207] The calculation of the direction angle from the starting point to the ending point specifically includes:

[0208] Obtain the starting coordinates of the optimal path coordinates of the endpoint Construct direction vector ( Based on the direction vector, calculate the direction angle θ from the starting point to the ending point, where θ = arctan( );

[0209] The method of generating corresponding directional guidance instructions based on the preset angle interval mapping relationship specifically includes: generating a straight-ahead instruction when the direction angle θ satisfies 80°≤θ≤110°; generating a right-turn instruction when the direction angle θ satisfies 0°≤θ<80°; generating a left-turn instruction when the direction angle θ satisfies 110°<θ≤180°; otherwise, it is determined to be an unreachable area.

[0210] The multi-level voice feedback module broadcasts directional guidance instructions determined in the current image every 2 seconds via the espeak voice engine, ensuring that users can obtain path guidance in a timely manner.

[0211] Obstacle recognition feedback: The multi-level voice feedback module receives the obstacle recognition results from the visual obstacle recognition module. The obstacle recognition results include the category label of the obstacle and the distance to the obstacle. The obstacle recognition results in the current image are broadcast by voice every 10 seconds (using the espeak voice engine). The obstacle recognition results fed back by the visual obstacle recognition module change slowly, so a low feedback frequency is used to avoid information overload.

[0212] (III) Stable Feedback Mechanism

[0213] To avoid false alarms caused by single recognition errors, this module introduces a stable feedback mechanism: for the received obstacle recognition result, the obstacle recognition result obtained by the current frame color RGB image is compared with the obstacle recognition result obtained by the previous frame color RGB image. Only after confirming that the same obstacle is identified as the same category at the same location can the judgment be passed and the voice broadcast be triggered.

[0214] This mechanism effectively reduces false alarms caused by camera noise or model misdetection.

[0215] (iv) Content of voice broadcast

[0216] When multiple obstacle avoidance planning feedback messages and obstacle recognition feedback messages are triggered at the same time, a priority arbitration mechanism is adopted:

[0217] First priority: obstacle avoidance information (path planning results), walking direction instructions;

[0218] Second priority: Obstacle identification results for hazardous obstacles;

[0219] Third priority: Obstacle identification results for more dangerous obstacles;

[0220] Fourth priority: Obstacle identification results for non-hazardous obstacles;

[0221] Within the same priority level from the second to the fourth priority level, information is broadcast sequentially from closest to furthest from obstacles. This mechanism ensures that critical information is delivered first and avoids information conflicts.

[0222] The broadcast content is "Ahead + Distance + Obstacle Type", as follows:

[0223] For dangerous obstacles: broadcast the obstacle category and distance information, such as "There is a car 3 meters ahead, please be careful to avoid it".

[0224] For more dangerous obstacles: broadcast the obstacle category and distance information, such as "There is a pedestrian 2 meters to the left".

[0225] For non-dangerous obstacles: only announce the obstacle category, such as "there is a stone block ahead", or do not announce it depending on the actual situation.

[0226] The preferred embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the scope of protection of the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, any person skilled in the art can make equivalent substitutions or changes based on the technical solution and inventive concept of the present invention within the scope of the technology disclosed in the present invention. These simple modifications are all within the scope of protection of the present invention.

Claims

1. A blind obstacle avoidance system based on deep learning and dynamic programming algorithms, characterized in that, include: The hardware control perception module is used to acquire environmental visual information output by the binocular depth camera, the environmental visual information including color RGB images and depth images; The visual obstacle recognition module uses a lightweight deep learning model to identify and classify obstacles based on environmental visual information, outputs obstacle category labels and bounding box information, and obtains obstacle distance by registering color RGB images and depth images. The occupied raster map generation module is used to convert depth information in a depth image into a two-dimensional occupied raster map and determine the raster state. The dynamic window obstacle avoidance planning module is used to plan obstacle avoidance paths in real time on the two-dimensional occupied grid map based on the dynamic window algorithm, and select the optimal path through the evaluation function. The multi-level voice feedback module is used to broadcast obstacle avoidance prompts and walking direction instructions to the user based on the obstacle hazard level and the optimal path.

2. The system according to claim 1, characterized in that, The visual obstacle recognition module uses the YOLOv5-Lite network as the target detection model, and its backbone network is a lightweight structure based on ShuffleNetV2.

3. The system according to claim 2, characterized in that, The training process of the target detection model is as follows: A. Construct an obstacle dataset, which includes multiple color RGB images captured by an Orbbec Gemini Pro camera, and obstacle annotation information corresponding to each color RGB image; The obstacle labeling information includes: the bounding box of the obstacle in the color RGB image, the bounding box coordinates, and the obstacle category label. The bounding box coordinates are composed of the coordinates of the upper left and lower right vertices of the bounding box in the pixel coordinate system of the color RGB image. B. Train the YOLOv5-Lite network using the obstacle dataset, that is, use the color RGB images in the obstacle dataset as the input of the YOLOv5-Lite network, and use the obstacle annotation information corresponding to each color RGB image as the expected output of the YOLOv5-Lite network to train the YOLOv5-Lite network. Stop training when the training rounds reach a preset number of rounds to obtain a trained object detection model.

4. The system according to claim 3, characterized in that, The preferred size of the training input image for the YOLOv5-Lite network is 320×320; After model training is completed, the obtained PyTorch format weight file is converted to ONNX format, and inference is performed on an embedded single-board computer platform using the ONNXRuntime inference engine to achieve model lightweighting.

5. The system according to claim 1, characterized in that, The visual obstacle recognition module acquires obstacle distance information specifically including: Calculate the pixel coordinates of the geometric center of the obstacle bounding box in the color RGB image; perform pixel-level registration and alignment between the color RGB image and the depth image; Based on the pixel coordinates of the geometric center, the depth value at the corresponding position in the indexed depth image is used as the obstacle distance.

6. The system according to claim 1, characterized in that, The specific content of the occupation grid map generation module in constructing the two-dimensional occupation grid map includes: S1. Generation of 3D point cloud: Generate a 3D point cloud based on the depth information in the depth image; S2. Two-dimensional top-view projection: Project the obstacle point cloud in the three-dimensional point cloud onto the horizontal plane, discard the height coordinate information of each point in the obstacle point cloud, and obtain the two-dimensional top-view coordinates corresponding to each three-dimensional point. S3. Raster Mapping: Based on the two-dimensional top-view coordinates, the corresponding two-dimensional projection points are assigned to the corresponding raster cells, thereby generating a two-dimensional occupied raster map; S4. Parameter settings for 2D occupied grid map: Map extent: raster resolution 0.1m × 0.1m; S5. Grid stabilization mechanism: Count the number of two-dimensional projection points falling into each grid cell. If the number of points is greater than or equal to a preset threshold, the grid cell is determined to be occupied; otherwise, it is in an idle state. The determination of the obstacle point cloud: points in the three-dimensional point cloud whose height coordinates exceed a preset height threshold form the obstacle point cloud, and only the two-dimensional top-view projection and raster mapping are performed on the obstacle point cloud.

7. The system according to claim 6, characterized in that, The update process of the two-dimensional occupied grid map is as follows: When the binocular depth camera moves, the following dynamic update operation is performed: The current coverage area of ​​the sliding window is determined based on the current position of the camera; Remove the historical map data of the current coverage area. The historical map data includes the number of points stored in the corresponding grid and the occupancy status information. The newly acquired 3D point cloud within the current coverage area is projected onto the horizontal plane. Based on the 2D top-view coordinates of each point after projection, the number of points falling into each grid is counted, and this number is recorded as the current number of points in the corresponding grid. Based on the current number of points in each grid, the occupancy status is redefined; thereby, the content of the two-dimensional occupied grid map is dynamically refreshed.

8. The system according to claim 6, characterized in that, The dynamic window obstacle avoidance planning module transforms the local path planning problem into an optimization problem in the velocity space, as detailed below: S1. A preset velocity window, wherein the velocity window includes a selectable set of linear velocities and a selectable set of angular velocities; S2. For a selected velocity combination, starting from the origin of the coordinate system on the two-dimensional occupied grid map, iteratively calculate K steps according to the kinematic model of the predicted trajectory to obtain a discrete trajectory point sequence, which is denoted as a predicted trajectory. S3. Iterate through the velocity combinations within the window to generate multiple predicted trajectories; S4. Filter valid trajectories for calculating the comprehensive score: For each predicted trajectory, iterate through all points on the trajectory and check whether the grid where each point is located is occupied; if any point on the trajectory falls into an occupied grid, the trajectory is marked as a collision trajectory and is not included in the calculation of the comprehensive score. S5. Calculate the overall score for each valid trajectory. : ; in Score the distance to the finish line: ; The coordinates of the endpoint of the valid trajectory on a two-dimensional occupied grid map; The coordinates of the destination point on the two-dimensional occupied grid map; Obstacle distance rating ; in Obstacle grid distance: ; Where n is the number of occupies in a two-dimensional grid map to predict the trajectory of the th element. Centered on a point, with a radius of... The number of grid cells occupied within the circular area. , To predict the trajectory of the first The raster index of the raster cell containing each point; , These represent the first and second digits within the circular region, respectively. A raster index that occupies a raster cell; Steering angle penalty split ; in The turning angle of the effective trajectory endpoint relative to the starting point; S6. Optimal Path Selection Calculate the comprehensive score of each valid trajectory and select the trajectory with the highest score as the obstacle avoidance path at the current moment.

9. The system according to claim 1, characterized in that, The multi-level voice feedback module is specifically configured as follows: The multi-level voice feedback module calculates the direction angle from the starting point to the ending point based on the received optimal path, and generates and broadcasts the corresponding direction guidance command based on the preset angle interval mapping relationship. The calculation of the direction angle from the starting point to the ending point specifically includes: Obtain the starting coordinates of the optimal path coordinates of the endpoint Construct direction vector ( Based on the direction vector, calculate the direction angle θ from the starting point to the ending point, where θ = arctan( ); The method of generating corresponding directional guidance instructions based on the preset angle interval mapping relationship specifically includes: generating a straight-ahead instruction when the direction angle θ satisfies 80°≤θ≤110°; generating a right-turn instruction when the direction angle θ satisfies 0°≤θ<80°; generating a left-turn instruction when the direction angle θ satisfies 110°<θ≤180°; otherwise, it is determined to be an unreachable area. The multi-level voice feedback module receives the obstacle recognition results from the visual obstacle recognition module and broadcasts them via voice; the obstacle recognition results include the obstacle's category label and the obstacle's distance.

10. The system according to claim 1, characterized in that, The hardware control and perception module includes an embedded single-board computer platform, a power supply module, and a binocular depth camera. The power supply module provides power to the embedded single-board computer platform and the binocular depth camera. The binocular depth camera acquires environmental visual information and embeds it into the embedded single-board computer platform, and then configures the system on the embedded single-board computer platform.