A blind guidance device based on multiple sensors

Through the guide device combined with multiple sensors and deep learning models, the existing guide device has solved the problem of single functions and high cost, and achieved safer and more efficient guide in complex outdoor environments.

CN114387584BActive Publication Date: 2025-06-03INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210049898.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-12-22
Filing Date
2022-01-17
Publication Date
2025-06-03
Estimated Expiration
2042-01-17

AI Technical Summary

Technical Problem

The existing blinding equipment has a single function and cannot effectively adapt to complex outdoor environments, resulting in limited application scenarios and high training and use costs.

Method used

A variety of sensors (such as global positioning modules, inertia measuring instruments, depth cameras and ultrasonic sensors) are used to combine deep learning models (such as ESPNetV2, YOLOv5 and PENet) for path planning, identify feasible areas and obstacles, generate optimal paths, and convey action instructions through human-computer interaction modules.

Benefits of technology

It realizes safer and more efficient blindness in complex outdoor environments, reduces training and use costs, and expands application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387584B_ABST
    Figure CN114387584B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent blind guiding device based on multiple sensors, including a variety of sensors for collecting the current position information, movement information and various environmental information of the blind; a path planning device; and a human-computer interaction module for converting the optimal path into an action instruction for guiding the blind to move. The present invention can greatly expand the application scenarios of existing wearable blind guiding devices and help the blind group travel more efficiently and safely in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, specifically to the technology of using multi-sensors for guiding the blind, and more specifically, to a blind guiding device based on multiple sensors. Background Art

[0002] How to help the blind move independently and efficiently in complex environments is a research field that has received much attention. When walking in outdoor environments, the blind must overcome many difficulties, such as avoiding obstacles and finding passable areas. Therefore, most blind people are reluctant to go out alone. To improve the quality of life of the blind population and help them travel more conveniently, there are currently various solutions, among which the blind cane and the guide dog are the two most common tools for assisting the blind in traveling. However, the function of the blind cane is too simple and can only help users detect some nearby obstacles by touching, so there are great safety hazards. The guide dog can assist the blind in completing various complex tasks such as avoiding obstacles and crossing intersections, and is a relatively good tool for assisting in traveling. However, the training and use costs of guide dogs are expensive, bringing a heavy economic burden to users. With the continuous development and maturity of related technologies such as robots, sensors, and artificial intelligence, various auxiliary blind guiding devices have been proposed one after another. For example, electronic blind canes, mechanical guide dogs, and wearable blind guiding devices. However, the functions of the current existing technologies are relatively single, can only complete some simple tasks, have low adaptability to complex outdoor environments, and the application scenarios are relatively limited. Summary of the Invention

[0003] Therefore, the object of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a blind guiding device based on multiple sensors.

[0004] The object of the present invention is achieved by the following technical solutions:

[0005] According to a first aspect of the present invention, there is provided a path planning device for guiding the blind, the path planning device being configured to: perform global path planning based on the destination and current position information of the blind person to obtain a global path; use a first deep learning model to identify a feasible area in the space where the blind person is located based on the color map collected by the depth camera to obtain a feasible area identification result; use a second deep learning model to perform object detection on the space where the blind person is located based on the color map collected by the depth camera to obtain an object detection result; use the object detection result and the depth map collected by the depth camera to determine a first detection result of obstacles and use a plurality of ultrasonic sensors to determine a second detection result of obstacles; map the space corresponding to the color map to a preset walking space grid, and determine a cost scaling factor of the corresponding grid in the walking space grid according to the feasible area identification result, the first detection result, and the second detection result; use a path planning algorithm to perform local path planning based on the global path and the cost scaling factor of the grid to obtain an optimal path for guiding the blind person to walk. Alternatively, map the space corresponding to the color map to a preset walking space grid, and determine a cost scaling factor of the corresponding grid in the walking space grid according to the feasible area identification result, the first detection result, the second detection result, and the correspondence between the depth map and the color map.

[0006] In some embodiments of the present invention, the first deep learning model is implemented using an improved ESPNetV2 model, wherein the improved ESPNetV2 model is obtained by adding a feature pyramid module at the end of the feature extraction module of the original ESPNetV2 model.

[0007] In some embodiments of the present invention, the improved ESPNetV2 model is trained in the following manner: use the collected road-related data set to train the improved ESPNetV2 model to segment and identify the feasible area in the sample image in the road data set, and output a feasible area identification result, wherein the road-related data set includes sample images and feasible area segmentation labels, and the feasible area segmentation labels include annotation values indicating that the corresponding areas of the sample images are blind paths, sidewalks, roadways, zebra crossings, and backgrounds; use cross entropy as a loss function to calculate the loss value of segmenting the feasible area according to the output feasible area identification result and the feasible area segmentation label, and update the parameters of the improved ESPNetV2 model according to the loss value of segmenting the feasible area.

[0008] In some embodiments of the present invention, the second deep learning model is implemented using an improved YOLOv5 model, wherein the improved YOLOv5 model is obtained by replacing the 3x3 ordinary convolution in the original YOLOv5 model with a depthwise separable convolution.

[0009] In some embodiments of the present invention, the second deep learning model is trained as follows: Use the COCO2017 dataset to train the improved YOLOv5 model to perform object detection on the sample images in the dataset, and output the object detection results; calculate the total loss of the object detection according to the output object detection results and object detection labels, and update the parameters of the improved YOLOv5 according to the total loss of the object detection, where the total loss of the object detection includes the classification prediction sub-loss and the detection box prediction sub-loss.

[0010] In some embodiments of the present invention, the first detection result includes: obstacles determined based on the object detection result of the image and obstacles determined by detecting the mutation of the depth value according to the depth map.

[0011] In some embodiments of the present invention, the depth map used for determining the obstacles by detecting the mutation of the depth value according to the depth map is the depth map after the associated depth map is completed by using the third deep learning model according to the color map collected by the depth camera.

[0012] In some embodiments of the present invention, the third deep learning model is implemented using the PENet model and is trained as follows: Use the KITTI-Depth dataset to train the PENet model to complete the associated depth map according to the color map, use the mean square error loss function to calculate the loss value according to the completed depth map and the depth label, and update the parameters of the PENet model according to the loss value.

[0013] In some embodiments of the present invention, when determining the cost scaling factor of the corresponding grid in the walking space grid, the relative degree to which the cost of the corresponding grid is reduced is as follows: blind path ≥ zebra crossing > sidewalk > roadway > obstacle, where the obstacles include pedestrians, bicycles, cars, traffic signs, and other ground obstacles.

[0014] In some embodiments of the present invention, when determining the cost scaling factor of the corresponding grid in the walking space grid, the basis for whether the corresponding grid is an obstacle is as follows: For long-distance obstacles with a distance greater than a predetermined distance threshold, the priority of the second detection result is higher than that of the first detection result; for short-distance obstacles with a distance less than or equal to the predetermined distance threshold, the priority of the first detection result is higher than that of the second detection result.

[0015] In some embodiments of the present invention, the path planning device is configured to: The number of frames of the color map collected by the depth camera used to update the feasible area recognition result per unit time is less than the number of frames of the color map collected by the depth camera used to update the object detection result per unit time.

[0016] According to a second aspect of the present invention, there is provided a blind guiding device based on multiple sensors, including: multiple sensors for collecting the current position information, movement information, and various environmental information of a blind person; a path planning device as described in the first aspect; and a human-computer interaction module for converting the optimal path into an action instruction for guiding the movement of the blind person.

[0017] In some embodiments of the present invention, the multiple sensors include: an inertial measurement unit for collecting the movement information of the blind person, where the movement information includes orientation; a global positioning module for collecting the current position information of the blind person; a depth camera and multiple ultrasonic sensors for collecting the environmental information of the space where the blind person is located, wherein the depth camera simultaneously collects a color map and a depth map that are spatially registered. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The following further describes the embodiments of the present invention with reference to the accompanying drawings, where:

[0019] Figure 1 FIG. is a schematic diagram of the physical object of the blind guiding device based on multiple sensors according to the embodiment of the present invention;

[0020] Figure 2 FIG. is a schematic diagram of the annotation result of pixel-level annotation of road data in the dataset according to the embodiment of the present invention;

[0021] Figure 3 FIG. is a schematic diagram of the 3D point cloud data according to the embodiment of the present invention;

[0022] Figure 4 FIG. is a visualization image of the background of the blind guiding device based on multiple sensors according to the embodiment of the present invention;

[0023] Figure 5 FIG. is a schematic diagram of the result of an experiment according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below through specific embodiments with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0025] As mentioned in the background art section, the functions of current existing technologies are relatively single, capable of only performing some simple tasks, with low adaptability to complex outdoor environments and limited application scenarios. The present invention collects the current position information, movement information, and various environmental information of the blind through multiple sensors; and performs path planning based on the destination, current position information, movement information, and various environmental information of the blind to obtain the optimal path. Among them, multiple deep learning models are used to determine the feasible areas and obstacles for path planning according to the corresponding environmental information respectively; and the optimal path is converted into an action instruction to guide the blind to act. Thus, path planning is carried out by means of the information collected by multiple sensors and the feasible areas and obstacles determined by multiple deep learning models to better assist the blind in moving in complex outdoor environments.

[0026] According to an embodiment of the present invention, the present invention provides a blind guiding device based on multiple sensors. The following will be described separately from three parts: device hardware, device software, and human-computer interaction.

[0027] (I) Device Hardware

[0028] According to an embodiment of the present invention, in order to perform better path planning, more comprehensive information needs to be collected. The present invention uses multiple sensors to collect the current position information, movement information, and various environmental information of the blind.

[0029] Figure 1 Fig. shows a physical schematic diagram of a blind guiding device based on multiple sensors according to an embodiment of the present invention, including multiple sensors, a processing module, a human-computer interaction module, a battery, or a combination thereof. The battery is used to supply power to the multiple sensors, the processing module, and the human-computer interaction module.

[0030] Preferably, the multiple sensors include a global positioning module, an inertial measurement unit, a depth camera, and multiple ultrasonic sensors.

[0031] According to an embodiment of the present invention, the global positioning module is used to collect the current position information of the blind. For example, the global positioning module uses a GPS module, a Beidou positioning module, a Galileo positioning module, or a combination thereof.

[0032] According to an embodiment of the present invention, the inertial measurement unit (IMU module) is used to collect the movement information of the blind, and the movement information includes speed and / or orientation. For example, the IMU module is used to obtain information such as the speed and orientation of the blind (user) to provide assistance to the path planning module.

[0033] According to an embodiment of the present invention, the depth camera is used to collect the environmental information of the space where the blind is located, where the environmental information includes a color map and a depth map registered in space. See Figure 1, the depth camera and the inertial measurement unit can be integrated into one device. For example, the depth camera uses a RealSense camera, in which the RGB camera captures high-definition color map information (i.e., the color map), the depth camera captures the depth information of the environment (i.e., the depth map), and the IMU module captures movement information. The depth camera can be worn on the head of the blind person (e.g., worn on the head through a helmet), chest or waist to sense the changing environmental information directly in front of the blind person.

[0034] According to an embodiment of the present invention, multiple ultrasonic sensors are used to measure the distance between the blind person and the obstacle. For example, 3-7 ultrasonic sensors are set. In one example, see Figure 1 , five ultrasonic sensors are installed on the helmet.

[0035] According to an embodiment of the present invention, the processing module (corresponding to the path planning device, in the present invention, the two terms can be equivalently replaced) is used to control all sensors, analyze environmental information, issue control instructions, and perform path planning. The selection of the processing module needs to comprehensively consider various requirements such as the performance and battery life of the system. As an example, Nvidia Jetson AGX Xavier is selected as the processing module.

[0036] According to an embodiment of the present invention, the blind guiding device based on multiple sensors further includes a human-computer interaction module. The human-computer interaction module can be, for example, a pair of headphones.

[0037] According to an embodiment of the present invention, the blind guiding device based on multiple sensors further includes a power supply. For example, a mobile battery is selected to supply power to various sensors, the processing module, and the human-computer interaction module.

[0038] (2) Device Software

[0039] The present invention includes device software related to environmental perception and path planning.

[0040] (1) Environmental Perception

[0041] According to an embodiment of the present invention, environmental perception refers to analyzing the environmental information collected by multiple sensors through an algorithm, including feasible area recognition and target detection. Preferably, the processing module performs feasible area recognition and target detection based on the video stream formed by the color map captured by the depth camera.

[0042] Feasible region detection is a core technology in environmental perception tasks, which helps blind people walk safely and independently in complex environments. In the blind guidance scenario, it is necessary to distinguish blind paths, sidewalks, roadways, and zebra crossings. When blind people walk outdoors, they should preferably choose blind paths, followed by sidewalks; when passing through intersections, they should walk on zebra crossings. Feasible region detection can be regarded as an image segmentation task; commonly used lightweight image segmentation models include: ENet, ShuffleNet, LiteSeg, etc. However, due to reasons such as accuracy and running speed, these models cannot be directly applied to the feasible region analysis task. Therefore, a lightweight image segmentation model (corresponding to the first deep learning model) is needed to complete this task. According to an embodiment of the present invention, based on the characteristics of the intelligent blind guidance project, the inventor selects ESPNetV2 as the basic model of the lightweight segmentation model and improves it. By adding a Feature Pyramid Network (FPN) at the end of the feature extraction model of the ESPNetV2 model, it helps the model perform multi-dimensional feature fusion, enabling the model to not only extract low-dimensional detailed information but also obtain high-dimensional semantic information, thereby improving the segmentation accuracy of the model. Among them, adding the feature pyramid module at the end of the feature extraction module of the original ESPNetV2 model means adding the feature pyramid module between the group convolution layer and the global average pooling layer of the original ESPNetV2 model; thus, the structure of the improved ESPNetV2 model is obtained. The output of the group convolution layer of the improved ESPNetV2 model is used as the input of the global average pooling layer (Global avg.pool) after being processed by the feature pyramid module.

[0043] According to an embodiment of the present invention, the improved ESPNetV2 model is trained as follows: The collected road-related dataset is used to train the improved ESPNetV2 model to segment and identify the feasible regions in the sample images in the road dataset, and the feasible region recognition results are output. Among them, the road-related dataset includes sample images and feasible region segmentation labels, and the feasible region segmentation labels include annotation values indicating that the corresponding regions of the sample images are blind paths, sidewalks, roadways, zebra crossings, and backgrounds; The cross-entropy is used as the loss function to calculate the loss value of segmenting the feasible region according to the output feasible region recognition results and the feasible region segmentation labels, and the parameters of the improved ESPNetV2 model are updated according to the loss value of segmenting the feasible region for training and system integration. Preferably, in order to make the algorithm better applicable to outdoor scenarios, the applicant has collected and annotated a large amount of road segmentation data for the feasible region analysis task. Pixel-level annotation is performed on the collected road data. Among them, the blind path, sidewalk, roadway, and zebra crossing are regarded as the foreground, and the rest are regarded as the background; The four foregrounds are respectively annotated in different annotation channels (for example, if a certain pixel belongs to a part of the blind path, the label value corresponding to this pixel is annotated as 1 in the annotation channel corresponding to the blind path, and the label values corresponding to this pixel in the other annotation channels are annotated as 0). The pixels other than the four foregrounds are annotated in the same annotation channel (that is, assuming that a certain pixel belongs to a part of the sky or a pedestrian, the label values corresponding to this pixel in the annotation channels corresponding to the four foregrounds are all annotated as 0, and the label value corresponding to this pixel in the annotation channel corresponding to the background is annotated as 1). The visualized annotation results are as Figure 2 shown, where Figure 2 the upper half is the sample image, and the lower half is the annotation result corresponding to the corresponding sample image. During training, the corresponding optimization algorithm and hyperparameters can be selected according to needs and / or experience. For example, the Adam optimization algorithm can be selected to optimize the model; It is preset to train for 200 epochs, and the initial learning rate of the model is set to 0.01 and adjusted in a cosine periodic change manner. The technical solution of this embodiment can at least achieve the following beneficial technical effects: After training the improved ESPNetV2 model, the obtained model can enhance the modeling ability of the segmentation algorithm for details, improve the segmentation accuracy of the model, more accurately segment the feasible region, thereby enhancing the adaptability to complex outdoor environments and better assisting the blind to walk outdoors.

[0044] It should be understood that in addition to the improved ESPNetV2 model described above, the first deep learning model can also be implemented using other lightweight deep learning models. For example, the ESPNetV1 model or the ESPNetV2 model can be directly used to achieve an effect close to that of the improved ESPNetV2 model. The training method can refer to the foregoing embodiments of the improved ESPNetV2 model.

[0045] When the blind population walks in a complex outdoor environment, in order to ensure the safety of users, the guiding device needs to detect the surrounding environment in real time. Therefore, the present invention needs to detect some common targets in the outdoor environment, including pedestrians, bicycles, cars, traffic signs, trees, traffic lights, and ground obstacles (such as flower beds, piles, trash cans, etc.). According to an embodiment of the present invention, the second deep learning model uses the color image collected by the depth camera to perform target detection on the space where the blind person is located, and obtains the target detection result. Preferably, the second deep learning model can be implemented using the improved YOLOv5 model, wherein the improved YOLOv5 model replaces the 3x3 ordinary convolution in the original YOLOv5 model with a depthwise separable convolution. The present invention adopts the improved YOLOv5 model, which can meet the requirements of the guiding device for the real-time performance and accuracy of the target detection algorithm, and further reduces the number of model parameters and the amount of computation by replacing the ordinary convolution in the YOLOv5 model with a depthwise separable convolution, realizing a lightweight target detection algorithm suitable for outdoor scenarios.

[0046] According to an embodiment of the present invention, the second deep learning model is trained as follows: The improved YOLOv5 model is trained using the COCO2017 dataset to perform target detection on the sample images in the dataset, and the target detection result is output; the total loss of the target detection is calculated according to the output target detection result and the target detection label, and the parameters of the improved YOLOv5 are updated according to the total loss of the target detection, wherein the total loss of the target detection includes a classification prediction sub-loss and a detection box prediction sub-loss. Preferably, the classification prediction sub-loss uses the binary cross-entropy loss function. That is, the binary cross-entropy loss function is used to optimize the classifier. Preferably, the detection box prediction sub-loss uses the GIoU loss function, that is, the GIoU loss function is used to optimize the detection box (bounding box) regressor. During training, the corresponding optimization algorithm and hyperparameters can be selected according to needs and / or experience. For example, it can be preset to train for 100 epochs, the initial learning rate of the model is set to 0.01 and the learning rate is adjusted in a cosine cycle change manner. Additionally, optionally, it is also feasible to implement the second deep learning model using the existing YOLOv5 model, and its training method can refer to the embodiment of training the improved YOLOv5 model.

[0047] According to an embodiment of the present invention, for traffic lights, when a blind person crosses the street, it is necessary to detect the color of the traffic light to assist the blind person in crossing the street. Preferably, after the traffic light is detected by the target detection algorithm, the color of the traffic light can be recognized to obtain the color of the traffic light. Alternatively, image samples and target detection labels of traffic lights in the red light, yellow light, and green light states can be added to the original COCO2017 dataset to obtain an improved COCO2017 dataset for training an improved YOLOv5 model to recognize the color of traffic lights. Or, an existing text recognition model can be set in the blind guide device to recognize the color of the traffic light and the value of its countdown.

[0048] Obstacle avoidance is a huge challenge that the blind population must face when walking outdoors. When a user walks independently on the sidewalk, they need to avoid different types of obstacles such as bicycles and pedestrians. How to determine the distance to these obstacles is the core of the obstacle avoidance task. Existing blind guide systems usually use ultrasonic sensors or depth cameras to detect the distance between the user and the obstacles; however, both of these methods have certain defects. The detection range of the ultrasonic sensor is small, and the depth camera is easily affected by the environment. To solve this problem, in the present invention, an ultrasonic sensor and a depth camera are used simultaneously, and a depth map completion algorithm is designed to enhance the ability of the blind guide device to obtain depth information.

[0049] According to an embodiment of the present invention, the processing module is configured to: perform target detection on the space where the blind person is located using a second deep learning model based on the color image collected by the depth camera to obtain a target detection result; determine a first detection result of the obstacle using the target detection result and the depth map collected by the depth camera. Preferably, the first detection result includes: an obstacle determined based on the target detection result of the image and an obstacle determined by detecting a sudden change in the depth value according to the depth map. The depth map used in the obstacle determined by detecting the sudden change in the depth value according to the depth map is a depth map obtained by complementing the associated depth map using a third deep learning model based on the color image collected by the depth camera. Preferably, the third deep learning model is implemented using the PENet model and is trained as follows: The PENet model is trained using the KITTI-Depth dataset to complement the associated depth map based on the color image, the mean square error loss function is used to calculate the loss value based on the complemented depth map and the depth label, and the parameters of the PENet model are updated according to the loss value. The KITTI-Depth data is a 3D point cloud data, including RGB images and corresponding 3D point cloud information, specifically as Figure 3 shown; where Figure 3The upper part is the original RGB image, and the lower part is the point cloud image after visualization. During training, corresponding optimization algorithms and hyperparameters can be selected according to needs and / or experience. For example, the Adam optimization algorithm can be selected to optimize the model. It is preset to train for 100 epochs, the initial learning rate of the algorithm is 0.00, and the mean square error (MSE) is used as the loss function. The technical solution of this embodiment can at least achieve the following beneficial technical effects: The present invention aims at the problem that some detection areas of a depth camera fail in environments such as strong light. Based on the PENet model, a lightweight depth map completion model is trained, thereby alleviating the impact on the depth camera, improving the accuracy of depth value detection, and better assisting blind people to travel.

[0050] According to an embodiment of the present invention, the first detection result includes obstacles determined by using the target detection result and the depth map collected by the depth camera. Preferably, the obstacles determined by using the target detection result include pedestrians, bicycles, motorcycles, cars, traffic signs, and other ground obstacles. Or, the obstacles determined by using the target detection result include: pedestrians, bicycles, cars, traffic signs, trees, traffic lights, and ground obstacles. In addition, for redundant identification and to identify some unknown obstacle types, preferably, the obstacles determined by using the depth map collected by the depth camera are obstacles determined by detecting depth value mutations in the depth map. For example, since target detection is for some common objects, and the types of objects involved outdoors are numerous and may not be in the preset categories. For some unknown obstacles, such as scattered bricks, waste, construction waste piles, or deep pits that appear on the street, they can be obstacles determined by detecting depth value mutations in the depth map. For example, the depth values of normal ground change smoothly. If there are suddenly some areas where the depth values increase or decrease significantly, it can be determined as an obstacle. Since the positions of the pixels in the depth map and the color map correspond to each other, the positions of the obstacles determined by using the target detection result and the depth map collected by the depth camera can be conveniently obtained. Thus, the missed detection of obstacles can be reduced, and the travel of blind people can be better assisted.

[0051] According to an embodiment of the present invention, since the walking speed of the blind group is relatively slow (about 0.2 m / s or even slower), the change of the feasible region is slow, but the change of surrounding targets is fast. To reduce the overall amount of calculated data while ensuring safety, according to an embodiment of the present invention, the processing module may further be configured to: within a unit time, the number of frames of the color map collected by the depth camera used for updating the feasible region recognition result is less than the number of frames of the color map collected by the depth camera used for updating the target detection result. That is: for the feasible region recognition, the present invention does not need to analyze each frame of data collected by the depth camera; for example, assuming that the parameter for collecting the color map by the RealSense camera is 30 frames per second (FPS); in the present invention, the processing module may be configured to: evenly and at intervals run the first deep learning model every second to perform feasible region recognition on 2 frames (or 4 frames, 6 frames, 10 frames, etc.) of the color map collected by the depth camera, and run the second deep learning model every second to perform target detection on each frame (or detect 20 frames, 25 frames, 28 frames, etc. per second) of the color map collected by the depth camera.

[0052] (2) Path planning

[0053] Based on the results of environmental perception according to the foregoing embodiments, path planning can be performed to assist the blind in traveling.

[0054] According to an embodiment of the present invention, path planning includes global path planning and local path planning.

[0055] According to an embodiment of the present invention, the processing module is configured to: perform global path planning based on the destination and current position information of the blind to obtain a global path; use the first deep learning model to perform feasible region recognition on the space where the blind is located according to the color map collected by the depth camera to obtain a feasible region recognition result; use the second deep learning model to perform target detection on the space where the blind is located according to the color map collected by the depth camera to obtain a target detection result; use the target detection result and the depth map collected by the depth camera to determine the first detection result of obstacles and use a plurality of ultrasonic sensors to determine the second detection result of obstacles; map the space corresponding to the color map to a preset walking space grid, and determine the cost scaling factor of the corresponding grid in the walking space grid according to the feasible region recognition result, the first detection result, and the second detection result; use a path planning algorithm to perform local path planning based on the global path and the cost scaling factor of the grid to obtain an optimal path for guiding the blind to walk. The technical solution of this embodiment can at least achieve the following beneficial technical effects: The present invention realizes the feasible region recognition result, the first detection result, and the second detection result by means of a deep learning model, and for the safety of the blind's travel, uses the corresponding cost scaling factor to scale the cost of the grid, thereby assisting the blind to walk more safely and efficiently.

[0056] According to an embodiment of the present invention, global path planning needs to be realized with the cooperation of a navigation application and a global positioning module. For example, it is completed by using the Baidu Map application (Baidu-Map API) and the GPS module in cooperation. The blind person can input the destination through the voice assistant of Baidu Map. The Baidu Map application plans the global path according to the destination of the blind person and the current position information measured by the GPS module, and indicates the local location to which the local path leads through the global path.

[0057] In the problem of the blind person's travel, local path planning is more complex. The guiding device should not only ensure the safety of the user but also select the optimal walking route. To solve this problem, the present invention regards local path planning as a special optimal path search problem. Common path search algorithms include: depth-first search algorithm, breadth-first search algorithm, Dijstra algorithm, etc. Among them, the depth-first search algorithm and the breadth-first search algorithm use a graph-based path search model, but the running efficiency is poor. The Dijstra algorithm is a heuristic path search algorithm. This method needs to calculate the total cost from each node to the starting point during the running process, and then traverse these nodes to select an optimal node as the path for the next movement; repeat the loop until reaching the end point. However, it is difficult for the Dijstra algorithm to be applicable to an environment with obstacles. According to an embodiment of the present invention, the present invention selects the optimized A-Star algorithm for local path planning.

[0058] The original A-Star algorithm calculates the comprehensive cost of each grid corresponding node according to the following formula:

[0059] f(n)=g(n)+h(n);

[0060] Among them, f(n) is the comprehensive cost of node n, g(n) is the cost of node n from the starting point, and h(n) is the estimated cost of node n from the end point. When the original A-Star algorithm runs, all nodes are regarded as unit size, and thus g(n), h(n), and f(n) are directly calculated. Finally, the node with the smallest comprehensive cost needs to be considered as the direction of the next movement.

[0061] To adapt to the travel situation of the blind person, according to an embodiment of the present invention, on the basis of the A-Star algorithm, the present invention adds a cost scaling factor to different grids. Preferably, when determining the cost scaling factor of the corresponding grid in the walking space grid, the relative degree by which the cost of the corresponding grid is reduced is as follows: blind path ≥ zebra crossing > sidewalk > roadway > obstacle, where the obstacles include pedestrians, bicycles, motorcycles, cars, traffic signs, and other ground obstacles.

[0062] According to an embodiment of the present invention, when determining the cost scaling factor of the corresponding grid in the walking space grid, the basis for whether the corresponding grid is an obstacle is as follows: for a long-distance obstacle whose distance is greater than a predetermined distance threshold, the priority of the first detection result is higher than that of the second detection result; for a short-distance obstacle whose distance is less than or equal to the predetermined distance threshold, the priority of the second detection result is higher than that of the first detection result. For example, assuming that the distance threshold is set to 1 m, if a short-distance obstacle within 1 m is detected according to the second detection result, the cost scaling factor of the corresponding grid will be considered to be set according to the second detection result; thus, the obstacle can be avoided. If the obstacle is outside 1 m, the second detection result will not be referred to, and the first detection result will be taken as the standard. Refer to Figure 1 , according to an embodiment of the present invention, multiple ultrasonic sensors are arranged at a certain angular interval to detect obstacles in different amplitude ranges in the front. For example, assuming that 5 ultrasonic sensors are arranged, if the set distance threshold is set to 1 m, the 5 sensors respectively correspond to two rows of grids on the side close to the blind person in the walking space network, and these two rows of grids respectively correspond to the 5 ultrasonic sensors according to the arranged angles (for example, the two rows of grids are divided into 5 parts from left to right and correspond to the 5 ultrasonic sensors arranged from left to right on the helmet in sequence). The technical solution of this embodiment can at least achieve the following beneficial technical effects: according to the different characteristics of the ultrasonic sensor and the depth camera, the present invention preferentially refers to the results detected by the corresponding sensors under different distance conditions, so as to more accurately determine the obstacle.

[0063] According to an embodiment of the present invention, in the optimized A-Star algorithm, the grids corresponding to the blind path, zebra crossing, sidewalk, and roadway are defined as feasible regions, and the grids corresponding to obstacles are defined as infeasible regions; the cost scaling factor of the grid representing the blind path is 3, the cost scaling factor of the grid representing the sidewalk is 2, the cost scaling factor of the grid representing the roadway is set to 1, and the cost scaling factor of the grid representing the zebra crossing is 3. Of course, it should be understood that the result of being recognized as an obstacle has a higher priority than the result of being recognized as a feasible region. For example: if a grid is recognized as both a feasible region and an obstacle at the same time, it is marked as an obstacle. The optimized A-Star algorithm calculates the comprehensive cost of each grid corresponding node according to the following formula: F(n)=G(n)+H(n), where F(n) represents the weighted comprehensive cost, G(n) represents the weighted cost of node n from the starting point, , H(n) represents the weighted estimated cost of node n from the end point, , where w represents the cost scaling factor of the grid, g(n) is the cost from node n to the starting point, and h(n) is the estimated cost from node n to the ending point. During local path planning, the optimized A-Star algorithm uses F(n), G(n), and H(n) for planning, and the rest of the logic can remain unchanged. The technical solution of this embodiment can at least achieve the following beneficial technical effects: By setting the cost scaling factor of the grid, the path planning algorithm can consider the characteristics of the blind from the aspect of walking safety, enabling them to walk outside more safely and efficiently.

[0064] Alternatively, according to an optional embodiment of the present invention, the cost scaling factor of the grid representing the blind path is 1, the cost scaling factor of the grid representing the sidewalk is 2, the cost scaling factor of the grid representing the roadway is set to 3, and the cost scaling factor of the grid representing the zebra crossing is 1. In this embodiment, the optimized A-Star algorithm calculates the comprehensive cost of each grid corresponding node according to the following formula: F(n)=G(n)+H(n), where F(n) represents the weighted comprehensive cost, G(n) represents the weighted cost from node n to the starting point, G(n)=w×g(n), H(n) represents the weighted estimated cost from node n to the ending point, H(n)=w×h(n), w represents the cost scaling factor of the grid, g(n) is the cost from node n to the starting point, and h(n) is the estimated cost from node n to the ending point.

[0065] According to an embodiment of the present invention, the color map can be directly mapped or adjusted to a specific size before mapping. For example, assume that the image size of the color map is adjusted to 512*512, and then the entire image is evenly divided into 32x32 = 1024 grids (corresponding to the grid); then, according to the result of environmental perception, different cost scaling factors are assigned to each grid; a local optimal path is planned according to the calculation result of the A-Star algorithm; then, according to the orientation information feedback by the inertial measurement unit, the blind person is guided to adjust the orientation so that the blind person can face the node to be moved forward. The technical solution of this embodiment can at least achieve the following beneficial technical effects: By comprehensively using algorithms corresponding to multiple models to determine the result, for example, using an object detection algorithm to locate obstacles such as pedestrians and vehicles and traffic lights ahead; using a feasible area recognition algorithm to sense the specific position of feasible areas such as zebra crossings, setting different cost scaling factors to improve the A-Star algorithm, and obtaining a planning model for determining the local optimal path, which can help users find the optimal moving route in real time; enabling the guiding device to better and more safely guide users to complete complex tasks of outdoor walking (such as passing through intersections).

[0066] (III) Human-computer interaction

[0067] According to an embodiment of the present invention, the present invention uses voice for human-computer interaction. The blind guide device synthesizes the results of environmental perception and path planning, generates a series of action instructions, and then conveys the action instructions to the blind person in the form of voice through headphones to guide the blind person to act. According to an embodiment of the present invention, the voice prompt module includes a plurality of preset voice instruction combinations, and the voice instruction combinations include: go straight, wait, turn left, turn right, stop, obstacle ahead, obstacle on the left, obstacle on the right, intersection ahead, red light, green light or a combination thereof. According to an embodiment of the present invention, the voice instruction combination can also refer to more detailed environmental information and be set more specifically. For example, the voice instruction combination includes: go straight for x meters, wait for x seconds, turn left by x degrees (some common angles can be preset at intervals, such as turning left 30°, 45°, 90°), turn right by x degrees (for example, turn right 30°, 45°, 90°), stop, accelerate, decelerate, obstacle ahead, obstacle on the left, obstacle on the right, step up ahead, step down ahead, intersection ahead, red light for x seconds, green light for x seconds. Preferably, the action information includes the current speed of the blind person, and it can be determined whether the blind person needs to accelerate or decelerate according to the current speed of the blind person and the environmental information.

[0068] The blind person makes corresponding actions according to the issued voice instructions, such as moving forward, turning left, etc. The inertial measurement unit will continuously monitor the behavior of the blind person. The processing module is configured to calculate and analyze the orientation and speed information of the blind person in real time, judge whether it matches the optimal path, and further adjust the actions of the blind person. As the blind person moves, the optimal path is continuously adjusted according to the updated current position information, action information and various environmental information until the destination of the blind person is reached.

[0069] For the sake of visualization, refer to Figure 4 , which shows the recognition results of the feasible area, target detection results, depth map, global path and IMU information (corresponding to the action information, including orientation and speed) in the background visualization during the navigation process in the experiment.

[0070] In addition, the applicant also conducted a simple experiment on the present invention, and the experimental results are as follows:

[0071] (1) Image segmentation experiment

[0072] In order to verify the performance of the proposed feasible area recognition algorithm (corresponding to the first deep learning model, using the improved ESPNetV2 model), the proposed algorithm was compared with the existing lightweight image segmentation algorithm on the collected road dataset. Specifically, mIOU and the model calculation amount were compared, and the experimental results are shown in Table 1.

[0073] Table 1

[0074] Model mIOU (%) FLOPs (B) ENet 80.15 3.4 MobileNet 84.23 12.6 Improved ESPNetV2 Model 90.46 2.4

[0075] The experimental results show that the improved ESPNetV2 model of the present invention has great advantages in terms of accuracy, and the computational complexity is much smaller than that of ENet and MobileNet. This result indicates that the improved ESPNetV2 model of the present invention is more suitable for the road segmentation task.

[0076] (2) Location navigation task

[0077] To verify the performance of the present invention, a specific location navigation task was also designed. The participants needed to walk from one location to another according to the guidance of this device; that is: guiding the user to walk from Figure 5 point A shown in the figure to point B. The test path is about 80 meters long. There are a total of 5 participants, among which P3 - P5 used this device, Ps represents participants with normal vision, and Pw represents participants who only used a white cane. The results are as Figure 5 shown. The time and speed of the test users are shown in Table 2.

[0078] Table 2

[0079] User P3 P4 P5 Pw Ps Success Rate 0.4 0.8 1.0 0.4 1.0 Average Speed 0.14 0.21 0.27 0.08 0.75

[0080] The experimental results show that the walking paths of the users using this device are similar to those of the participants with normal vision, much better than those of the participants who only used a white cane, and the walking path of participant P5 is more efficient. This experimental result demonstrates the effectiveness of this device.

[0081] (3) Passing through the obstacle area

[0082] The applicant also designed a group of experiments for passing through the obstacle area. In this task, the test users should pass through the obstacle area at the fastest speed while reducing the contact with the obstacles. There are a total of 7 participants in this experiment, among which P1 - P5 used this device, Ps represents participants with normal vision, and Pw represents participants who only used a white cane. The passing time and the number of times of touching the obstacles of the testers were recorded. The results are shown in Table 3.

[0083] Table 3

[0084] User P1 P2 P3 P4 P5 Pw Ps Success Rate 0.5 0.7 0.7 0.6 0.9 0.6 1.0 Average Touch 2.4 3.0 2.2 3.6 2.0 6.0 0 Average Time 183 192 173 181 166 228 31

[0085] It can be seen from the results that the performance of this device is far better than that of the white cane device. Among them, the success rate of user P5 in passing through the obstacle area is 0.9, and the number of times of touching the obstacles is also small, indicating that this device plays a certain role in the obstacle avoidance task.

[0086] (4) Experiment for passing through intersections

[0087] This device integrates a variety of deep learning algorithms and can help users complete complex tasks of passing through intersections. To test this function, an experiment of passing through intersections was also designed. There were 6 participants in this experiment, among which P1 - P5 used this device, and Ps represents participants with normal vision. We recorded the passing time and success rate of the testers, and the experimental results are shown in Table 4.

[0088] Table 4

[0089] User P1 P2 P3 P4 P5 Ps Success Rate 0.6 0.4 0.6 0.8 1.0 1.0 Average Time 51.9 55.7 49.6 51.8 41.9 9.7

[0090] It can be seen from the results that with the help of this device, the highest success rate of users passing through intersections is 100%, and the lowest is 0.4. The results show that this device has the ability to handle complex task navigation tasks. After multiple trainings and adaptations by users, better experimental results can be obtained.

[0091] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even the order can be changed, as long as the required functions can be achieved.

[0092] The present invention can be a system, a method, and / or a computer program product. The computer program product can include a computer - readable storage medium having computer - readable program instructions thereon for causing a processor to implement various aspects of the present invention.

[0093] The computer - readable storage medium can be a tangible device that retains and stores instructions for use by an instruction - execution device. The computer - readable storage medium can, for example, include but is not limited to an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non - exhaustive list) of the computer - readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read - only memory (ROM), an erasable programmable read - only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read - only memory (CD - ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing.

[0094] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A path planning device for guiding the blind, characterized in that, the path planning device is configured to: perform global path planning based on the destination and current position information of the blind person to obtain a global path; use a first deep learning model to identify the feasible area of the space where the blind person is located according to the color map collected by the depth camera to obtain a feasible area identification result; use a second deep learning model to perform object detection on the space where the blind person is located according to the color map collected by the depth camera to obtain an object detection result; use the object detection result and the depth map collected by the depth camera to determine the first detection result of obstacles and use multiple ultrasonic sensors to determine the second detection result of obstacles; map the space corresponding to the color map to a preset walking space grid, and determine the cost scaling factor of the corresponding grid in the walking space grid according to the feasible area identification result, the first detection result and the second detection result. When determining the cost scaling factor of the corresponding grid in the walking space grid, the relative degree of reduction of the cost of the corresponding grid is as follows: blind path ≥ zebra crossing > sidewalk > roadway > obstacle; use a path planning algorithm to perform local path planning according to the global path and the cost scaling factor of the grid to obtain an optimal path for guiding the blind person to walk.

2. The path planning device according to claim 1, characterized in that, the first deep learning model is implemented by using an improved ESPNetV2 model, wherein the improved ESPNetV2 model is obtained by adding a feature pyramid module at the end of the feature extraction module of the original ESPNetV2 model.

3. The path planning device according to claim 2, characterized in that, the improved ESPNetV2 model is trained as follows: use the collected road-related data set to train the improved ESPNetV2 model to segment and identify the feasible area in the sample image in the road data set, and output the feasible area identification result. The road-related data set includes sample images and feasible area segmentation labels, and the feasible area segmentation labels include annotation values indicating that the corresponding areas of the sample images are blind paths, sidewalks, roadways, zebra crossings, and backgrounds; use cross entropy as the loss function to calculate the loss value of segmenting the feasible area according to the output feasible area identification result and the feasible area segmentation label, and update the parameters of the improved ESPNetV2 model according to the loss value of segmenting the feasible area.

4. The path planning device according to claim 1, characterized in that, the second deep learning model is implemented by using an improved YOLOv5 model, wherein the improved YOLOv5 model replaces the 3x3 ordinary convolution in the original YOLOv5 model with a depthwise separable convolution.

5. The path planning device according to claim 4, characterized in that, the second deep learning model is trained as follows: use the COCO2017 data set to train the improved YOLOv5 model to perform object detection on the sample images in the data set and output the object detection result; Calculate the total loss of object detection based on the output object detection results and object detection labels, and update the parameters of the improved YOLOv5 according to the total loss of object detection, where the total loss of object detection includes a classification prediction sub-loss and a detection box prediction sub-loss.

6. The path planning device according to claim 1, wherein, The first detection result includes: obstacles determined based on the object detection result of the image and obstacles determined by detecting the mutation of the depth value according to the depth map.

7. The path planning device according to claim 6, wherein, The depth map used for determining the obstacles by detecting the mutation of the depth value according to the depth map is the depth map after the depth map associated with the color map collected by the depth camera is completed by using a third deep learning model.

8. The path planning device according to claim 7, wherein, The third deep learning model is implemented using the PENet model and is trained as follows: Use the KITTI-Depth dataset to train the PENet model to complete the depth map associated with the color map, Calculate the loss value according to the completed depth map and the depth label using the mean square error loss function, and update the parameters of the PENet model according to the loss value.

9. The path planning device according to claim 1, wherein, Obstacles include pedestrians, bicycles, cars, traffic signs, and other ground obstacles.

10. The path planning device according to claim 9, wherein, When determining the cost scaling factor of the corresponding grid in the walking space grid, the basis for whether the corresponding grid is an obstacle is as follows: For long-distance obstacles with a distance greater than a predetermined distance threshold, the priority of the second detection result is higher than that of the first detection result; For short-distance obstacles with a distance less than or equal to the predetermined distance threshold, the priority of the first detection result is higher than that of the second detection result.

11. The path planning device according to any one of claims 1-10, wherein, The path planning device is configured to: the number of frames of the color map collected by the depth camera used to update the feasible area recognition result per unit time is less than the number of frames of the color map collected by the depth camera used to update the object detection result per unit time.

12. A multi-sensor based blind guiding device, wherein, Comprising: Multiple sensors for collecting the current position information, movement information, and various environmental information of the blind; The path planning device according to any one of claims 1-11; and A human-computer interaction module for converting the optimal path into an action instruction for guiding the blind to move.

13. The multi-sensor based blind guiding device according to claim 12, wherein, The multiple sensors include: An inertial measurement unit for collecting the movement information of the blind, and the movement information includes the orientation; A global positioning module for collecting the current position information of the blind; A depth camera and multiple ultrasonic sensors for collecting the environmental information of the space where the blind is located, wherein the depth camera simultaneously collects a color map and a depth map that are spatially registered.