Automatic driving obstacle avoidance method, computer readable storage medium, controller and vehicle

Through multimodal sensors and pre-trained object detection models, combined with downsampling, clustering and feature extraction processing, the problem of obstacle detection difficulty of autonomous vehicles in complex environments is solved, and higher obstacle detection accuracy and obstacle avoidance reliability are achieved.

CN120028807APending Publication Date: 2025-05-23ZHEJIANG GEELY HLDG GRP CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510094499.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Autonomous driving vehicles face huge challenges in complex and changing environments, especially because factors such as ambient light and camera angle affect image quality, which increases the difficulty of obstacle detection.

Method used

Multimodal sensors (camera, lidar, millimeter wave radar) are used to collect data, use pre-trained target detection models to detect obstacles, and process multimodal data through downsampling, clustering and feature extraction, and fuse information from different sensors to improve the accuracy of obstacle information.

Benefits of technology

It improves the accuracy of obstacle detection and the reliability of obstacle avoidance, can accurately identify and avoid obstacles in complex environments, and enhances the safety of autonomous driving vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120028807A_ABST
    Figure CN120028807A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving obstacle avoidance method, a computer readable storage medium, a controller and a vehicle, and the method comprises the steps: obtaining multi-modal data collected by a multi-modal sensor, the multi-modal data at least comprising a camera image, laser radar data and millimeter wave radar data; the method comprises the following steps: performing obstacle detection on a camera image by using a pre-trained target detection model, performing clustering operation on laser radar data, and performing feature extraction on millimeter wave radar data to obtain an obstacle prediction set, a clustering set and a feature extraction set; determining initial obstacle information according to the obstacle prediction set and the clustering set, and determining target obstacle information according to the initial obstacle information and the feature extraction set; and making an obstacle avoidance decision according to the target obstacle information and the current state of the vehicle, and controlling the vehicle to perform obstacle avoidance driving according to the obstacle avoidance decision. The method has the advantages of accurate obstacle detection and reliable obstacle avoidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to an autonomous driving obstacle avoidance method and a computer-readable storage medium, a controller, and a vehicle. Background Art

[0002] With the rapid development of science and technology, autonomous driving technology is gradually becoming a reality. However, in practical applications, autonomous vehicles face complex and changing environmental challenges. There are many types of obstacles on the road, including vehicles, pedestrians, road signs, etc., which poses a huge challenge to the obstacle avoidance algorithm of autonomous vehicles. At the same time, factors such as ambient lighting and camera angle will also affect image quality, further increasing the difficulty of obstacle detection. Summary of the invention

[0003] The present invention aims to solve at least one of the technical problems in the related art to a certain extent. To this end, one object of the present invention is to propose an automatic driving obstacle avoidance method, which has the advantages of accurate obstacle detection and reliable obstacle avoidance.

[0004] A second object of the present invention is to provide a computer-readable storage medium.

[0005] The third object of the present invention is to provide a controller.

[0006] A fourth object of the present invention is to provide a vehicle.

[0007] To achieve the above-mentioned purpose, an embodiment of the first aspect of the present invention proposes an automatic driving obstacle avoidance method for a vehicle, wherein a multimodal sensor is provided on the vehicle, and the method comprises: obtaining multimodal data collected by the multimodal sensor, wherein the multimodal data comprises at least a camera image, a lidar data and a millimeter-wave radar data; while performing obstacle detection on the camera image using a pre-trained target detection model, downsampling and clustering operations are performed on the lidar data, and feature extraction is performed on the millimeter-wave radar data to obtain an obstacle prediction set, a clustering set and a feature extraction set; determining initial obstacle information according to the obstacle prediction set and the clustering set, and determining target obstacle information according to the initial obstacle information and the feature extraction set; making an obstacle avoidance decision according to the target obstacle information and the current state of the vehicle, and controlling the vehicle to perform obstacle avoidance driving according to the obstacle avoidance decision.

[0008] According to the automatic driving obstacle avoidance method of the embodiment of the present invention, after obstacle detection is performed on the acquired multimodal data, the obstacle detection results based on the camera image, the obstacle clustering results based on the lidar data, and the obstacle feature extraction results based on the millimeter wave radar data are fused to improve the accuracy of the obtained target obstacle information, and an obstacle avoidance decision is made according to the target obstacle information and the current state of the vehicle, and the vehicle is controlled to perform obstacle avoidance driving according to the obstacle avoidance decision, thereby improving the reliability of obstacle avoidance.

[0009] In addition, the automatic driving obstacle avoidance method proposed in the above embodiment of the present invention may also have the following additional technical features:

[0010] According to one embodiment of the present invention, determining the initial obstacle information based on the obstacle prediction set and the clustering set includes: projecting the clustering results in the clustering set to the camera image, and matching them with the obstacle prediction results in the obstacle prediction set to generate a matching set and an initial unmatched set; performing redundancy removal processing on the obstacle prediction results in the initial unmatched set to obtain a target unmatched set; and using the obstacle prediction results in the matching set and the target unmatched set as the initial obstacle information.

[0011] According to an embodiment of the present invention, the projecting the clustering results in the clustering set to the camera image and matching them with the obstacle prediction results in the obstacle prediction set to generate a matching set and an initial unmatched set includes: sorting the clustering results in the obstacle clustering set in ascending order according to the depth information of the clustering results; projecting the clustering results with the minimum depth in the sorted obstacle clustering set to the camera image, and searching the obstacle prediction set for an obstacle prediction result whose center distance is less than a preset distance threshold, and placing the found target obstacle prediction result in the camera image; The obstacle prediction result with the largest confidence is added to the matching result set; while deleting the clustering result of the minimum depth from the obstacle clustering set, the obstacle prediction result added to the matching result set is deleted from the obstacle prediction set, the obstacle clustering set and the obstacle prediction set are updated, and the step of projecting the clustering result of the minimum depth in the sorted obstacle clustering set to the camera image is returned to be executed until all clustering results in the obstacle clustering set are traversed, and the remaining obstacle prediction results in the target obstacle prediction results are added to the initial unmatched set.

[0012] According to one embodiment of the present invention, performing redundancy removal processing on the obstacle prediction results in the initial unmatched set to obtain the target unmatched set includes: calculating an intersection-and-union ratio between each obstacle prediction result in the matching result set and each obstacle prediction result in the initial unmatched set; if there is an intersection-and-union ratio greater than a preset threshold, deleting the obstacle prediction result corresponding to the intersection-and-union ratio from the initial unmatched set, performing non-maximum suppression processing on the obstacle prediction results in the unmatched set obtained after the deletion, to obtain the target unmatched set; if there is no intersection-and-union ratio greater than the preset threshold, performing non-maximum suppression processing on the obstacle prediction results in the initial unmatched set, to obtain the target unmatched set.

[0013] According to one embodiment of the present invention, the target detection model includes a backbone network, a neck network and a detection head, a first feature extraction module is arranged between the first CBL module of the neck network and the first splicing module, wherein the first CBL module is connected to the first output end of the backbone network, a second feature extraction module is arranged between the second CBL module of the neck network and the second splicing module, wherein the second CBL module is connected to the second output end of the backbone network, a third feature extraction module is arranged between the CSPSPP module of the backbone network and the third CBL module of the neck network, wherein the feature extraction module includes a first convolution branch, a second convolution branch and a bottleneck attention submodule, the input ends of the first convolution branch and the second convolution branch are connected to the input end of the feature extraction module, the output ends of the first convolution branch and the second convolution branch are connected to the input end of the bottleneck attention submodule after element-wise addition operation, the output end of the bottleneck attention submodule is connected to the output end of the feature extraction module, the first convolution branch includes a first convolution layer, and the second convolution branch includes a second convolution layer, a deformable convolution layer and a third convolution layer connected in sequence.

[0014] According to one embodiment of the present invention, making an obstacle avoidance decision based on the target obstacle information and the current state of the vehicle includes: performing path planning based on the target obstacle information and the current state of the vehicle to obtain a planned path; and determining the obstacle avoidance decision based on the planned path and the current state of the vehicle.

[0015] According to one embodiment of the present invention, before acquiring the multimodal data collected by the multimodal sensor, the method further includes: acquiring environmental information of the current environment of the vehicle; using an adaptive optimization algorithm to determine the degree of dependence on the camera image, the lidar data, and the millimeter-wave radar data according to the environmental information, so as to determine the initial obstacle information according to the obstacle prediction set and the clustering set, and determine the target obstacle information according to the initial obstacle information and the feature extraction set, determine the initial obstacle information according to the degree of dependence corresponding to the camera image, the degree of dependence corresponding to the lidar data, the obstacle prediction set, and the clustering set, and determine the target obstacle information according to the degree of dependence corresponding to the millimeter-wave radar data, the initial obstacle information, and the feature extraction set.

[0016] To achieve the above-mentioned purpose, the second aspect of the present invention proposes a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the automatic driving obstacle avoidance method proposed in the first aspect of the present invention is implemented.

[0017] To achieve the above-mentioned purpose, the third aspect of the present invention proposes a controller, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the automatic driving obstacle avoidance method proposed in the first aspect of the present invention is implemented.

[0018] To achieve the above-mentioned objective, a fourth aspect of the present invention provides a vehicle, comprising a multimodal sensor and a controller as provided in the third aspect of the present invention.

[0019] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flow chart of an automatic driving obstacle avoidance method according to an embodiment of the present invention;

[0021] Figure 2 is a structural block diagram of a target detection model according to an embodiment of the present invention;

[0022] Figure 3 is a structural block diagram of a feature extraction module according to an embodiment of the present invention;

[0023] Figure 4 is a flow chart of determining a target obstacle according to an embodiment of the present invention;

[0024] Figure 5 is a flow chart of generating a matching set and an initial unmatched set according to an embodiment of the present invention;

[0025] Figure 6 is a flow chart of obtaining a target unmatched set according to an embodiment of the present invention;

[0026] Figure 7 is a flow chart of making obstacle avoidance decisions according to one embodiment of the present invention;

[0027] Figure 8 is a flow chart of an automatic driving obstacle avoidance method according to a specific embodiment of the present invention;

[0028] Fig. 9 is a structural block diagram of a controller according to an embodiment of the present invention;

[0029] Fig.10 is a schematic diagram of a vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0031] The following is a detailed description of the automatic driving obstacle avoidance method, computer-readable storage medium, controller, and vehicle in accordance with an embodiment of the present invention in conjunction with the accompanying drawings and specific implementation methods.

[0032] The automatic driving obstacle avoidance method of the embodiment of the present invention is used for a vehicle, and a multimodal sensor is provided on the vehicle.

[0033] In the embodiment of the present invention, the multimodal sensor provided on the vehicle is not limited to a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor, a laser radar sensor, or a millimeter wave radar.

[0034] Figure 1 FIG. 1 is a flow chart of an automatic driving obstacle avoidance method according to an embodiment of the present invention. Figure 1 As shown, the automatic driving obstacle avoidance method may include:

[0035] S101, acquiring multimodal data collected by a multimodal sensor, wherein the multimodal data at least includes camera images, laser radar data, and millimeter wave radar data.

[0036] Specifically, multiple sensors such as cameras, lidar sensors, and millimeter-wave radar sensors are used to simultaneously collect different types of data about the vehicle's surroundings. Cameras can obtain rich visual information to identify the category and color of objects, etc. Lidar sensors can provide high-precision three-dimensional point cloud data to detect the shape and distance of objects. Millimeter-wave radar sensors can provide the speed and distance of targets.

[0037] In order to maintain the time synchronization of the data collected by each sensor, a high-precision clock generator can be used to send a synchronization acquisition signal to each sensor. After obtaining the data collected by each sensor, an interpolation algorithm is used to synchronize different types of data to reduce data deviation. The data collected by different sensors at the same time point can be further compared to repeatedly verify the time synchronization, so as to accurately integrate the different types of data collected by each sensor.

[0038] S102, while using a pre-trained target detection model to detect obstacles on the camera image, downsample and cluster the lidar data, and extract features from the millimeter wave radar data to obtain an obstacle prediction set, a clustering set, and a feature extraction set.

[0039] Specifically, the camera image can be used to detect obstacles using a pre-trained target detection model to obtain an obstacle prediction set for each obstacle output by the pre-trained target detection model. The laser radar data can be clustered to obtain a cluster set for each obstacle. The millimeter wave radar data can be feature extracted to obtain a feature extraction set for each obstacle.

[0040] S103, determining initial obstacle information according to the obstacle prediction set and the clustering set, and determining target obstacle information according to the initial obstacle information and the feature extraction set.

[0041] In order to improve the accuracy of the acquired obstacle information, the embodiment of the present invention fuses the obstacle detection results based on the camera image, the obstacle clustering results based on the lidar data, and the obstacle feature extraction results based on the millimeter wave radar data.

[0042] Specifically, to improve the accuracy of obstacle detection, the embodiment of the present invention determines initial obstacle information, such as obstacle color information, obstacle shape information, and obstacle category information, based on obstacle prediction results in the obstacle prediction set and obstacle clustering results in the clustering set. According to the determined initial obstacle information and obstacle feature extraction results in the feature extraction set, target obstacle information, such as obstacle color information, obstacle shape information, obstacle category information, obstacle distance information, and obstacle speed information, is determined.

[0043] In the embodiment of the present invention, the target obstacle information includes at least one of obstacle color information, obstacle shape information, obstacle category information, obstacle distance information, and obstacle speed information.

[0044] S104, making an obstacle avoidance decision according to the target obstacle information and the current state of the vehicle, and controlling the vehicle to perform obstacle avoidance driving according to the obstacle avoidance decision.

[0045] Specifically, a path is planned based on target obstacle information and the current state of the vehicle, an obstacle avoidance decision is made based on the planned path, and the vehicle is controlled to perform obstacle avoidance driving based on the determined obstacle avoidance decision, thereby improving obstacle avoidance reliability.

[0046] The automatic driving obstacle avoidance method of the embodiment of the present invention performs obstacle detection on the acquired multimodal data, and then fuses the obstacle detection results based on the camera image, the obstacle clustering results based on the lidar data, and the obstacle feature extraction results based on the millimeter wave radar data to improve the accuracy of the obtained target obstacle information, makes an obstacle avoidance decision based on the target obstacle information and the current state of the vehicle, controls the vehicle to perform obstacle avoidance driving according to the obstacle avoidance decision, and improves the reliability of obstacle avoidance.

[0047] In one embodiment of the present invention, Figure 2 and Figure 3 As shown, the target detection model may include a backbone network, a neck network and a detection head, a first feature extraction module is arranged between the first CBL module of the neck network and the first splicing module, wherein the first CBL module is connected to the first output end of the backbone network, a second feature extraction module is arranged between the second CBL module of the neck network and the second splicing module, wherein the second CBL module is connected to the second output end of the backbone network, a third feature extraction module is arranged between the CSPSPP module of the backbone network and the third CBL module of the neck network, wherein the feature extraction module includes a first convolution branch, a second convolution branch and a bottleneck attention submodule, the input ends of the first convolution branch and the second convolution branch are connected to the input end of the feature extraction module, the output ends of the first convolution branch and the second convolution branch are connected to the input end of the bottleneck attention submodule after element-wise addition operation, the output end of the bottleneck attention submodule is connected to the output end of the feature extraction module, the first convolution branch includes a first convolution layer, and the second convolution branch includes a second convolution layer, a deformable convolution layer and a third convolution layer connected in sequence.

[0048] The target detection model in the embodiment of the present invention adopts the YOLOv7-tiny model. In order to improve the detection accuracy of the network model in complex environments, the embodiment of the present invention introduces a feature extraction module DCN-BAM based on the attention mechanism and deformable convolution into the neck network of the YOLOv7-tiny model. Figure 2 .

[0049] Specifically, Figure 2 As shown, when the feature extraction module DCN-BAM is integrated into the YOLOv7-tiny model, the feature extraction module is set between the first CBL module and the first splicing module of the neck network, the feature extraction module is set between the second CBL module and the second splicing module of the neck network, and the feature extraction module is set between the CSPSPP module of the backbone network and the third CBL module of the neck network to improve the detection accuracy of the YOLOv7-tiny model in complex environments.

[0050] Specifically, Figure 3 As shown, the feature extraction module DCN-BAM includes a first convolution branch, a second convolution branch and a bottleneck attention module (Bottleneck Attention Module, BAM). The first convolution branch includes a first convolution layer with a convolution kernel of 1×1, and the second convolution branch includes a second convolution layer with a convolution kernel of 1×1, a deformable convolutional network (DCN) with a convolution kernel of 3×3, and a third convolution layer with a convolution kernel of 1×1, which are connected in sequence. The first convolution branch and the second convolution branch respectively extract features from the feature map input to the feature extraction module to obtain a first feature map and a second feature map. The first feature map and the second feature map are added element-wise and then input to the bottleneck attention module, which extracts features from the feature map after the element-wise addition operation to capture important target features.

[0051] In this embodiment, the deformable convolution layer can change the sampling position of the convolution kernel to adapt to the changes in the target shape of the road obstacle. The bottleneck attention submodule can assign different weights to different spatial positions and channels of the feature map to capture important target characteristics and suppress the interference of complex environments.

[0052] In the embodiment of the present invention, according to the new data continuously collected during the driving process of the vehicle, the target detection model can be updated online using SVRG (Stochastic Variance Reduced Gradient) to adapt to different environments and conditions.

[0053] In one embodiment of the present invention, before using a pre-trained target detection model to perform obstacle detection on a camera image, the autonomous driving obstacle avoidance method further includes: performing color correction and / or distortion correction processing on the camera image.

[0054] Specifically, before the camera image is input into the pre-trained target detection model, the camera image is subjected to color correction and / or distortion correction and other processing to improve the clarity and accuracy of the camera image, thereby further improving the accuracy of the obstacle prediction results output by the pre-trained target detection model.

[0055] It should be noted that the lidar data collected by the lidar sensor is three-dimensional point cloud data. Downsampling and clustering operations on the lidar data can reduce the amount of data and extract key features.

[0056] Specifically, when downsampling the LiDAR data, the three-dimensional space is divided into cubes (i.e., voxel grids) of equal size. All point cloud data are traversed, and the points contained in each voxel grid are counted. Then, a simple strategy is usually used to determine the representative point of each voxel grid, such as taking the average value of the coordinates of all points in the voxel grid (i.e., the center of gravity) as the representative point, or randomly selecting a point as the representative point. Finally, only these representative points are retained to form the downsampled point cloud data, thereby achieving the purpose of reducing the amount of data.

[0057] For example, a LiDAR sensor scans a road traffic scene containing streets, vehicles, and pedestrians, and the amount of LiDAR data obtained is very large. If the side length of the voxel grid is set to 0.2 meters, the entire three-dimensional space is divided into many small cubes with a side length of 0.2 meters. For example, in a certain voxel grid, there are 10 points representing a part of a vehicle. By taking the center of gravity point as the representative point, the original 10 points are finally replaced by this one representative point.

[0058] Downsampling can better preserve the spatial distribution structure of the point cloud and the approximate shape of the object, so that the downsampled point cloud data can still reflect the contours and relative position relationships of the objects in the scene, while significantly reducing the amount of data and facilitating subsequent processing.

[0059] Specifically, when clustering the downsampled LiDAR data, select any unclustered point from the downsampled point cloud data as the starting point, traverse the remaining points, and if the Euclidean distance between a point and the starting point is less than the set threshold, the point is divided into the same cluster as the starting point. Then, with the newly added point as the center, repeat the above traversal and division process again until no new points can be added to the cluster. Then, select a new starting point from the unclustered points, and repeat the entire clustering process until all points are divided into corresponding clusters.

[0060] For example, in a parking lot scene, the point cloud obtained by LiDAR scanning includes multiple cars and some surrounding facilities, etc. By setting the distance threshold to 0.3 meters for Euclidean distance clustering, the point clouds belonging to the same car can be clustered together, and the point clouds of different cars are clustered into different classes. At the same time, the point clouds of objects such as street lights and signboards are also clustered into their own classes.

[0061] In one embodiment of the present invention, a convolutional neural network (CNN) and / or a generative adversarial network (GAN) are used to extract features from millimeter wave radar data.

[0062] It should be noted that the dual method of using convolutional neural network CNN combined with generative adversarial network GAN to perform noise reduction, filtering and other processing on the collected millimeter-wave radar data can remove noise and outliers and improve data quality. Among them, when using the dual method of convolutional neural network CNN combined with generative adversarial network GAN, the convolutional neural network CNN can be used to purify the feature data to a certain extent, and then the purified feature data is provided as input to the generator of the generative adversarial network GAN. Based on the features processed by CNN, the generator further generates data that is more in line with the real noise-free and outlier-free distribution, and then the discriminator performs identification, and the two continue to train against each other, and finally obtain high-quality data.

[0063] It should be noted that for data with spatial structure such as millimeter-wave radar data and audio sensor data, as well as data with sequence structure and meaningful local features, a dual method of convolutional neural network CNN combined with generative adversarial network GAN can be used for noise reduction and filtering. For other simple data, convolutional neural network CNN or generative adversarial network GAN can be selected separately for noise reduction and filtering according to the characteristics of the data.

[0064] In one embodiment of the present invention, Figure 4 As shown, determining the target obstacle according to the obstacle prediction set and the clustering set may include:

[0065] S201, projecting the clustering results in the clustering set to the camera image, and matching them with the obstacle prediction results in the obstacle prediction set to generate a matching set and an initial unmatched set;

[0066] S202, performing redundancy removal processing on obstacle prediction results in the initial unmatched set to obtain a target unmatched set;

[0067] S203: Use obstacle prediction results in the matching set and the target unmatched set as target obstacles.

[0068] Although the laser radar sensor can measure the depth information and shape information of the obscured target, due to the inaccurate size of the measured size, the embodiment of the present invention fuses the clustering results (three-dimensional space object information (such as the actual position and shape)) obtained based on the laser radar data with the obstacle prediction results detected by the target detection model from the camera image, so as to more comprehensively and accurately perceive the environment around the vehicle and improve the accuracy of the determined target obstacles. For example, based on the clustering of the laser radar data, it is determined that there is an object in front, and its approximate position is known through its centroid coordinates, and the target detection model recognizes from the camera image that the object is a pedestrian and its specific position in the image. By fusing these two types of information, the actual distance, orientation and appearance characteristics of the pedestrian relative to the vehicle can be determined more accurately, which helps the vehicle to better plan the path and make decisions (such as whether to slow down, avoid, etc.).

[0069] Specifically, in order to perceive the environment around the vehicle more comprehensively and accurately, the clustering results in the clustering set are projected onto the camera image and matched with the obstacle prediction results in the obstacle prediction set. The obstacle prediction results in the obstacle prediction set that match the clustering results are added to the matching set, and the obstacle prediction results in the obstacle prediction set that do not match the clustering results are added to the initial unmatched set. The obstacle prediction results in the initial unmatched set are de-redundanted, and the obstacle prediction results in the matching set and the target unmatched set obtained by de-redundanting are used as target obstacles.

[0070] It should be noted that the obstacle prediction results in the obstacle prediction set are obtained based on the camera image, and the clustering results in the clustering set are projected onto the camera image to facilitate matching and fusion of the clustering results and the obstacle prediction results.

[0071] In one embodiment of the present invention, Figure 5 As shown, the clustering results in the clustering set are projected onto the camera image and matched with the obstacle prediction results in the obstacle prediction set to generate a matching set and an initial unmatched set, which may include:

[0072] S301, sorting the clustering results in the obstacle clustering set in ascending order according to the depth information of the clustering results;

[0073] S302, projecting the clustering result with the minimum depth in the sorted obstacle clustering set to the camera image, and searching the obstacle prediction result with a distance from its center less than a preset distance threshold from the obstacle prediction set, and adding the obstacle prediction result with the highest confidence among the found target obstacle prediction results to the matching result set;

[0074] S303, while deleting the clustering result of the minimum depth from the obstacle clustering set, the obstacle prediction results added to the matching result set are deleted from the obstacle prediction set, the obstacle clustering set and the obstacle prediction set are updated, and the step of projecting the clustering result of the minimum depth in the sorted obstacle clustering set to the camera image is returned to be executed until all the clustering results in the obstacle clustering set are traversed, and the remaining obstacle prediction results in the target obstacle prediction results are added to the initial unmatched set.

[0075] The clustering result in the embodiment of the present invention may include depth information and shape information of the obstacle.

[0076] The obstacle prediction result in the embodiment of the present invention may include the target category of the obstacle, location information (coordinate information in the camera image coordinate system) and the confidence level corresponding to the target category.

[0077] The embodiment of the present invention sorts the clustering results in the obstacle clustering set in ascending order of depth information, projects the clustering results in the obstacle clustering set to the camera image in turn, and then matches and fuses them with the obstacle prediction results in the obstacle prediction set.

[0078] Specifically, the clustering results in the obstacle clustering set are sorted in ascending order according to the depth information of the clustering results, and the clustering result with the minimum depth in the sorted obstacle clustering set is projected onto the camera image. The obstacle prediction result whose center distance with the clustering result with the minimum depth is less than a preset distance threshold is searched from the obstacle prediction set, and the obstacle prediction result with the highest confidence is added to the matching result set.

[0079] The clustering results based on LiDAR data are mainly based on the point cloud data in three-dimensional space to obtain the location information (depth information) and approximate shape information of the object, while the obstacle prediction results output by the target detection model are based on the target category, location information and confidence obtained from the camera image data. By searching for obstacle prediction results whose distance from the center of the clustering result is less than the preset distance threshold, the target information from two different sensors and processing methods can be associated. Searching for obstacle prediction results whose distance from the center of the clustering result is less than the preset distance threshold in the obstacle prediction set is based on the distance condition. The obstacle prediction results that are closer to the clustering results in spatial position are selected, which means that those targets that are more likely to represent the same actual object as the current LiDAR cluster are selected from the many obstacle prediction results, achieving a preliminary association.

[0080] Among the obstacle prediction results that have been filtered out and whose distance is less than the threshold, further check the confidence corresponding to each target category (the confidence reflects the reliability of the network model in detecting and judging the category of the obstacle, usually a value between 0 and 1, the higher the value, the more reliable it is). Select the prediction result with the highest confidence. It should be noted that the higher the confidence, the more accurate and reliable the network model's judgment is, and it is more likely to accurately correspond to the real object represented by the lidar cluster.

[0081] The obstacle prediction result with the highest confidence is added to the matching result set. The matching result set is a data set that finally integrates the lidar clustering information and the accurate prediction information of the network model. The matching result set stores the target-related information that has been screened and optimized and is considered to be accurately associated. In the future, more in-depth analysis, decision-making and other operations can be performed based on this set.

[0082] The clustering result with the minimum depth is deleted from the obstacle clustering set, and the obstacle prediction results added to the matching result set are deleted from the obstacle prediction set. The obstacle clustering set and the obstacle prediction set are updated, and the above process is repeated until all the lidar clustering results are traversed to obtain the matching result set. The remaining obstacle prediction results in the target obstacle prediction results are added to the initial unmatched set to obtain the initial unmatched set.

[0083] Among them, the obstacle prediction results that have been selected into the matching result set are deleted from the obstacle prediction set. On the one hand, this can avoid the subsequent repeated processing of the same matched target and improve the processing efficiency; on the other hand, it also ensures the consistency of the data, so that the remaining obstacle prediction set contains obstacle prediction results that have not yet participated in the matching or have failed to match, which is convenient for continuing to perform associated matching operations with other lidar clusters.

[0084] It should be noted that the camera image records the visual information in the scene in the form of a two-dimensional pixel matrix, including the color, texture, shape (from a two-dimensional perspective) of the object. The shooting perspective and scene range of the CMOS image sensor overlap with the area scanned by the LiDAR sensor, and there is an image area on the camera image that corresponds to the clustering result obtained based on the LiDAR data.

[0085] In one embodiment of the present invention, Figure 6 As shown, the obstacle prediction results in the initial unmatched set are processed to remove redundancy, and the target unmatched set is obtained, including:

[0086] S401, calculating the intersection-and-union ratio of each obstacle prediction result in the matching result set and each obstacle prediction result in the initial unmatched set;

[0087] S402, if there is an intersection-and-union ratio greater than a preset threshold, the obstacle prediction result corresponding to the intersection-and-union ratio is deleted from the initial unmatched set, and a non-maximum suppression process is performed on the obstacle prediction results in the unmatched set obtained after the deletion to obtain a target unmatched set;

[0088] S403: If there is no intersection-over-union ratio greater than a preset threshold, non-maximum suppression processing is performed on the obstacle prediction results in the initial unmatched set to obtain a target unmatched set.

[0089] Specifically, the prediction results in the initial unmatched set that are redundant with the matching result set are deleted, and the intersection over union (IoU) of each obstacle prediction result in the matching result set and each obstacle prediction result in the initial unmatched set is calculated.

[0090] If there is an intersection-and-union ratio greater than a preset threshold, the redundant obstacle prediction results with an intersection-and-union ratio greater than the preset threshold are deleted from the initial unmatched set to obtain the initial unmatched set with the first redundancy removed (the unmatched set obtained after deletion). The obstacle prediction results in the unmatched set obtained after deletion are subjected to non-maximum suppression processing using the non-maximum suppression algorithm (NMS) to delete the redundant results that are not detected by the laser radar in the initial unmatched set with the first redundancy removed, and obtain the final initial unmatched set with the redundancy removed (the target unmatched set).

[0091] If there is no intersection-over-union ratio greater than the preset threshold, the non-maximum suppression algorithm is used to perform non-maximum suppression on the obstacle prediction results in the initial unmatched set to obtain the target unmatched set.

[0092] It should be noted that the intersection over union (IoU) can measure the degree of overlap between obstacle prediction results. When the intersection over union (IoU) is greater than the preset threshold, it means that the unmatched obstacle prediction results have a high overlap with the obstacle prediction results already in the matching result set, and it is likely to be a repeated detection of the same object. Deleting these redundant prediction results can avoid erroneous judgments or unnecessary calculations due to the interference of multiple similar results in subsequent processing (such as target tracking, target classification refinement, etc.). If these redundant prediction results are not deleted, the same actual object may have multiple different representations in the matching result set, which will reduce the matching accuracy. For example, in an autonomous driving scenario, for a vehicle in front, if there are multiple different but highly overlapping prediction results for the vehicle in the matching result set, the vehicle will be confused about the judgment of the target in front, such as the inability to accurately determine the exact position and boundary of the vehicle during path planning. By deleting redundant prediction results, each prediction result in the matching result set can correspond to an actual object more accurately, improving the consistency and accuracy of the entire set.

[0093] In one embodiment of the present invention, Figure 7 As shown in the figure, obstacle avoidance decisions are made based on the target obstacle information and the current state of the vehicle, including:

[0094] S501, performing path planning according to target obstacle information and the current state of the vehicle to obtain a planned path;

[0095] S502: Determine an obstacle avoidance decision based on the planned path and the current state of the vehicle.

[0096] Specifically, the position, size, speed and other information of obstacles around the vehicle can be determined based on the target obstacle information. The target tracking algorithm is used to track the detected target obstacle based on the target obstacle information, and the motion state of the target obstacle is understood in real time. According to the motion state of the target obstacle and the current state of the vehicle, path planning is performed to determine a safe driving path. Obstacle avoidance decisions are made based on the path planning results and the current state of the vehicle. Obstacle avoidance decisions include operations such as acceleration, deceleration, and steering to avoid collisions with obstacles.

[0097] It should be noted that while determining the obstacle avoidance decision based on the planned path and the current state of the vehicle, the decision priority in different situations is considered, such as taking braking measures first in an emergency to ensure the safety of the vehicle.

[0098] The embodiment of the present invention adopts model predictive control (MPC) to perform model prediction based on target obstacle information, and uses a discretized system model to predict the state of the vehicle in the future based on the current state information and control input of the vehicle during driving, so as to perform path planning.

[0099] It is feasible to use the formula X = Φ*x 0 +Γ*U to predict the future state. Among them, X is the predicted state sequence, Φ is the state transfer matrix, x 0 is the state at the current moment, Γ is the control input matrix, and U is the control input sequence.

[0100] The control input matrix Γ is mainly related to the dynamic model of the vehicle. It describes the degree of influence of the control input on the change of the vehicle state. The values ​​of its elements are determined based on the physical characteristics of the vehicle (such as mass, inertia, friction between tires and the ground, etc.) and the characteristics of the vehicle control system (such as the response characteristics of the power system, the transmission ratio of the steering system, etc.).

[0101] It should be noted that the vehicle’s current state x 0 It is closely related to the information after modal fusion and is a comprehensive state of the information after modal fusion at the current moment. The information after modal fusion contains the data fusion results of the vehicle's own multiple sensors (such as lidar, camera, millimeter wave radar, inertial measurement unit, etc.), which are used to accurately describe the current comprehensive state of the vehicle. Some of the data in the modal fusion information can be used to generate control input U. For example, in an autonomous driving scenario, the vehicle's acceleration, deceleration (controlling the throttle or brake), and steering (controlling the steering wheel) operations can be determined through the fused environmental perception information (such as whether there are obstacles ahead, the curvature of the road, etc.), and these operations constitute the control input sequence U.

[0102] The control input sequence U is determined by the control inputs at the current moment and in the future while the vehicle is driving. In practical applications, it is a sequence of control input variables containing multiple time steps. For example, in the scenario of autonomous driving, the control inputs include throttle opening, brake force, steering wheel angle, etc. If the state of the vehicle is predicted in the next N time steps, then the control input sequence U = [μ 0 ,μ 1 ,……,μ N-1 ], where μ 0 is the control input at the current moment (such as the current throttle opening), μ 1is the control input for the next time step plan, and so on. These control inputs are determined based on the vehicle's driving goals (such as maintaining a certain speed, following the vehicle in front, turning, etc.) and the current environmental perception (such as obstacles ahead that need to be avoided, the road narrowing and the need to adjust the speed, etc.).

[0103] In one embodiment of the present invention, Figure 8 As shown, before acquiring the multimodal data collected by the multimodal sensor, the method further includes:

[0104] S601, obtaining environmental information of the vehicle's current environment;

[0105] S602, using an adaptive optimization algorithm, determining the degree of dependence on the camera image, the lidar data, and the millimeter-wave radar data according to the environmental information, so that when determining the initial obstacle information according to the obstacle prediction set and the clustering set, and determining the target obstacle information according to the initial obstacle information and the feature extraction set, the initial obstacle information is determined according to the degree of dependence corresponding to the camera image, the degree of dependence corresponding to the lidar data, the obstacle prediction set, and the clustering set, and the target obstacle information is determined according to the degree of dependence corresponding to the millimeter-wave radar data, the initial obstacle information, and the feature extraction set.

[0106] Specifically, the environmental information of the vehicle's current environment is obtained, and an adaptive optimization algorithm is used to dynamically adjust the degree of dependence on the parameters of different types of sensors and the fusion strategy of different types of data under different environmental conditions and driving scenarios, so as to further enhance the accuracy of obstacle detection.

[0107] It should be noted that adaptive optimization enables decisions to be adjusted according to real-time changes. When the vehicle is driving, factors such as road conditions, weather, and traffic flow are constantly changing. The adaptive optimization algorithm can adjust the obstacle avoidance strategy in real time to ensure that the decision is always in line with the current best safety plan. For example, when the road surface is slippery on rainy days, the algorithm will automatically reduce the speed and increase the braking distance to improve the reliability of obstacle avoidance decisions. For example, under strong light, the camera image may be affected. The degree of reliance on lidar data can be automatically adjusted to ensure continuous and accurate detection of obstacles.

[0108] The adaptive optimization algorithm in the embodiment of the present invention can also cope with various emergencies and abnormal conditions, such as the sudden appearance of obstacles and sensor failures, and maintain the stable operation and safe driving of the vehicle through adaptive optimization. For example, when a sensor fails, the algorithm can automatically adjust the fusion strategy and rely on other sensors that are working normally to continue obstacle detection and obstacle avoidance decisions, ensuring that the vehicle will not lose safety due to a single sensor failure.

[0109] The adaptive optimization algorithm in the embodiment of the present invention can also use dynamic convolution to adaptively adjust the parameters of the obstacle avoidance algorithm according to the driving state of the vehicle and environmental changes. For example, the range and sensitivity of obstacle detection can be adjusted according to the vehicle speed, and the constraints of path planning can be adjusted according to the road conditions.

[0110] The embodiment of the present invention establishes performance evaluation indicators (including collision rate, driving efficiency, comfort, etc.) to evaluate the performance of the automatic driving obstacle avoidance method of the embodiment of the present invention in real time. Based on the established performance evaluation indicators, a large number of simulation tests are carried out on the provided automatic driving obstacle avoidance method in the Gazebo virtual environment platform to simulate different traffic scenarios, weather conditions and road conditions, verifying the effectiveness and reliability of the automatic driving obstacle avoidance method of the embodiment of the present invention. Tests are carried out on actual vehicles, including closed field tests and open road tests, to verify the performance of the automatic driving obstacle avoidance method of the embodiment of the present invention in a real environment. Safety verifications such as fault injection tests and extreme working condition tests are carried out on the obstacle avoidance algorithm to ensure the safety of the vehicle under various circumstances.

[0111] The automatic driving obstacle avoidance method of the embodiment of the present invention enhances the reliability of decision-making. Multimodal fusion provides a richer source of information for obstacle avoidance decisions. By comprehensively considering the data of multiple sensors, the threat level of obstacles and the safety status of the surrounding environment can be more comprehensively evaluated, so as to make more reliable decisions. For example, when facing a sudden appearance of a pedestrian, the distance information detected by the lidar, the pedestrian movement recognized by the camera, and the relative speed measured by the millimeter-wave radar can be combined to quickly determine whether emergency braking or steering avoidance is required.

[0112] The automatic driving obstacle avoidance method of the embodiment of the present invention improves driving safety. Accurate obstacle detection and reliable decision-making can effectively avoid collision accidents and greatly improve the driving safety of automatic driving vehicles. Whether driving fast on the highway or shuttling through city streets, the obstacle avoidance algorithm with multimodal fusion and adaptive optimization can detect potential dangers in time and take appropriate measures to avoid them, providing higher safety protection for passengers and pedestrians.

[0113] The automatic driving obstacle avoidance method of the embodiment of the present invention improves the obstacle detection accuracy. Specifically, the fusion of multimodal data can take advantage of the advantages of different sensors. The camera can capture rich texture and color features, the lidar can provide accurate three-dimensional spatial information, and the millimeter-wave radar can work stably in bad weather. The fused multimodal data can more accurately detect obstacles of various shapes, sizes and materials, and reduce missed detections and false detections. For example, in a complex urban environment, for irregularly shaped obstacles, such as temporary construction facilities, multimodal fusion can combine the visual features of the camera and the geometric information of the lidar to more comprehensively identify their position and contours and improve detection accuracy.

[0114] The automatic driving obstacle avoidance method of the embodiment of the present invention improves driving efficiency and comfort. Accurate obstacle avoidance decisions can not only ensure safety, but also optimize the vehicle's driving path and improve driving efficiency. The algorithm can choose the fastest and smoothest path to avoid obstacles, reduce unnecessary deceleration and detours, and thus shorten driving time. For example, in the case of traffic congestion, the algorithm can accurately determine the position and movement trend of surrounding vehicles through multimodal fusion, find the best time to insert, and improve the vehicle's traffic efficiency. At the same time, reasonable obstacle avoidance decisions can also reduce the vehicle's rapid acceleration, rapid deceleration, and frequent steering, and improve riding comfort. Passengers will not feel severe shaking and bumps in the car, which improves the user experience of automatic driving.

[0115] The automatic driving obstacle avoidance method of the embodiment of the present invention is adaptable to different environments and working conditions. The automatic driving obstacle avoidance algorithm with multimodal fusion and adaptive optimization has strong adaptability and can work normally in various environments and working conditions. Whether it is day or night, sunny or rainy, urban roads or country roads, the algorithm can be adjusted and optimized according to the actual situation to ensure the safe driving of the vehicle. For different types of obstacles, such as static obstacles (such as roadblocks, telephone poles, etc.) and dynamic obstacles (such as pedestrians, vehicles, etc.), the algorithm can also adopt different obstacle avoidance strategies to improve the ability to cope with various complex situations.

[0116] The automatic driving obstacle avoidance method of the embodiment of the present invention adds a feature extraction module based on the attention mechanism and deformable convolution to the original YOLOv7-tiny model, enhances the original YOLOv7-tiny model's ability to express obstacle target features in complex environmental scenarios, performs obstacle detection based on the fusion of laser radar, machine vision, and millimeter-wave radar, improves obstacle detection accuracy, and uses an adaptive optimization algorithm to optimize obstacle avoidance paths and decisions in real time according to different traffic scenarios and environmental conditions, thereby improving the algorithm's adaptability and robustness. Through multimodal fusion and adaptive optimization, the performance and reliability of the obstacle avoidance algorithm are improved.

[0117] The present invention provides a computer-readable storage medium.

[0118] In this embodiment, a computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, the automatic driving obstacle avoidance method as described above is implemented.

[0119] The invention provides a controller.

[0120] In this embodiment, the controller may include a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the automatic driving obstacle avoidance method as described above is implemented.

[0121] Fig. 9 4 is a structural block diagram of a controller according to an embodiment of the present invention.

[0122] like Fig. 9 As shown, the controller 500 includes: a processor 501 and a memory 503. The processor 501 and the memory 503 are connected, such as through a bus 502. Optionally, the controller 500 may also include a transceiver 504. It should be noted that in actual applications, the transceiver 504 is not limited to one, and the structure of the controller 500 does not constitute a limitation on the embodiments of the present invention.

[0123] Processor 501 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present invention. Processor 501 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0124] The bus 502 may include a path to transmit information between the above components. The bus 502 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 502 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 9 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0125] The memory 503 is used to store a computer program corresponding to the automatic driving obstacle avoidance method of the above embodiment of the present invention, and the computer program is controlled and executed by the processor 501. The processor 501 is used to execute the computer program stored in the memory 503 to implement the content shown in the above method embodiment.

[0126] Among them, the controller 500 includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig. 9 The controller 500 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0127] The computer-readable storage medium and controller in the embodiment of the present invention use the above automatic driving obstacle avoidance method to detect obstacles based on the fusion of laser radar, machine vision and millimeter wave radar, improve the obstacle detection accuracy, and use an adaptive optimization algorithm to optimize the obstacle avoidance path and decision in real time according to different traffic scenes and environmental conditions, thereby improving the adaptability and robustness of the algorithm. Through multimodal fusion and adaptive optimization, the performance and reliability of the obstacle avoidance algorithm are improved.

[0128] The present invention provides a vehicle.

[0129] Fig.10 FIG. 1 is a schematic diagram of a vehicle according to an embodiment of the present invention. Fig.10 As shown, vehicle 1000 may include multimodal sensor 100 and controller 500 as described above.

[0130] The vehicle in the embodiment of the present invention uses the above automatic driving obstacle avoidance method to detect obstacles based on the fusion of laser radar, machine vision and millimeter wave radar to improve the obstacle detection accuracy. It uses an adaptive optimization algorithm to optimize the obstacle avoidance path and decision in real time according to different traffic scenes and environmental conditions to improve the adaptability and robustness of the algorithm. Through multimodal fusion and adaptive optimization, the performance and reliability of the obstacle avoidance algorithm are improved.

[0131] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0132] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0133] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0134] In the description of the present invention, it is to be understood that the terms “center”, “longitudinal”, “lateral”, “length”, “width”, “thickness”, “up”, “down”, “front”, “back”, “left”, “right”, “vertical”, “horizontal”, “top”, “bottom”, “inside”, “outside”, “clockwise”, “counterclockwise”, “axial”, “radial”, “circumferential”, etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0135] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0136] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0137] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, a first feature being "above", "above" or "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below", "below" or "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.

[0138] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. An automatic driving obstacle avoidance method, characterized in that: For a vehicle, wherein a multimodal sensor is provided on the vehicle, the method comprising: Acquire multimodal data collected by the multimodal sensor, wherein the multimodal data includes at least camera images, laser radar data, and millimeter wave radar data; While performing obstacle detection on the camera image using a pre-trained target detection model, downsampling and clustering operations are performed on the laser radar data, and feature extraction is performed on the millimeter wave radar data to obtain an obstacle prediction set, a clustering set, and a feature extraction set; determining initial obstacle information according to the obstacle prediction set and the clustering set, and determining target obstacle information according to the initial obstacle information and the feature extraction set; An obstacle avoidance decision is made according to the target obstacle information and the current state of the vehicle, and the vehicle is controlled to perform obstacle avoidance driving according to the obstacle avoidance decision.

2. The automatic driving obstacle avoidance method according to claim 1, characterized in that: The determining the initial obstacle information according to the obstacle prediction set and the clustering set includes: Projecting the clustering results in the clustering set to the camera image, and matching them with the obstacle prediction results in the obstacle prediction set to generate a matching set and an initial unmatched set; Performing redundancy removal processing on obstacle prediction results in the initial unmatched set to obtain a target unmatched set; The obstacle prediction results in the matching set and the target unmatched set are used as the initial obstacle information.

3. The automatic driving obstacle avoidance method according to claim 2, characterized in that: The projecting the clustering results in the clustering set to the camera image and matching them with the obstacle prediction results in the obstacle prediction set to generate a matching set and an initial unmatched set includes: Sorting the clustering results in the obstacle clustering set in ascending order according to the depth information of the clustering results; Projecting the clustering result with the minimum depth in the sorted obstacle clustering set to the camera image, searching the obstacle prediction result whose center distance is less than a preset distance threshold from the obstacle prediction set, and adding the obstacle prediction result with the highest confidence among the found target obstacle prediction results to the matching result set; While deleting the clustering result of the minimum depth from the obstacle clustering set, the obstacle prediction results added to the matching result set are deleted from the obstacle prediction set, the obstacle clustering set and the obstacle prediction set are updated, and the step of projecting the clustering result of the minimum depth in the sorted obstacle clustering set to the camera image is returned to be executed until all clustering results in the obstacle clustering set are traversed, and the remaining obstacle prediction results in the target obstacle prediction results are added to the initial unmatched set.

4. The automatic driving obstacle avoidance method according to claim 2, characterized in that: The step of performing redundancy removal processing on the obstacle prediction results in the initial unmatched set to obtain a target unmatched set includes: Calculating an intersection-and-union ratio of each obstacle prediction result in the matching result set and each obstacle prediction result in the initial unmatched set; If there is an intersection-and-union ratio greater than a preset threshold, the obstacle prediction result corresponding to the intersection-and-union ratio is deleted from the initial unmatched set, and a non-maximum suppression process is performed on the obstacle prediction results in the unmatched set obtained after the deletion to obtain the target unmatched set; If there is no intersection-over-union ratio greater than a preset threshold, non-maximum suppression processing is performed on the obstacle prediction results in the initial unmatched set to obtain the target unmatched set.

5. The automatic driving obstacle avoidance method according to claim 1, characterized in that: The target detection model includes a backbone network, a neck network and a detection head, a first feature extraction module is arranged between a first CBL module of the neck network and a first splicing module, wherein the first CBL module is connected to a first output end of the backbone network, a second feature extraction module is arranged between a second CBL module of the neck network and a second splicing module, wherein the second CBL module is connected to a second output end of the backbone network, a third feature extraction module is arranged between a CSPSPP module of the backbone network and a third CBL module of the neck network, wherein the feature extraction module includes a first convolution branch, a second convolution branch and a bottleneck attention submodule, the input ends of the first convolution branch and the second convolution branch are connected to the input end of the feature extraction module, the output ends of the first convolution branch and the second convolution branch are connected to the input end of the bottleneck attention submodule after element-wise addition operation, the output end of the bottleneck attention submodule is connected to the output end of the feature extraction module, the first convolution branch includes a first convolution layer, and the second convolution branch includes a second convolution layer, a deformable convolution layer and a third convolution layer connected in sequence.

6. The automatic driving obstacle avoidance method according to claim 1, characterized in that: The making of an obstacle avoidance decision according to the target obstacle information and the current state of the vehicle includes: Performing path planning according to the target obstacle information and the current state of the vehicle to obtain a planned path; The obstacle avoidance decision is determined based on the planned path and the current state of the vehicle.

7. The automatic driving obstacle avoidance method according to claim 1, characterized in that: Before acquiring the multimodal data collected by the multimodal sensor, the method further includes: Obtaining environmental information of the vehicle's current environment; An adaptive optimization algorithm is used to determine the degree of dependence on the camera image, the lidar data, and the millimeter-wave radar data according to the environmental information, so that when initial obstacle information is determined according to the obstacle prediction set and the clustering set, and target obstacle information is determined according to the initial obstacle information and the feature extraction set, the initial obstacle information is determined according to the degree of dependence corresponding to the camera image, the degree of dependence corresponding to the lidar data, the obstacle prediction set, and the clustering set, and the target obstacle information is determined according to the degree of dependence corresponding to the millimeter-wave radar data, the initial obstacle information, and the feature extraction set.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the automatic driving obstacle avoidance method as described in any one of claims 1 to 7 is implemented.

9. A controller, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that: When the computer program is executed by the processor, the automatic driving obstacle avoidance method as described in any one of claims 1 to 7 is implemented.

10. A vehicle, characterized in that: Comprising a multimodal sensor and a controller as claimed in claim 9.