Obstacle detection method, vehicle control method, device, equipment and storage medium

By combining semantic segmentation and obstacle detection models with the positional relationship between candidate obstacles and traffic semantic image blocks, the problem of vehicles having difficulty recognizing obstacles is solved, achieving more efficient obstacle detection and vehicle control.

CN117315624BActive Publication Date: 2026-05-19BEIJING JINGDONG QIANSHITECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING JINGDONG QIANSHITECHNOLOGY CO LTD
Filing Date
2023-10-08
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In related technologies, vehicles have difficulty accurately identifying obstacles in the driving environment, resulting in low efficiency of autonomous driving.

Method used

By inputting the driving environment image into the semantic segmentation layer and outputting an environmental semantic image, obstacle detection is performed. The positional relationship between the candidate obstacle detection results and the traffic semantic image blocks is used, combined with the traffic obstacle detection model, to output the traffic obstacle detection results.

Benefits of technology

It improves the accuracy of obstacle detection and the precision of subsequent vehicle control, reduces the amount of data processing, and enhances vehicle driving efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315624B_ABST
    Figure CN117315624B_ABST
Patent Text Reader

Abstract

The disclosure provides an obstacle detection method, a vehicle control method, an apparatus, a device and a storage medium, and relates to the fields of artificial intelligence, intelligent driving and smart logistics. The detection method comprises: inputting a driving environment image related to a vehicle into a semantic segmentation layer to output an environment semantic image; performing obstacle detection on the environment semantic image to obtain at least one initial obstacle detection result; determining, from the at least one initial obstacle detection result, a candidate obstacle detection result at least partially coinciding with a passing semantic image block according to a positional relationship between the initial obstacle detection result and the passing semantic image block; and inputting a candidate obstacle image block associated with the candidate obstacle detection result in the driving environment image into a passing obstacle detection model to output a passing obstacle detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of artificial intelligence, intelligent driving, and smart logistics, and more specifically, to an obstacle detection method, a vehicle control method, an apparatus, an equipment, and a storage medium. Background Technology

[0002] With the rapid development of technology, autonomous driving functions are widely used in passenger cars, logistics vehicles, and automated inspection vehicles, enabling vehicles to perform autonomous driving functions such as automatic parking and automatic inspection based on the generated movement path, thereby improving the driving efficiency and operational efficiency of vehicles.

[0003] In realizing the concept disclosed herein, the inventors discovered at least the following problems in the related technology: the related vehicles have difficulty accurately identifying obstacles in the driving environment, resulting in low driving efficiency of the vehicles during automatic driving. Summary of the Invention

[0004] In view of this, the present disclosure provides an obstacle detection method, a vehicle control method, an apparatus, an equipment, and a storage medium.

[0005] One aspect of this disclosure provides an obstacle detection method, comprising:

[0006] The vehicle-related driving environment image is input into the semantic segmentation layer, and the output is an environmental semantic image. The environmental semantic image includes a traffic semantic image block that represents a traffic area in the driving environment. The traffic area is an area suitable for the vehicle to pass through.

[0007] Obstacle detection is performed on the above environmental semantic image to obtain at least one initial obstacle detection result;

[0008] Based on the positional relationship between the initial obstacle detection results and the aforementioned traffic semantic image blocks, candidate obstacle detection results that at least partially overlap with the aforementioned traffic semantic image blocks are determined from at least one of the initial obstacle detection results; and

[0009] The candidate obstacle image blocks associated with the above candidate obstacle detection results in the above driving environment image are input into the traffic obstacle detection model, and the traffic obstacle detection results are output.

[0010] According to embodiments of this disclosure, the above-mentioned obstacle detection of the environmental semantic image to obtain at least one initial obstacle detection result includes:

[0011] The environmental semantic image is input into the first obstacle detection model, and the first detection result is output. The first detection result includes multiple first obstacle classification results and a first confidence level corresponding to each of the multiple first obstacle classification results.

[0012] If multiple first confidence levels meet preset conditions, the multiple first confidence levels are processed according to the information entropy algorithm to obtain the first confidence entropy; and

[0013] Based on the first confidence entropy, the first detection result is determined as the initial obstacle detection result.

[0014] According to embodiments of this disclosure, the semantic segmentation layer is included in the second obstacle detection model, which further includes an image adversarial generative network layer, a difference evaluation layer, and an abnormal obstacle detection layer.

[0015] The above-mentioned obstacle detection of the environmental semantic image to obtain at least one initial obstacle detection result includes:

[0016] The aforementioned environmental semantic image is input into the aforementioned image adversarial generative network layer, and the output is a predicted driving environment image;

[0017] The above driving environment image and the above predicted driving environment image are input into the above difference evaluation layer, and the image difference information is output.

[0018] Based on the above-mentioned abnormal obstacle detection layer, the above-mentioned image difference information, the above-mentioned environmental semantic image, the above-mentioned driving environment image and the above-mentioned predicted driving environment image are processed to obtain at least one of the above-mentioned initial obstacle detection results.

[0019] According to embodiments of this disclosure, the semantic segmentation layer further outputs environmental semantic discreteness information, which characterizes the semantic prediction probability discreteness of pixels in the environmental semantic image. The abnormal obstacle detection layer includes a first convolutional sub-layer, a second convolutional sub-layer, a first fusion sub-layer, and an image decoding sub-layer.

[0020] The above-mentioned processing of the image difference information, the driving environment image, and the predicted driving environment image based on the above-mentioned abnormal obstacle detection layer includes:

[0021] The above-mentioned environmental semantic image, the above-mentioned driving environment image and the above-mentioned predicted driving environment image are input into the above-mentioned first convolutional sub-layer, and the environmental semantic image features, driving environment image features and predicted driving environment image features are output.

[0022] The concatenation result of the above-mentioned environmental semantic image features, the above-mentioned driving environment image features and the above-mentioned predicted driving environment image features is input into the above-mentioned second convolutional sub-layer to obtain the first intermediate fusion feature;

[0023] The first intermediate fusion feature, the image difference information, and the environmental semantic discreteness information are input into the first fusion sub-layer to obtain the second intermediate fusion feature; and

[0024] The second intermediate fusion feature is input into the image decoding sub-layer, and at least one initial obstacle detection result is output.

[0025] According to embodiments of this disclosure, the above-mentioned inputting candidate obstacle image blocks associated with the above-mentioned candidate obstacle detection results from the driving environment image into the traffic obstacle detection model and outputting traffic obstacle detection results includes:

[0026] The above-mentioned candidate obstacle image blocks are input into the obstacle recognition layer of the above-mentioned obstacle detection model, and the candidate obstacle category is output. The above-mentioned obstacle detection model also includes a passage condition prediction layer.

[0027] The candidate obstacle categories are input into the traffic condition prediction layer, which outputs predicted traffic condition parameters corresponding to the candidate obstacle detection results; and

[0028] Based on the predicted passage condition parameters, the above candidate obstacle detection results are determined as the above passage obstacle detection results.

[0029] According to embodiments of this disclosure, the obstacle detection method further includes:

[0030] Based on the above obstacle detection results, an obstacle detection message is generated; and

[0031] Send the aforementioned obstacle detection message to the target client.

[0032] Another aspect of this disclosure provides a vehicle control method, comprising:

[0033] Collect images of the vehicle's driving environment;

[0034] The obstacle detection method provided in this embodiment processes the above-mentioned driving environment image to obtain a traffic obstacle detection result; and

[0035] Based on the above obstacle detection results, control the above vehicles to perform movement operations.

[0036] Another aspect of this disclosure provides an obstacle detection device, comprising:

[0037] An environmental semantic image acquisition module is used to input a vehicle-related driving environment image into a semantic segmentation layer and output an environmental semantic image, wherein the environmental semantic image includes a traffic semantic image block representing a traffic area in the driving environment, and the traffic area is an area suitable for the vehicle to pass through.

[0038] The initial obstacle detection result acquisition module is used to perform obstacle detection on the above environmental semantic image and obtain at least one initial obstacle detection result;

[0039] The candidate obstacle detection result acquisition module is configured to determine, based on the positional relationship between the initial obstacle detection results and the access semantic image patch, candidate obstacle detection results that at least partially overlap with the access semantic image patch from at least one of the initial obstacle detection results; and

[0040] The obstacle detection result acquisition module is used to input the candidate obstacle image blocks associated with the candidate obstacle detection results in the above driving environment image into the obstacle detection model and output the obstacle detection result.

[0041] Another aspect of this disclosure provides a vehicle control device, comprising:

[0042] The acquisition module is used to acquire images of the vehicle's driving environment.

[0043] An obstacle detection module is configured to process the driving environment image according to any one of claims 1 to 6, and obtain a traffic obstacle detection result; and

[0044] The control module is used to control the vehicle to perform movement operations based on the above-mentioned obstacle detection results.

[0045] Another aspect of this disclosure provides an electronic device comprising:

[0046] One or more processors;

[0047] Memory, used to store one or more programs.

[0048] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described above.

[0049] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the method described above.

[0050] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, are used to implement the method described above.

[0051] According to embodiments of this disclosure, by determining candidate obstacle detection results that at least partially overlap with the traffic semantic image block from the initial obstacle detection results, initial obstacle detection results unrelated to the vehicle's traffic area in the driving environment image can be removed, thereby reducing the amount of data that the subsequent traffic obstacle detection model needs to process. At the same time, by leveraging the image classification performance of the traffic obstacle detection model to obtain traffic obstacle detection results, the prediction accuracy for traffic conditions involving crushing obstacles can be improved, solving the technical problem of low obstacle detection accuracy and achieving the technical effect of improving the accuracy of subsequent vehicle control. Attached Figure Description

[0052] The above and other objects, features, and advantages of this disclosure will become clearer from the following description of embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0053] Figure 1 This illustration schematically shows an exemplary system architecture to which obstacle detection methods and apparatus can be applied according to embodiments of the present disclosure;

[0054] Figure 2 A flowchart illustrating an obstacle detection method according to an embodiment of the present disclosure is shown schematically.

[0055] Figure 3 A schematic diagram of a second obstacle detection model according to an embodiment of the present disclosure is shown.

[0056] Figure 4 A schematic diagram of an abnormal obstacle detection layer according to an embodiment of the present disclosure is shown.

[0057] Figure 5 A schematic diagram of a traffic obstacle detection model according to an embodiment of the present disclosure is shown.

[0058] Figure 6 A flowchart illustrating a vehicle control method according to an embodiment of the present disclosure is shown schematically.

[0059] Figure 7 This diagram schematically illustrates an application scenario of a vehicle control method according to an embodiment of the present disclosure.

[0060] Figure 8 A block diagram of an obstacle detection device according to an embodiment of the present disclosure is shown schematically;

[0061] Figure 9 A block diagram of a vehicle control device according to an embodiment of the present disclosure is schematically shown; and

[0062] Figure 10 A block diagram of an electronic device suitable for implementing an obstacle detection method and a vehicle control method according to embodiments of the present disclosure is shown schematically. Detailed Implementation

[0063] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0064] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0065] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0066] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).

[0067] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.

[0068] In the embodiments disclosed herein, user authorization or consent is obtained before acquiring or collecting user personal information.

[0069] In applications such as autonomous driving and smart logistics, vehicles with autonomous driving capabilities, such as unmanned vehicles and intelligent handling robots, can use detection devices such as cameras and LiDAR to collect information about their driving environment. By processing the collected environmental information, they can generate obstacle detection results, which instruct the vehicle to replan its driving path for detected obstacles, thereby improving driving safety and efficiency. However, the inventors discovered that the accuracy of obstacle detection during vehicle operation is low, and there is a tendency for missed or false detections, which negatively impacts driving efficiency and safety.

[0070] Embodiments of this disclosure provide an obstacle detection method, comprising: inputting a vehicle-related driving environment image into a semantic segmentation layer to output an environmental semantic image, wherein the environmental semantic image includes a traffic semantic image block representing a traffic area in the driving environment, the traffic area being an area suitable for vehicle passage; performing obstacle detection on the environmental semantic image to obtain at least one initial obstacle detection result; determining, based on the positional relationship between the initial obstacle detection result and the traffic semantic image block, a candidate obstacle detection result that at least partially overlaps with the traffic semantic image block from the at least one initial obstacle detection result; and inputting the candidate obstacle image block associated with the candidate obstacle detection result in the driving environment image into a traffic obstacle detection model to output a traffic obstacle detection result.

[0071] According to embodiments of this disclosure, by identifying candidate obstacle detection results that at least partially overlap with the traffic semantic image block from the initial obstacle detection results, initial obstacle detection results unrelated to the vehicle's traffic area in the driving environment image can be removed, thereby reducing the amount of data that the subsequent traffic obstacle detection model needs to process. At the same time, by leveraging the image classification performance of the traffic obstacle detection model to obtain traffic obstacle detection results, the prediction accuracy for traffic conditions involving crushing obstacles can be improved, obstacle detection accuracy can be increased, and the precision of subsequent vehicle control can be enhanced.

[0072] Figure 1 An exemplary system architecture 100, in which obstacle detection methods and apparatus can be applied according to embodiments of this disclosure, is illustrated schematically. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0073] like Figure 1As shown, the system architecture 100 according to this embodiment may include vehicles 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between vehicles 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0074] Users can use vehicles 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Processors can be installed on vehicles 101, 102, and 103, and various communication client applications can also be installed, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software (for example only).

[0075] Vehicles 101, 102, and 103 can be any type of vehicle with autonomous driving or intelligent driving assistance functions, including but not limited to passenger cars, intelligent handling robots, unmanned inspection vehicles, etc.

[0076] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using vehicles 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the vehicles.

[0077] It should be noted that the obstacle detection method provided in this embodiment can generally be executed by vehicles 101, 102, or 103, or by other vehicles different from vehicles 101, 102, or 103. Correspondingly, the obstacle detection device provided in this embodiment can generally be installed in vehicles 101, 102, or 103, or in other vehicles different from vehicles 101, 102, or 103. Alternatively, the obstacle detection method provided in this embodiment can also be executed by server 105. Correspondingly, the obstacle detection device provided in this embodiment can also be installed in server 105. The obstacle detection method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with vehicles 101, 102, 103, and / or server 105. Correspondingly, the obstacle detection device provided in this embodiment can also be installed in a server or server cluster that is different from server 105 and capable of communicating with vehicles 101, 102, 103, and / or server 105.

[0078] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0079] Figure 2 A flowchart illustrating an obstacle detection method according to an embodiment of the present disclosure is shown schematically.

[0080] like Figure 2 As shown, the obstacle detection method includes operations S210 to S240.

[0081] In operation S210, the driving environment image related to the vehicle is input to the semantic segmentation layer, and the environment semantic image is output. The environment semantic image includes a traffic semantic image block that represents the traffic area in the driving environment. The traffic area is an area suitable for vehicle traffic.

[0082] According to embodiments of this disclosure, the vehicle can be any type of transportation, such as a passenger car, a truck, an automated guided vehicle, an intelligent inspection robot, etc. Embodiments of this disclosure do not limit the type of vehicle. The driving environment image can characterize the vehicle's driving environment, such as roads, buildings, pedestrians, traffic signs, etc.

[0083] According to embodiments of this disclosure, the semantic segmentation layer can be constructed based on an image semantic segmentation algorithm, and the environmental semantic image can be an image obtained by performing image semantic segmentation on a driving environment image. The environmental semantic image can contain image regions representing any type of driving environment information such as buildings, roads, and obstacles, obtained after image semantic segmentation.

[0084] According to embodiments of this disclosure, the traffic semantic image block can be a semantic image block in the environmental semantic image that corresponds to the area where the vehicle can drive. The area where the vehicle can drive can be, for example, a lane area, a vehicle operation area, etc.

[0085] In operation S220, obstacle detection is performed on the environmental semantic image to obtain at least one initial obstacle detection result.

[0086] According to the embodiments of this disclosure, the initial obstacle detection result can be any type of detection result such as a detection frame, obstacle type, size, or image block corresponding to an obstacle in the driving environment. The embodiments of this disclosure do not limit the specific type of the initial obstacle detection result.

[0087] According to embodiments of this disclosure, initial obstacle detection results can be obtained based on any type of neural network algorithm. For example, environmental semantic images can be processed based on a multilayer perceptron algorithm, but it is not limited to this. Initial obstacle detection results can also be obtained based on other types of neural network algorithms. Embodiments of this disclosure do not limit this.

[0088] In operation S230, based on the positional relationship between the initial obstacle detection results and the passage semantic image block, candidate obstacle detection results that at least partially overlap with the passage semantic image block are determined from at least one initial obstacle detection result.

[0089] According to embodiments of this disclosure, the positional relationship between the initial obstacle detection result and the passage semantic image block can be determined by the pixel position corresponding to the initial obstacle detection result and the pixel position corresponding to the passage semantic image block. However, it is not limited to this; the positional relationship between the initial obstacle detection result and the passage semantic image block can also be determined based on other methods, such as determining the positional relationship based on the geographical coordinates of the obstacle corresponding to the initial obstacle detection result and the geographical location of the passage area. Embodiments of this disclosure do not limit the specific method for determining the positional relationship between the initial obstacle detection result and the passage semantic image block.

[0090] According to embodiments of this disclosure, candidate obstacle detection results that at least partially overlap with the traffic semantic image block can characterize obstacles that affect the vehicle's driving. Therefore, determining candidate obstacle detection results from at least one initial obstacle detection result can at least remove initial obstacle detection results that do not affect the vehicle's driving, thereby reducing the amount of data processing required for subsequent traffic obstacle detection models and saving computational overhead generated during the obstacle detection process.

[0091] In operation S240, the candidate obstacle image blocks associated with the candidate obstacle detection results in the driving environment image are input into the traffic obstacle detection model, and the traffic obstacle detection results are output.

[0092] According to embodiments of this disclosure, the obstacle detection model can be constructed based on any type of neural network algorithm, such as an attention network algorithm, or other types of neural network algorithms. The embodiments of this disclosure do not limit the specific algorithm type for constructing the obstacle detection model.

[0093] According to embodiments of this disclosure, the obstacle detection results can indicate the passage conditions for a vehicle to run over the obstacle. The obstacle detection results can include detection results corresponding to obstacles that a vehicle can pass over, such as empty cardboard boxes, tree branches, etc. Therefore, based on the obstacle detection results, the vehicle can be instructed to run over obstacles in the driving environment to improve vehicle efficiency. Alternatively, the obstacle detection results can also instruct the vehicle to detour around obstacles, thereby avoiding collisions with signs or other obstacles that could cause safety accidents.

[0094] According to an embodiment of this disclosure, operation S220, which performs obstacle detection on the environmental semantic image to obtain at least one initial obstacle detection result, may include the following operations.

[0095] An environmental semantic image is input into a first obstacle detection model, which outputs a first detection result. The first detection result includes multiple first obstacle classification results and a first confidence level corresponding to each of the multiple first obstacle classification results. If the multiple first confidence levels meet preset conditions, the multiple first confidence levels are processed according to the information entropy algorithm to obtain a first confidence entropy. Based on the first confidence entropy, the first detection result is determined as the initial obstacle detection result.

[0096] According to embodiments of this disclosure, the first obstacle detection model can be constructed based on an object detection algorithm, such as an attention network algorithm, or it can be constructed based on other types of neural network algorithms, such as residual network algorithms. Embodiments of this disclosure do not limit the specific algorithm type used to construct the first obstacle detection model.

[0097] According to embodiments of this disclosure, the first obstacle classification result can characterize the category of the obstacle, such as pedestrians, bicycles, roadside trees, etc.

[0098] According to embodiments of this disclosure, a first confidence level can characterize the predicted probability of a first obstacle classification result. A situation where a preset condition is met can characterize a situation where it is difficult to determine the obstacle classification result using multiple first confidence levels. Whether the preset condition is met can be determined based on the difference values ​​between multiple first confidence levels. For example, setting the difference value between multiple first confidence levels to be less than or equal to a preset value can be used to determine that the preset condition is met. Alternatively, setting the difference value between M out of N first confidence levels to be less than or equal to a preset value can also be used to determine that the preset condition is met, where N ≥ M > 1. Embodiments of this disclosure do not limit the specific method of setting the preset condition.

[0099] According to embodiments of this disclosure, by processing multiple first confidence levels using an information entropy algorithm, the resulting first confidence entropy can characterize the uncertainty of the first obstacle detection model's classification result for the first obstacle. Thus, when the first confidence entropy meets a preset entropy threshold, the first detection result can be determined as the initial obstacle detection result. This can at least improve the initial detection of abnormal obstacle categories, avoiding the problem of missed or false detections due to difficulty in determining the obstacle category. Furthermore, the subsequent traffic obstacle detection model can further accurately detect the category of obstacles in the driving environment image, thereby improving obstacle detection accuracy.

[0100] According to embodiments of this disclosure, a semantic segmentation layer is included in a second obstacle detection model, which further includes an image adversarial generative network layer, a difference evaluation layer, and an abnormal obstacle detection layer.

[0101] Figure 3 A schematic diagram of a second obstacle detection model according to an embodiment of the present disclosure is shown.

[0102] like Figure 3 As shown, the second obstacle detection model 300 may include a semantic segmentation layer 310, an image adversarial generative network layer 320, a difference evaluation layer 330, and an abnormal obstacle detection layer 340.

[0103] According to embodiments of this disclosure, performing obstacle detection on an environmental semantic image to obtain at least one initial obstacle detection result may further include the following operations.

[0104] The environmental semantic image is input into the image adversarial generative network layer, which outputs a predicted driving environment image. The driving environment image and the predicted driving environment image are input into the difference evaluation layer, which outputs image difference information. The abnormal obstacle detection layer processes the image difference information, environmental semantic image, driving environment image and predicted driving environment image to obtain at least one initial obstacle detection result.

[0105] like Figure 3 As shown, the driving environment image 301 can be input into the semantic segmentation layer 310, and the output is an environment semantic image 302. The environment semantic image 302 is then input into the image adversarial generative network layer 310, and the output is a predicted driving environment image 303. The image adversarial generative network layer 320 can be constructed based on the cGAN (Conditional Generative Adversarial Network) algorithm. Based on the trained image adversarial generative network layer 320, it can generate a relatively realistic predicted driving environment image with pixel-to-pixel correspondence according to the semantic mapping relationship between the semantic environment image 302 and the driving environment image 301, thereby realizing the resynthesis of the driving environment image.

[0106] like Figure 3As shown, the driving environment image 301 and the predicted driving environment image 303 can also be input into the difference evaluation layer 330, which outputs image difference information 304. The difference evaluation layer 330 can be constructed based on the Learned Perceptual Image Patch Similarity (LPIPS) algorithm. The predicted driving environment image 303 may miss basic information such as the color and appearance of objects in the driving environment, such as obstacles, in the driving environment image 301. By inputting the driving environment image 301 and the predicted driving environment image 303 into the difference evaluation layer 330, the pixel values ​​of the driving environment image 301 and the predicted driving environment image 303 can be compared, and the image difference information 304 can be used to characterize the perceptual difference between the driving environment image 301 and the predicted driving environment image 303, thereby improving the detection of misclassification and anomaly classification of obstacles.

[0107] like Figure 3 As shown, based on the image difference information, environmental semantic image, driving environment image and predicted driving environment image processed by the abnormal obstacle detection layer, the image difference information 304, environmental semantic image 302, driving environment image 301 and predicted driving environment image 303 can be input into the abnormal obstacle detection layer 340 to obtain at least one initial obstacle detection result 305.

[0108] According to embodiments of this disclosure, the abnormal obstacle detection layer can be constructed based on an image encoder and an image decoder. The image encoder can generate an image code representing the driving environment by at least fusing image difference information, environmental semantic image, driving environment image and predicted driving environment image, and the image is decoded by the decoder to obtain accurate initial obstacle detection results.

[0109] According to embodiments of this disclosure, the semantic segmentation layer also outputs environmental semantic discreteness information, which characterizes the semantic prediction probability discreteness of pixels in the environmental semantic image. The abnormal obstacle detection layer includes a first convolutional sub-layer, a second convolutional sub-layer, a first fusion sub-layer, and an image decoding sub-layer.

[0110] According to embodiments of this disclosure, environmental semantic discreteness information may include pixel category prediction probability uncertainty information and the difference distance between pixel category prediction probabilities. For example, pixel category prediction probability uncertainty information can be represented by the following formula (1).

[0111]

[0112] In formula (1), Hx represents the uncertainty information of pixel category prediction probability, p(c) represents the prediction probability corresponding to the pixel in the environmental semantic image, and c represents the prediction probability corresponding to the pixel (i.e. pixel category prediction probability).

[0113] For example, the difference distance between pixel category prediction probabilities can also be represented by the following formula (2).

[0114]

[0115] In formula (2), Dx represents the difference distance between pixel category prediction probabilities, and p(c') represents the average pixel category prediction probabilities of multiple pixels in the environmental semantic image.

[0116] According to embodiments of this disclosure, the uncertainty of the prediction of obstacles in the environmental semantic image can be quantitatively represented by environmental semantic discreteness information. This allows the abnormal obstacle detection layer to fully learn the uncertainty information of the prediction of obstacles, thereby improving the detection accuracy of abnormal obstacles that are difficult to classify, and thus improving the accuracy of obstacle detection.

[0117] Figure 4 A schematic diagram of an abnormal obstacle detection layer according to an embodiment of the present disclosure is shown.

[0118] like Figure 4 As shown, the abnormal obstacle detection layer 400 may include a first convolutional sublayer 410, a second convolutional sublayer 420, a first fusion sublayer 430, and an image decoding sublayer 440.

[0119] According to embodiments of this disclosure, processing image difference information, driving environment image, and predicted driving environment image based on the abnormal obstacle detection layer may include the following operations.

[0120] The environmental semantic image, the driving environment image, and the predicted driving environment image are input into the first convolutional sub-layer, which outputs environmental semantic image features, driving environment image features, and predicted driving environment image features. The concatenated result of the environmental semantic image features, driving environment image features, and predicted driving environment image features is input into the second convolutional sub-layer to obtain the first intermediate fusion feature. The first intermediate fusion feature, image difference information, and environmental semantic discreteness information are input into the first fusion sub-layer to obtain the second intermediate fusion feature. The second intermediate fusion feature is input into the image decoding sub-layer to output at least one initial obstacle detection result.

[0121] like Figure 4As shown, the environmental semantic image 401, the driving environment image 402, and the predicted driving environment image 403 can be input into the first convolutional sub-layer 410, which outputs environmental semantic image features 401', driving environment image features 402', and predicted driving environment image features 403'. The first convolutional sub-layer 410 can be constructed based on a convolutional neural network algorithm, such as a VGG (Visual Geometry Group) network structure, to extract the depth image features of each image from the environmental semantic image, the driving environment image, and the predicted driving environment image.

[0122] like Figure 4 As shown, the concatenated result of environmental semantic image features 401', driving environment image features 402', and predicted driving environment image features 403' can be input into the second convolutional sub-layer 420, and the first intermediate fusion feature 406 can be output. The second convolutional sub-layer 420 can be constructed based on a convolutional neural network algorithm, for example, it can be constructed based on a single-layer convolutional neural network layer.

[0123] like Figure 4 As shown, the first intermediate fusion feature 406, along with the environmental semantic discreteness information 404 and the image difference information 405, are input into the first fusion sub-layer 430, and the second intermediate fusion feature 407 is output.

[0124] According to embodiments of this disclosure, the first fusion sublayer 430 can calculate the feature correlation between the first intermediate fusion feature 406 and the environmental semantic discreteness information 404, as well as the feature correlation between the first intermediate fusion feature 406 and the image difference information 405. This allows for a more accurate characterization of the feature data with category prediction uncertainty in the first intermediate fusion feature 406. Consequently, the image decoding sublayer can accurately predict the abnormal segmentation regions in the environmental semantic image, reducing the number of missed detections in the image semantic segmentation stage. This improves the recall quantity and recall accuracy for abnormal obstacle types, such as low obstacles with long-tail data characteristics and other unknown types of obstacles, thereby reducing the problem of missed detections for abnormal obstacle types.

[0125] According to embodiments of this disclosure, the first fusion sub-layer can be constructed based on a correlation algorithm, such as a correlation coefficient algorithm. However, it is not limited to this; the first fusion sub-layer can also be constructed based on an attention network algorithm. Embodiments of this disclosure do not limit the specific algorithm type used to construct the first fusion sub-layer.

[0126] According to an embodiment of this disclosure, operation S240, inputting candidate obstacle image blocks associated with candidate obstacle detection results from the driving environment image into the traffic obstacle detection model, and outputting traffic obstacle detection results may include the following operations.

[0127] The candidate obstacle image patch is input into the obstacle recognition layer of the obstacle detection model, and the candidate obstacle category is output. The obstacle detection model also includes a passage condition prediction layer. The candidate obstacle category is input into the passage condition prediction layer, and the predicted passage condition parameters corresponding to the candidate obstacle detection results are output. Based on the predicted passage condition parameters, the candidate obstacle detection results are determined as the passage obstacle detection results.

[0128] According to embodiments of this disclosure, candidate obstacle image blocks corresponding to the candidate obstacle detection results can be segmented from a driving environment image based on the candidate obstacle detection results. The candidate obstacle image blocks can represent candidate obstacles.

[0129] According to embodiments of this disclosure, the obstacle recognition layer can be constructed based on an image classification algorithm, such as a multilayer perceptron algorithm or a convolutional neural network algorithm. The embodiments of this disclosure do not limit the specific algorithm type for constructing the obstacle recognition layer.

[0130] According to embodiments of this disclosure, the traffic condition prediction layer can process the identified candidate obstacle categories to obtain a traffic obstacle detection result indicating whether a vehicle can drive over the candidate obstacle. The traffic obstacle detection result can, for example, label pixels in the driving environment image corresponding to abnormal obstacle types as 1 and 0. Pixels labeled as 1 indicate that the vehicle can drive over the obstacle, while pixels labeled as 0 indicate that the vehicle needs to detour around the obstacle.

[0131] According to embodiments of this disclosure, the obstacle detection model can be used to further accurately classify candidate obstacles that represent abnormal types, so as to accurately identify obstacles such as manhole covers and empty cardboard boxes in the passage area that can be crushed and passed, thereby improving the accuracy of subsequent vehicle motion control and improving vehicle driving efficiency.

[0132] Figure 5 A schematic diagram of a traffic obstacle detection model according to an embodiment of the present disclosure is shown.

[0133] like Figure 5As shown, the obstacle detection model 500 may include an obstacle recognition layer 510 and a passage condition prediction layer 520. Candidate obstacle image patches 501 can be input to the obstacle recognition layer of the obstacle detection model 500, outputting candidate obstacle categories 502. The candidate obstacle categories 502 are then input to the passage condition prediction layer 520, outputting predicted passage condition parameters 503 corresponding to the candidate obstacle detection results.

[0134] According to embodiments of this disclosure, the obstacle detection method may further include the following operations.

[0135] Based on the obstacle detection results, an obstacle detection message is generated; and the obstacle detection message is sent to the target client.

[0136] According to embodiments of this disclosure, the obstacle detection results can accurately identify obstacles such as manhole covers and empty cardboard boxes that can be crushed in the passage area. By sending an obstacle detection message containing the obstacle detection results to a target client or server related to vehicle route planning, the target client or server can improve the accuracy of vehicle route planning and distribute the generated vehicle route to one or more other vehicles, thereby improving the passage efficiency of multiple vehicles.

[0137] Figure 6 A flowchart illustrating a vehicle control method according to an embodiment of the present disclosure is shown schematically.

[0138] like Figure 6 As shown, the vehicle control method may include operations S610 to S630.

[0139] When operating the S610, images of the driving environment related to the vehicle are acquired.

[0140] In operation S620, the driving environment image is processed according to the obstacle detection method provided in the embodiments of this disclosure to obtain the obstacle detection result.

[0141] When operating the S630, the vehicle is controlled to perform movement operations based on the obstacle detection results.

[0142] According to embodiments of this disclosure, the vehicle is controlled to perform movement operations based on the obstacle detection results. For example, a vehicle planning path may be generated based on the obstacle detection results, and the vehicle movement may be controlled based on the vehicle planning path.

[0143] Figure 7 The diagram illustrates an application scenario of a vehicle control method according to an embodiment of the present disclosure.

[0144] like Figure 7As shown, in this application scenario 700, the vehicle 710 can acquire images of the driving environment through an image acquisition device to obtain a driving environment image. Simultaneously, the vehicle 710 can also process the driving environment image based on the obstacle detection method provided in this embodiment to obtain a traffic obstacle detection result. The traffic obstacle detection result can indicate that the obstacle 711 is an empty cardboard box in the driving path. Based on the traffic obstacle detection result, a driving path that passes over the obstacle 711 can be generated, thus allowing the vehicle 710 to perform movement operations according to the driving path.

[0145] Figure 8 A block diagram of an obstacle detection apparatus according to an embodiment of the present disclosure is shown schematically.

[0146] like Figure 8 As shown, the obstacle detection device 800 includes an environmental semantic image acquisition module 810, an initial obstacle detection result acquisition module 820, a candidate obstacle detection result acquisition module 830, and a passage obstacle detection result acquisition module 840.

[0147] The environmental semantic image acquisition module 810 is used to input vehicle-related driving environment images into the semantic segmentation layer and output environmental semantic images. The environmental semantic images include traffic semantic image blocks that represent traffic areas in the driving environment, and the traffic areas are areas suitable for vehicle traffic.

[0148] The initial obstacle detection result acquisition module 820 is used to perform obstacle detection on the environmental semantic image and obtain at least one initial obstacle detection result.

[0149] The candidate obstacle detection result acquisition module 830 is used to determine, from at least one initial obstacle detection result, a candidate obstacle detection result that at least partially overlaps with the passage semantic image block, based on the positional relationship between the initial obstacle detection result and the passage semantic image block.

[0150] The obstacle detection result acquisition module 840 is used to input the candidate obstacle image blocks associated with the candidate obstacle detection results in the driving environment image into the obstacle detection model and output the obstacle detection results.

[0151] According to embodiments of this disclosure, the initial obstacle detection result acquisition module includes: a first detection result acquisition submodule, a first confidence entropy acquisition submodule, and a first initial obstacle detection result acquisition submodule.

[0152] The first detection result acquisition submodule is used to input the environmental semantic image into the first obstacle detection model and output the first detection result. The first detection result includes multiple first obstacle classification results and a first confidence level corresponding to each of the multiple first obstacle classification results.

[0153] The first confidence entropy acquisition submodule is used to process multiple first confidences according to the information entropy algorithm to obtain the first confidence entropy when multiple first confidences meet preset conditions.

[0154] The first initial obstacle detection result acquisition submodule is used to determine the first detection result as the initial obstacle detection result based on the first confidence entropy.

[0155] According to embodiments of this disclosure, a semantic segmentation layer is included in a second obstacle detection model, which further includes an image adversarial generative network layer, a difference evaluation layer, and an abnormal obstacle detection layer.

[0156] The initial obstacle detection result acquisition module includes: a predicted driving environment image generation submodule, an image difference information generation submodule, and a second initial obstacle detection result acquisition submodule.

[0157] The Predicted Driving Environment Image Generation Submodule is used to predict driving environment images by inputting the environmental semantic image into the Image Generative Adversarial Network layer and outputting the predicted driving environment image.

[0158] The image difference information generation submodule is used to input the driving environment image and the predicted driving environment image into the difference evaluation layer and output image difference information;

[0159] The second initial obstacle detection result acquisition submodule is used to obtain at least one initial obstacle detection result based on the image difference information processed by the abnormal obstacle detection layer, the environmental semantic image, the driving environment image, and the predicted driving environment image.

[0160] According to embodiments of this disclosure, the semantic segmentation layer also outputs environmental semantic discreteness information, which characterizes the semantic prediction probability discreteness of pixels in the environmental semantic image. The abnormal obstacle detection layer includes a first convolutional sub-layer, a second convolutional sub-layer, a first fusion sub-layer, and an image decoding sub-layer.

[0161] The second initial obstacle detection result acquisition submodule includes: a feature extraction unit, a first intermediate fusion feature acquisition unit, a second intermediate fusion feature acquisition unit, and an initial obstacle detection result acquisition unit.

[0162] The feature extraction unit is used to input the environmental semantic image, the driving environment image, and the predicted driving environment image into the first convolutional sub-layer, and output the environmental semantic image features, the driving environment image features, and the predicted driving environment image features.

[0163] The first intermediate fusion feature acquisition unit is used to input the concatenation result of environmental semantic image features, driving environment image features and predicted driving environment image features into the second convolutional sub-layer to obtain the first intermediate fusion feature.

[0164] The second intermediate fusion feature acquisition unit is used to input the first intermediate fusion feature, image difference information and environmental semantic discreteness information into the first fusion sub-layer to obtain the second intermediate fusion feature.

[0165] The initial obstacle detection result acquisition unit is used to input the second intermediate fusion feature into the image decoding sub-layer and output at least one initial obstacle detection result.

[0166] According to embodiments of this disclosure, the obstacle detection result acquisition module includes: a candidate obstacle category acquisition submodule, a predicted passage condition parameter acquisition submodule, and an obstacle detection result acquisition submodule.

[0167] The candidate obstacle category acquisition submodule is used to input candidate obstacle image patches into the obstacle recognition layer of the access obstacle detection model and output candidate obstacle categories. The access obstacle detection model also includes a access condition prediction layer.

[0168] The predicted passage condition parameter acquisition submodule is used to input candidate obstacle categories into the passage condition prediction layer and output predicted passage condition parameters corresponding to the candidate obstacle detection results.

[0169] The obstacle detection result acquisition submodule is used to determine the candidate obstacle detection results as the actual obstacle detection results based on the predicted passage condition parameters.

[0170] According to embodiments of this disclosure, the obstacle detection device further includes a message generation module and a message sending module.

[0171] The message generation module is used to generate obstacle detection messages based on the obstacle detection results.

[0172] The message sending module is used to send obstacle detection messages to the target client.

[0173] Figure 9 A block diagram of a vehicle control device according to an embodiment of the present disclosure is shown schematically.

[0174] like Figure 9 As shown, the vehicle control device 900 includes: a data acquisition module 910, an obstacle detection module 920, and a control module 930.

[0175] The acquisition module 910 is used to acquire images of the driving environment related to the vehicle.

[0176] The obstacle detection module 920 is used to process the driving environment image according to the obstacle detection method provided in the embodiments of this disclosure to obtain the obstacle detection result.

[0177] The control module 930 is used to control the vehicle to perform movement operations based on the obstacle detection results.

[0178] Any one or more of the modules, submodules, and units according to embodiments of this disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, and units according to embodiments of this disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, and units according to embodiments of this disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented by hardware or firmware in any other reasonable manner by integrating or packaging the circuitry, or implemented in any one of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of this disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.

[0179] For example, any and multiple modules of the environmental semantic image acquisition module 810, the initial obstacle detection result acquisition module 820, the candidate obstacle detection result acquisition module 830, and the passage obstacle detection result acquisition module 840, or the acquisition module 910, the obstacle detection module 920, and the control module 930, can be combined into one module / submodule / unit, or any one of these modules / submodules / units can be split into multiple modules / submodules / units. Alternatively, at least some of the functions of one or more of these modules / submodules / units can be combined with at least some of the functions of other modules / submodules / units and implemented in one module / submodule / unit. According to embodiments of this disclosure, at least one of the environmental semantic image acquisition module 810, the initial obstacle detection result acquisition module 820, the candidate obstacle detection result acquisition module 830, and the passage obstacle detection result acquisition module 840, or the acquisition module 910, the obstacle detection module 920, and the control module 930, can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of them. Alternatively, at least one of the following modules may be implemented, at least partially, as a computer program module: environmental semantic image acquisition module 810, initial obstacle detection result acquisition module 820, candidate obstacle detection result acquisition module 830, and passage obstacle detection result acquisition module 840; or acquisition module 910, obstacle detection module 920, and control module 930. When the computer program module is run, it can perform the corresponding function.

[0180] It should be noted that the obstacle detection device part in the embodiments of this disclosure corresponds to the obstacle detection method part in the embodiments of this disclosure. For a detailed description of the obstacle detection device part, please refer to the obstacle detection method part, which will not be repeated here.

[0181] It should be noted that the vehicle control device part in the embodiments of this disclosure corresponds to the vehicle control method part in the embodiments of this disclosure. For a detailed description of the vehicle control device part, please refer to the vehicle control method part, which will not be repeated here.

[0182] Figure 10 A block diagram of an electronic device suitable for implementing an obstacle detection method and a vehicle control method according to embodiments of the present disclosure is shown schematically. Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0183] like Figure 10 As shown, an electronic device 1000 according to an embodiment of the present disclosure includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage portion 1008 into a random access memory (RAM) 1003. The processor 1001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1001 may also include onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0184] RAM 1003 stores various programs and data required for the operation of electronic device 1000. Processor 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Processor 1001 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 1002 and / or RAM 1003. It should be noted that the programs may also be stored in one or more memories other than ROM 1002 and RAM 1003. Processor 1001 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0185] According to embodiments of this disclosure, the electronic device 1000 may further include an input / output (I / O) interface 1005, which is also connected to a bus 1004. The system 1000 may also include one or more of the following components connected to the input / output (I / O) interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output (I / O) interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1010 as needed so that computer programs read from it can be installed into the storage section 1008 as needed.

[0186] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by processor 1001, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0187] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0188] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0189] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 1002 and / or RAM 1003 described above and / or one or more memories other than ROM 1002 and RAM 1003.

[0190] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of this disclosure.

[0191] When the computer program is executed by the processor 1001, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0192] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1009, and / or installed from a removable medium 1011. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0193] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features recited in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not expressly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0195] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. An obstacle detection method, comprising: The vehicle-related driving environment image is input into the semantic segmentation layer, and the output is an environmental semantic image, wherein the environmental semantic image includes a traffic semantic image block representing a traffic area in the driving environment, and the traffic area is an area suitable for the vehicle to pass through; Obstacle detection is performed on the environmental semantic image to obtain at least one initial obstacle detection result; Based on the positional relationship between the initial obstacle detection results and the access semantic image patch, candidate obstacle detection results that at least partially overlap with the access semantic image patch are determined from at least one of the initial obstacle detection results; and The candidate obstacle image patch associated with the candidate obstacle detection result in the driving environment image is input into the traffic obstacle detection model, and the traffic obstacle detection result is output. The semantic segmentation layer is included in the second obstacle detection model, which further includes an image adversarial generative network layer, a difference evaluation layer, and an abnormal obstacle detection layer. The step of performing obstacle detection on the environmental semantic image to obtain at least one initial obstacle detection result includes: The environmental semantic image is input into the image adversarial generative network layer, and a predicted driving environment image is output. The image adversarial generative network layer generates the predicted driving environment image with a pixel-to-pixel correspondence based on the semantic mapping relationship between the environmental semantic image and the driving environment image. The driving environment image and the predicted driving environment image are input to the difference evaluation layer, which outputs image difference information characterizing the perceived difference between the driving environment image and the predicted driving environment image. The difference evaluation layer compares the row pixel values ​​of the driving environment image and the predicted driving environment image. The abnormal obstacle detection layer processes the image difference information, the environmental semantic image, the driving environment image, and the predicted driving environment image to obtain at least one initial obstacle detection result.

2. The obstacle detection method according to claim 1, wherein, The step of performing obstacle detection on the environmental semantic image to obtain at least one initial obstacle detection result includes: The environmental semantic image is input into a first obstacle detection model, and a first detection result is output. The first detection result includes multiple first obstacle classification results and a first confidence level corresponding to each of the multiple first obstacle classification results. When multiple first confidence levels meet preset conditions, the multiple first confidence levels are processed according to the information entropy algorithm to obtain the first confidence entropy; and Based on the first confidence entropy, the first detection result is determined as the initial obstacle detection result.

3. The obstacle detection method according to claim 1, wherein, The semantic segmentation layer also outputs environmental semantic discreteness information, which characterizes the semantic prediction probability discreteness of pixels in the environmental semantic image. The abnormal obstacle detection layer includes a first convolutional sub-layer, a second convolutional sub-layer, a first fusion sub-layer, and an image decoding sub-layer. The step of processing the image difference information, the driving environment image, and the predicted driving environment image according to the abnormal obstacle detection layer includes: The environmental semantic image, the driving environment image, and the predicted driving environment image are input into the first convolutional sub-layer, and the environmental semantic image features, driving environment image features, and predicted driving environment image features are output. The concatenation result of the environmental semantic image features, the driving environment image features, and the predicted driving environment image features is input into the second convolutional sub-layer to obtain the first intermediate fusion feature; The first intermediate fusion feature, the image difference information, and the environmental semantic discreteness information are input into the first fusion sub-layer to obtain the second intermediate fusion feature; and The second intermediate fusion feature is input into the image decoding sub-layer, and at least one of the initial obstacle detection results is output.

4. The obstacle detection method according to claim 1, wherein, The step of inputting candidate obstacle image blocks associated with the candidate obstacle detection results from the driving environment image into the traffic obstacle detection model and outputting traffic obstacle detection results includes: The candidate obstacle image blocks are input into the obstacle recognition layer of the passage obstacle detection model, and the candidate obstacle category is output. The passage obstacle detection model also includes a passage condition prediction layer. The candidate obstacle categories are input into the traffic condition prediction layer, and predicted traffic condition parameters corresponding to the candidate obstacle detection results are output; and Based on the predicted passage condition parameters, the candidate obstacle detection results are determined as the passage obstacle detection results.

5. The obstacle detection method according to claim 1, further comprising: Based on the obstacle detection results, an obstacle detection message is generated; as well as The obstacle detection message is sent to the target client.

6. A vehicle control method, comprising: Collect images of the vehicle's driving environment; The obstacle detection method according to any one of claims 1 to 5 processes the driving environment image to obtain a traffic obstacle detection result; and Based on the obstacle detection results, the vehicle is controlled to perform movement operations.

7. An obstacle detection device, comprising: An environmental semantic image acquisition module is used to input a vehicle-related driving environment image into a semantic segmentation layer and output an environmental semantic image, wherein the environmental semantic image includes a traffic semantic image block representing a traffic area in the driving environment, and the traffic area is an area suitable for the vehicle to pass through; An initial obstacle detection result acquisition module is used to perform obstacle detection on the environmental semantic image and obtain at least one initial obstacle detection result; A candidate obstacle detection result acquisition module is configured to determine, based on the positional relationship between the initial obstacle detection results and the passage semantic image patch, candidate obstacle detection results that at least partially overlap with the passage semantic image patch from at least one of the initial obstacle detection results; and The obstacle detection result acquisition module is used to input candidate obstacle image blocks associated with the candidate obstacle detection results from the driving environment image into the obstacle detection model, and output the obstacle detection results. The semantic segmentation layer is included in the second obstacle detection model, which further includes an image adversarial generative network layer, a difference evaluation layer, and an abnormal obstacle detection layer. The initial obstacle detection result acquisition module includes: The predicted driving environment image generation submodule is used to input an environmental semantic image into the image adversarial generative network layer and output a predicted driving environment image. The image adversarial generative network layer generates the predicted driving environment image with a pixel-to-pixel correspondence based on the semantic mapping relationship between the environmental semantic image and the driving environment image. The image difference information generation submodule is used to input the driving environment image and the predicted driving environment image into the difference evaluation layer, and output image difference information that characterizes the perceived difference between the driving environment image and the predicted driving environment image, wherein the difference evaluation layer compares the row pixel values ​​of the driving environment image and the predicted driving environment image; The second initial obstacle detection result acquisition submodule is used to process the image difference information, the environmental semantic image, the driving environment image and the predicted driving environment image according to the abnormal obstacle detection layer to obtain at least one initial obstacle detection result.

8. A vehicle control device, comprising: The acquisition module is used to acquire images of the vehicle's driving environment. An obstacle detection module is used to process the driving environment image according to any one of claims 1 to 5 to obtain a traffic obstacle detection result; as well as The control module is used to control the vehicle to perform movement operations based on the detection results of the obstacles.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 6.

10. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 6.

11. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 6.