Target detection method and device, vehicle and medium

By using a prediction model in a vehicle to predict edge points of the target to be detected in the driving image, the problem of low speed bump detection accuracy in the prior art is solved, and higher detection accuracy and effect are achieved.

CN120220082APending Publication Date: 2025-06-27BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311828053.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When identifying speed bumps in uneven roads, the prior art is limited by factors such as the lack of obvious characteristics of the center point of the speed bump and camera distortion, resulting in low detection accuracy.

Method used

By acquiring the travel image collected by the vehicle during driving, processing is performed to obtain the image to be detected, and then inputting the image to be detected into the prediction model, predicting the edge point of the target to be detected, and obtaining the position information of the edge point based on the predicted image to improve detection accuracy.

Benefits of technology

By predicting the edge points of the detection target, speed bumps and other uneven road surface characteristics can be more accurately identified, improving the accuracy and effect of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220082A_ABST
    Figure CN120220082A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method and device, a vehicle and a medium, and relates to the technical field of vehicles, and the method comprises the steps: firstly obtaining a driving image collected in the driving process of the vehicle, carrying out the processing of the driving image, obtaining a to-be-detected image containing a to-be-detected target, inputting the to-be-detected image into a prediction model, and carrying out the prediction of the to-be-detected target. Obtaining a prediction image for predicting an edge point of a to-be-detected target in the to-be-detected image, obtaining a first position of the edge point in a first coordinate system where the to-be-detected image is located according to the prediction image, and obtaining a second position of the edge point in a second coordinate system where the to-be-detected image is located according to the first position of the edge point in the prediction image; according to the method and the device, the edge points of the to-be-detected target are predicted, so that the prediction of the to-be-detected target focuses on the local features, and the detection accuracy of the to-be-detected target can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of vehicles, and particularly to an object detection method, device, vehicle and medium. Background Art

[0002] During the driving of a vehicle, relevant algorithms are used to identify uneven road surfaces such as speed bumps and potholed roads, so that the vehicle can make adaptive adjustments in advance, enabling the vehicle to drive over the uneven road surface more smoothly, thereby obtaining a better driving experience. For some uneven road surfaces such as speed bumps, the characteristics of their center points are not obvious. At the same time, due to camera distortion in the camera for photographing the road surface, the above reasons can easily lead to low detection accuracy. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides an object detection method, device, vehicle and medium.

[0004] According to the first aspect of the embodiments of the present disclosure, an object detection method is provided. The method includes: acquiring a driving image collected during the driving of a vehicle, and processing the driving image to obtain a to-be-detected image, where the to-be-detected image is an image including a to-be-detected object; inputting the to-be-detected image into a prediction model to obtain a prediction image for predicting edge points of the to-be-detected object in the to-be-detected image; according to the prediction image, acquiring a first position of the edge points in a first coordinate system where the to-be-detected image is located; and according to the first position where the edge points in the prediction image are located, acquiring a second position of the to-be-detected object in a second coordinate system where the driving image is located.

[0005] Optionally, the number of the prediction images is a preset number, and the preset number of prediction images includes position information of the edge points. The step of, according to the prediction image, acquiring a first position of the edge points in a first coordinate system where the to-be-detected image is located includes: obtaining the first position of the edge points in the first coordinate system where the to-be-detected image is located according to the position information of the edge points in the preset number of prediction images.

[0006] Optionally, the preset number of predicted images includes: a confidence map, a first direction offset map, and a second direction offset map; wherein, the size of the confidence map, the size of the first direction offset map, and the size of the second direction offset map are all the same as the size of the image to be detected, the confidence map includes the confidence corresponding to each pixel among a plurality of pixels on the image to be detected and the first coordinate position of each pixel in the first coordinate system where the image to be detected is located, the confidence corresponding to each pixel is used to represent the confidence of the pixel belonging to the edge of the target to be detected, the first coordinate position includes the first coordinate on the first coordinate axis in the first coordinate system and the second coordinate on the second coordinate axis in the first coordinate system, the first direction offset map includes the third coordinate of each pixel on the first coordinate axis, and the second direction offset map includes the fourth coordinate of each pixel on the second coordinate axis; the obtaining of the first position of the edge point in the first coordinate system where the image to be detected is located according to the position information of the edge point in the preset number of predicted images includes: taking the pixel points on the confidence map with a confidence greater than the preset confidence as edge points; obtaining the first coordinate position of the edge point in the first coordinate system from the confidence map, obtaining the third coordinate of the edge point on the first coordinate axis from the first direction offset map, and obtaining the fourth coordinate of the edge point on the second coordinate axis from the second direction offset map, wherein the first coordinate position includes the first coordinate of the edge point on the first coordinate axis and the second coordinate of the edge point on the second coordinate axis; obtaining the fifth coordinate on the first coordinate axis according to the first coordinate and the third coordinate of the edge point, and obtaining the sixth coordinate on the second coordinate axis according to the second coordinate and the fourth coordinate of the edge point; determining the first position of the edge point in the first coordinate system according to the fifth coordinate and the sixth coordinate of the edge point in the first coordinate system.

[0007] Optionally, the number of the edge points is multiple. Obtaining the second position of the target to be detected in the second coordinate system where the driving image is located according to the first position where the edge points in the predicted image are located includes: converting, through a first coordinate conversion relationship, the first position of each edge point among the multiple edge points in the first coordinate system where the predicted image is located into a third position in the second coordinate system where the driving image is located, where the first coordinate conversion relationship is the conversion relationship between the first coordinate system and the second coordinate system; forming the contour of the target to be detected according to the pixel points corresponding to the third position on the driving image, and obtaining the fourth position where the center point of the contour is located; obtaining the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located.

[0008] Optionally, the method further includes: obtaining, through a preset function, a bounding box that encloses the contour, where the bounding box is a rectangle; obtaining the side length of the bounding box according to the fifth positions of two opposite end points on the bounding box in the second coordinate system; the obtaining the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located includes: obtaining the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located and the side length of the bounding box.

[0009] Optionally, the prediction model is obtained in the following manner: obtaining a sample image to be recognized, where the sample image to be recognized is an image containing a sample target; constructing a sample predicted image according to the sample image to be recognized, where the sample predicted image includes the position information of the pixel points at the edge of the sample target; training an initial model according to the sample image to be recognized to obtain the prediction model.

[0010] Optionally, constructing a sample prediction image based on the sample image to be recognized includes: constructing a confidence map label, a first direction offset map label, and a second direction offset map label based on the sample image to be recognized, where the size of the confidence map label, the size of the first direction offset map label, and the size of the second direction offset map label are all the same as the size of the sample image to be recognized, the confidence map label includes the confidence corresponding to each pixel among a plurality of pixels on the sample image to be recognized and the coordinate position of each pixel in a first coordinate system where the sample image to be recognized is located, the confidence corresponding to each pixel is used to characterize the pixel confidence of the pixel belonging to the edge of the sample target, the first direction offset map label includes the coordinate position of each pixel on the first coordinate axis, and the second direction offset map label includes the coordinate position of each pixel on the second coordinate axis.

[0011] Optionally, training the initial model based on the sample image to be recognized to obtain the prediction model includes: inputting the sample image to be recognized into the initial model to obtain an intermediate confidence map, a first intermediate offset map label, and a second intermediate offset map label; obtaining a first loss function based on the confidence map label and the intermediate confidence map, and obtaining a second loss function based on the first intermediate offset map label, the second intermediate offset map label, the first direction offset map label, and the second direction offset map label; training the initial model based on the first loss function and the second loss function to obtain the prediction model.

[0012] Optionally, obtaining the sample image to be recognized includes: obtaining a driving sample image collected during the driving of the vehicle; processing the driving sample image to obtain the sample image to be recognized.

[0013] Optionally, the size of the sample image to be recognized is smaller than the size of the driving sample image, and the confidence map label is constructed in the following manner: constructing an initial confidence map label, where the pixel values of all pixels on the initial confidence map label are 0; labeling the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; dividing the driving sample image into a preset number of pixel regions, where each pixel region in the preset number of pixel regions corresponds to each pixel on the initial confidence map label; taking the pixel region where any contour point on the contour line is located as a first target pixel region; setting the pixel value of the pixel on the initial confidence map label corresponding to the first target pixel region to 1 to obtain the confidence map label.

[0014] Optionally, the size of the sample image to be recognized is smaller than the size of the driving sample image. The first direction offset map label is constructed in the following manner: construct a first initial offset map label, wherein the pixel values of all pixel points on the first initial offset map label are 0; label the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; divide the driving sample image into a preset number of pixel regions, wherein each pixel region in the preset number of pixel regions corresponds to each pixel point on the first initial offset map label; take the pixel region where any contour point on the contour line is located as a second target pixel region; set the pixel value of the pixel on the first initial offset map label corresponding to the second target pixel region to a first target value to obtain the first direction offset map label.

[0015] Optionally, the first target value is obtained in the following manner: obtain a first distance from each contour point in the second target region to the center point of the second target region to obtain a plurality of first distances; calculate the average value of the plurality of first distances to obtain a first average distance; take the first average distance as the first target value.

[0016] Optionally, the size of the sample image to be recognized is smaller than the size of the driving sample image. The second direction offset map label is constructed in the following manner: construct a second initial offset map label, wherein the pixel values of all pixel points on the second initial offset map label are 0; label the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; divide the driving sample image into a preset number of pixel regions, wherein each pixel region in the preset number of pixel regions corresponds to each pixel point on the second initial offset map label; take the pixel region where any contour point on the contour line is located as a third target pixel region; set the pixel value of the pixel point on the second initial offset map label corresponding to the third target pixel region to a second target value to obtain the second direction offset map label.

[0017] Optionally, the second target value is obtained in the following manner: obtain a second distance from each contour point in the third target region to the center point of the third target region to obtain a plurality of second distances; calculate the average value of the plurality of second distances to obtain a second average distance; take the second average distance as the second target value.

[0018] Optionally, the processing of the driving image to obtain the image to be detected includes: performing a cropping process on the driving image to obtain the image to be detected.

[0019] Optionally, the processing of the driving image to obtain a to-be-detected image includes: performing downsampling processing on the driving image to obtain the to-be-detected image.

[0020] Optionally, the processing of the driving image to obtain a to-be-detected image includes: performing cropping processing on the driving image to obtain a cropped image; performing downsampling processing on the cropped image to obtain the to-be-detected image.

[0021] Optionally, the method further includes: through a second coordinate conversion relationship, converting the second position of the to-be-detected target in a second coordinate system where the driving image is located into a sixth position in a third coordinate system corresponding to the vehicle, where the second coordinate conversion relationship is the conversion relationship between the second coordinate system and the third coordinate system.

[0022] According to a second aspect of the embodiments of the present disclosure, a target detection device is provided. The device includes: a processing module, configured to acquire a driving image collected during the driving of a vehicle and process the driving image to obtain a to-be-detected image, where the to-be-detected image is an image including a to-be-detected target; a prediction module, configured to input the to-be-detected image into a prediction model to obtain a prediction image for predicting edge points of the to-be-detected target in the to-be-detected image; a first position acquisition module, configured to acquire a first position of the edge points in a first coordinate system where the to-be-detected image is located according to the prediction image; a second position acquisition module, configured to obtain a second position of the to-be-detected target in a second coordinate system where the driving image is located according to the first position where the edge points in the prediction image are located.

[0023] According to a third aspect of the embodiments of the present disclosure, a vehicle is provided. The vehicle includes: a processor;

[0024] a memory for storing instructions executable by the processor; where the processor is configured to implement the steps of the target detection method described in the first aspect when executing the instructions.

[0025] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the target detection method provided in the first aspect of the present disclosure are implemented.

[0026] The present disclosure provides an object detection method, apparatus, vehicle, and medium. First, driving images collected during the driving of a vehicle are obtained, and the driving images are processed to obtain a to-be-detected image including a to-be-detected object. Then, the to-be-detected image is input into a prediction model to obtain a prediction image for predicting edge points of the to-be-detected object in the to-be-detected image. Next, according to the prediction image, a first position of the edge points in a first coordinate system where the to-be-detected image is located is obtained, and according to the first position of the edge points in the prediction image, a second position of the to-be-detected object in a second coordinate system where the driving image is located is obtained. In the present disclosure, by predicting the edge points of the to-be-detected object, the prediction of the to-be-detected object pays attention to local features, which can improve the accuracy of detecting the to-be-detected object.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0029] Figure 1 is an image captured by a camera under distortion conditions;

[0030] Figure 2 is a schematic diagram of an identification result in the related art;

[0031] Figure 3 is a schematic diagram of an identification result in the related art;

[0032] Figure 4 is a flowchart of an object detection method shown according to an exemplary embodiment;

[0033] Figure 5 is a flowchart of an object detection method shown according to another exemplary embodiment;

[0034] Figure 6 is a schematic diagram of a driving image;

[0035] Figure 7 is a schematic diagram of a to-be-detected image;

[0036] Figure 8 is a flowchart of an object detection method shown according to another exemplary embodiment;

[0037] Figure 9 is a flowchart of an object detection method shown according to another exemplary embodiment;

[0038] Figure 10It is a block diagram of an object detection device shown according to an exemplary embodiment;

[0039] Figure 11 It is a block diagram of a vehicle applied to an object detection method shown according to an exemplary embodiment. Detailed implementation manners

[0040] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0041] It should be noted that all actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.

[0042] During the driving of the vehicle, uneven road surfaces such as speed bumps and potholed roads are identified through relevant algorithms, so that the vehicle can make adaptive adjustments in advance, enabling the vehicle to drive over the uneven road surface more smoothly, thereby obtaining a better driving experience. For some uneven road surfaces, the characteristics of their center points are not obvious, and at the same time, due to the camera distortion of the camera for photographing the road surface, Figure 1 is an image captured by the camera under distortion conditions. The above reasons can easily lead to low detection accuracy.

[0043] For speed bumps on uneven roads, a speed bump detection algorithm is usually used to locate the position of the speed bump in front of the road, providing information about the speed bump on the front road surface for the "magic carpet" function in related technologies. The "magic carpet" function uses sensing systems such as cameras and lidar to detect the road surface ahead. When an uneven road surface is detected, the suspension actuator makes an adaptive adjustment in advance, enabling the vehicle to drive more smoothly over the uneven road surface and obtaining the best comfort experience. Among them, the speed bump detection algorithm provides important support for the "magic carpet" function and has important business value. Two related technologies are provided below. One is to use the full-image segmentation method to segment the speed bump contour. Specifically, the model architecture of the speed bump detection method based on full-image segmentation can usually be divided into two parts: an encoder and a decoder. During training, the images containing speed bumps collected by the real vehicle and the manually annotated segmentation masks are used as inputs. The encoder is responsible for reducing the size of the feature map and capturing higher-level semantic information, while the decoder is used to restore the spatial information. During inference, only the image taken by the real vehicle camera needs to be input, and the speed bump contour can be segmented. The latest segmentation model algorithm (such as the deeplabv3plus algorithm) uses deeplabv3 as the encoder module and uses a simple and effective decoder module. This model can control the resolution of the encoder features extracted through atrous convolution, thereby balancing accuracy and running time.

[0044] The disadvantage of this related technology is that the segmentation scheme model has many parameters, a large amount of computation, high deployment resource consumption, and high time consumption. It cannot meet the requirements of high frame rate and low latency for real vehicle deployment.

[0045] The other is to use the object detection algorithm to predict the center point and the two bottom end points of the speed bump. Specifically:

[0046] The speed bump detection scheme based on object detection adopts the idea of a classic object detection algorithm (such as CenterNet). The center point of the speed bump, as well as the x and y direction offsets from the center point to the two bottom end points of the lower left and lower right, are predicted respectively, and finally the positions of the center point and the two bottom end points of the speed bump are obtained. During training, the images containing speed bumps collected by the real vehicle and the coordinates of the center point of the speed bump, the lower left bottom end point, and the lower right bottom end point manually annotated are used to calculate the x and y direction offsets from the center point to the two bottom end points of the lower left and lower right as the labels for network training. During inference, the image taken by the real vehicle camera is input, and the model finds the local maximum confidence response on the feature map output at the end of the network, retains the responses with a confidence higher than a certain threshold, each response corresponds to a speed bump, and the coordinates of the center point and the two end points of the bottom edge of each speed bump are calculated as the detection results of the speed bump.

[0047] The disadvantages of this related technology are as follows: due to the long and slender characteristics of the speed bump, the characteristics of the center point itself are not obvious. At the same time, due to the existence of camera distortion, the overly long speed bump does not satisfy the straight-line frame constraint in the image. For the above reasons, the detection rates of overly long speed bumps and distorted speed bumps are relatively low. At the same time, limited by the receptive field of the network, this technology may not be able to accurately predict the unique center point. For example, multiple center points may be predicted for the same speed bump, which is extremely likely to cause the problem of a speed bump being detected multiple times, as Figure 2 shown. Or a straight-line frame does not enclose the entire speed bump, resulting in the problem of missed detection of the speed bump or being considered as multiple speed bumps, as Figure 3 shown.

[0048] The present disclosure provides an object detection method. Please refer to Figure 1 wherein the object detection method can be applied to the object detection device 500 shown in Figure 10 and the vehicle 600 shown in Figure 11 as well as a computer-readable storage medium. In this embodiment, taking the application to a vehicle as an example, the vehicle can be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The following will elaborate in detail on the process shown in Figure 1 The object detection method specifically may include the following steps:

[0049] Step S110: Obtain a driving image collected during the driving of the vehicle, and process the driving image to obtain a to-be-detected image, where the to-be-detected image is an image containing a to-be-detected target.

[0050] During the driving of the vehicle, a driving image during the driving process is collected through a camera installed on the vehicle. The driving image contains a to-be-detected target, and the to-be-detected target can be in a long and slender shape. The to-be-detected target can be a speed bump, a lamp post of a street lamp, a lamp post of a traffic light, etc. The collected driving image is processed to obtain a to-be-detected image. Processing the driving image can include noise reduction processing, cropping processing, downsampling processing, etc. Among them, after the driving image undergoes cropping processing and downsampling processing, the size of the obtained to-be-detected image is smaller than the size of the original driving image.

[0051] Step S120: Input the to-be-detected image into a prediction model to obtain a prediction image for predicting the edge points of the to-be-detected target in the to-be-detected image.

[0052] The prediction model is a pre-trained model used to predict the edge points of the to-be-detected target in the to-be-detected image containing the to-be-detected target to obtain a prediction image, where the size of the to-be-detected image is the same as the size of the prediction image.

[0053] In one implementation, the prediction model can be stored locally in the vehicle. When prediction is required, the prediction model can be quickly obtained from the vehicle locally, and the prediction model is used to identify the image to be detected to obtain a prediction image. In this implementation, since the prediction model is stored locally in the vehicle, the vehicle can quickly obtain the prediction model for prediction, which improves the prediction rate.

[0054] In another implementation, the prediction model can be stored in a server communicatively connected to the vehicle. When prediction is required, the vehicle sends a request to the server connected to it. In response to this request, the server sends the prediction model to the vehicle, and the vehicle uses the prediction model to identify the image to be detected to obtain a prediction image. In this implementation, since the prediction model is stored in the server, the occupancy of the local memory of the vehicle by the prediction model can be reduced.

[0055] Step S130: According to the prediction image, obtain the first position of the edge point in the first coordinate system where the image to be detected is located.

[0056] Optionally, the position information of the edge points of the object to be detected is included in the prediction image. According to the position information on the prediction image, the first position of the edge point in the first coordinate system where the image to be detected is located is thus obtained.

[0057] Since the size of the prediction image is the same as the size of the image to be detected, the same position on the prediction image and the image to be detected corresponds to the same pixel point. It can be understood that if the pixel point on the prediction image is an edge point, then the pixel point at the same position as the edge point of the prediction image on the image to be detected also belongs to the edge point on the object to be detected. The position information of the edge points of the object to be detected is included in the prediction image. Since the pixel point at the same position as the edge point on the prediction image on the image to be detected is the edge point, the position information of the edge point on the prediction image is the first position of the edge point in the first coordinate system where the image to be detected is located.

[0058] Step S140: According to the first position where the edge point in the prediction image is located, obtain the second position of the object to be detected in the second coordinate system where the driving image is located.

[0059] In one implementation, according to the first position where the edge point in the prediction image is located and the processing method of the driving image in step S110, the second position of the object to be detected in the second coordinate system where the driving image is located is obtained.

[0060] In the object detection method provided in this embodiment, first, a driving image collected during the driving of the vehicle is obtained, and the driving image is processed to obtain a to-be-detected image including the to-be-detected object. Then, the to-be-detected image is input into a prediction model to obtain a prediction image for predicting the edge points of the to-be-detected object in the to-be-detected image. Next, according to the prediction image, the first position of the edge points in the first coordinate system where the to-be-detected image is located is obtained, and according to the first position of the edge points in the prediction image, the second position of the to-be-detected object in the second coordinate system where the driving image is located is obtained. In this embodiment, by predicting the edge points of the to-be-detected object, the prediction of the to-be-detected object pays attention to local features, effectively avoiding the problem of low accuracy in detecting the to-be-detected object caused by factors such as the unclear center point feature of the to-be-detected object and the distortion of the speed bump captured due to camera distortion, thereby improving the accuracy of detecting the to-be-detected object.

[0061] Optionally, before using the prediction model for prediction, the prediction model is pre-trained. The prediction model is obtained through the following steps:

[0062] Step S210: Obtain a to-be-recognized sample image, where the to-be-recognized sample image is an image including a sample object.

[0063] In one implementation, images including sample objects taken by various vehicles during driving are stored in an image library. For example, the sample object can be slender in shape. For example, the sample object can be a speed bump, a lamp post of a street lamp, a lamp post of a traffic light, etc. Images are obtained from the image library, and the obtained images are directly used as the to-be-recognized sample images.

[0064] In another implementation, a driving sample image collected during the driving of the vehicle is obtained, where the driving sample image can be an image including a sample object obtained by the vehicle from the image library. Then, the driving sample image is processed to obtain the to-be-recognized sample image. The processing of the driving sample image can include noise reduction processing, cropping processing, downsampling processing, etc.

[0065] Step S220: Construct a sample prediction image according to the to-be-recognized sample image, where the sample prediction image includes the position information of the pixel points at the edge of the sample object.

[0066] The number of sample prediction images may be a preset number. In one implementation, the preset number is 3. The 3 sample prediction images include a confidence map label, a first direction offset map label, and a second direction offset map label for constructing the sample image to be recognized. For example, the 3 sample prediction images can be constructed in the following manner: constructing a confidence map label, a first direction offset map label, and a second direction offset map label according to the sample image to be recognized, wherein the size of the confidence map label, the size of the first direction offset map label, and the size of the second direction offset map label are all the same as the size of the sample image to be recognized. The confidence map label includes the confidence corresponding to each pixel point among multiple pixel points on the sample image to be recognized and the coordinate position of each pixel point in the first coordinate system where the sample image to be recognized is located. The confidence corresponding to each pixel point is used to represent the confidence of the pixel point belonging to the pixel points on the edge of the sample target. The first direction offset map label includes the coordinate position of each pixel point on the first coordinate axis, and the second direction offset map label includes the coordinate position of each pixel point on the second coordinate axis.

[0067] Optionally, when the size of the sample image to be recognized is smaller than the size of the driving sample image, for example, after downsampling or cropping the driving sample image, the size of the obtained sample image to be recognized is smaller than the original driving image. The confidence map label is constructed in the following manner: constructing an initial confidence map label, wherein the pixel values of all pixel points on the initial confidence map label are 0; annotating the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; dividing the driving sample image into a preset number of pixel regions, wherein each pixel region in the preset number of pixel regions corresponds to each pixel point on the initial confidence map label. For example, the size of the sample image to be recognized is 240*96, and the size of the driving sample image is 2880*930. Since 2880 / 240 is 12 and 930 / 96 rounded up is 10, then each pixel point on the sample image to be recognized corresponds to a pixel region on 12*10 pixel regions. It can be understood that since the size of the sample image to be recognized is smaller than the size of the driving sample image, a pixel point on the sample image to be recognized corresponds to a pixel region when mapped to the driving sample image; taking the pixel region where any contour point on the contour line is located as the first target pixel region; setting the pixel value of the pixel on the initial confidence map label corresponding to the first target pixel region to 1 to obtain the confidence map label.

[0068] Optionally, when the size of the sample image to be recognized is smaller than the size of the driving sample image, the first direction offset map label is constructed in the following manner: construct a first initial offset map label, wherein the pixel values of all pixel points on the first initial offset map label are 0; label the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; divide the driving sample image into a preset number of pixel regions, wherein each pixel region in the preset number of pixel regions corresponds to each pixel point in the first initial offset map label; take the pixel region where any contour point on the contour line is located as a second target pixel region; set the pixel value of the pixel corresponding to the second target pixel region on the first initial offset map label to a first target value to obtain the first direction offset map label.

[0069] Wherein, the first target value is obtained in the following manner: obtain a first distance from each contour point in the second target region to the center point of the second target region to obtain a plurality of first distances; calculate the average value of the plurality of first distances to obtain a first average distance; take the first average distance as the first target value. Optionally, the first average distance can also be normalized to the interval [-0.5, 0.5] to obtain the first target value.

[0070] Optionally, when the size of the sample image to be recognized is smaller than the size of the driving sample image, the second direction offset map label is constructed in the following manner: construct a second initial offset map label, wherein the pixel values of all pixel points on the second initial offset map label are 0; label the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; divide the driving sample image into a preset number of pixel regions, wherein each pixel region in the preset number of pixel regions corresponds to each pixel point in the second initial offset map label; take the pixel region where any contour point on the contour line is located as a third target pixel region; set the pixel value of the pixel point corresponding to the third target pixel region on the second initial offset map label to a second target value to obtain the second direction offset map label.

[0071] Wherein, the second target value is obtained in the following manner: obtain a second distance from each contour point in the third target region to the center point of the third target region to obtain a plurality of second distances; calculate the average value of the plurality of second distances to obtain a second average distance; take the second average distance as the second target value.

[0072] It should be noted that the number of sample prediction images is not limited to the above 3. The number of sample prediction images can also be 2. For example, the sample prediction images include the confidence map label and the first direction offset map label constructed from the sample image to be recognized, or the sample prediction images include the confidence map label and the second direction offset map label constructed from the sample image to be recognized. The number of sample prediction images can also be 1. For example, the sample prediction image includes the confidence map label constructed from the sample image to be recognized.

[0073] Optionally, the second average distance can also be normalized to the interval [-0.5, 0.5] to obtain the second target value.

[0074] Optionally, through the above construction of samples, a training set composed of 50,000 groups of training data can be obtained for training the model.

[0075] Step S230: Train the initial model according to the sample image to be recognized to obtain the prediction model.

[0076] Iteratively train the initial model with the sample image to be recognized to obtain the prediction model. Among them, the initial model can be the FASTResNet18 network.

[0077] In one implementation, step S230 includes: inputting the sample image to be recognized into the initial model to obtain an intermediate confidence map, a first intermediate offset map label, and a second intermediate offset map label; obtaining a first loss function according to the confidence map label and the intermediate confidence map. For example, the first loss function can be a cross-entropy loss function; and obtaining a second loss function according to the first intermediate offset map label, the second intermediate offset map label, the first direction offset map label, and the second direction offset map label. For example, the second loss function can be an L1 loss function; training the initial model according to the first loss function and the second loss function to obtain the prediction model. Among them, the first loss function is used to supervise the confidence map label, and the second loss function is used to supervise the first intermediate offset map label and the second intermediate offset map label until the loss function converges, and finally the prediction model is obtained.

[0078] In this implementation, training the prediction model facilitates subsequent prediction of the edge points of the target to be detected, so as to improve the accuracy of the final detection of the detection target.

[0079] In one implementation, step S110 can be in the following manner: perform cropping processing on the driving image to obtain the image to be detected. For example, taking the deceleration strip as the target to be detected, the driving image is as Figure 6As shown, the size of the driving image is 2880*1860, where 2880 is the length of the driving image, and 1860 is the width or height of the driving image. The shearing process can be to shear the driving image in the height direction. For example, keep 1 / 2 of the driving image in the height direction, and use the image of the remaining ground part as the image to be detected, such as Figure 7 As shown, the size of the image to be detected is 2880*930.

[0080] In this embodiment, by performing a cropping process on the driving image and using a smaller-sized image to be detected for the prediction of the prediction model, the computational amount of the prediction model can be reduced, and the prediction image can be obtained quickly.

[0081] In another embodiment, step S110 can be as follows: perform a downsampling process on the driving image to obtain the image to be detected. For example, perform 4-fold downsampling on the driving image.

[0082] In another embodiment, step S110 can be as follows: perform a cropping process on the driving image to obtain a cropped image; for example, perform a cropping process on a driving image with a size of 2880*1860 to obtain a cropped image with a size of 2880*930. Then perform a downsampling process on the cropped image to obtain the image to be detected. For example, perform 4-fold downsampling on the cropped image to obtain an image to be detected with a size of 240*960.

[0083] In one embodiment, the number of the prediction images is a preset number, and the position information of the edge points is included in the preset number of prediction images. Step S130 can be as follows: obtain the first position of the edge point in the first coordinate system where the image to be detected is located according to the position information of the edge point in the preset number of prediction images.

[0084] As a way, the preset quantity can be 3, and the 3 predicted images include: a confidence map, a first direction offset map, and a second direction offset map; wherein, the size of the confidence map, the size of the first direction offset map, and the size of the second direction offset map are all the same as the size of the image to be detected. The confidence map includes the confidence corresponding to each pixel point among the multiple pixel points on the image to be detected and the first coordinate position of each pixel point in the first coordinate system where the image to be detected is located. The confidence corresponding to each pixel point is used to represent the confidence of the pixel point belonging to the edge of the target to be detected. The first coordinate position includes the first coordinate on the first coordinate axis in the first coordinate system and the second coordinate on the second coordinate axis in the first coordinate system. The first direction offset map includes the third coordinate of each pixel point on the first coordinate axis, and the second direction offset map includes the fourth coordinate of each pixel point on the second coordinate axis. Optionally, the coordinates on the confidence map are integers, and the coordinates on the two maps of the first direction offset map and the second direction offset map are numbers less than 1.

[0085] Obtaining the first position of the edge point in the first coordinate system where the image to be detected is located according to the predicted image includes: taking the pixel points on the confidence map with a confidence greater than the preset confidence as edge points, where the preset confidence can be 0.65; obtaining the first coordinate position of the edge point in the first coordinate system from the confidence map. For example, if the first coordinate system is an x-y coordinate system and the first coordinate position is (x1, y1), obtaining the third coordinate (x3, 0) of the edge point on the first coordinate axis from the first direction offset map, and obtaining the fourth coordinate (0, y4) of the edge point on the second coordinate axis from the second direction offset map. Wherein, the first coordinate position includes the first coordinate (x1, 0) of the edge point on the first coordinate axis and the second coordinate (0, y1) of the edge point on the second coordinate axis; obtaining the fifth coordinate (x1 + x3, 0) on the first coordinate axis according to the first coordinate (x1, 0) and the third coordinate (x3, 0) of the edge point, and obtaining the sixth coordinate (0, y1 + y4) on the second coordinate axis according to the second coordinate (0, y1) and the fourth coordinate (0, y4) of the edge point; determining the first position (x1 + x3, y1 + y4) of the edge point in the first coordinate system according to the fifth coordinate (x1 + x3, 0) and the sixth coordinate (0, y1 + y4) of the edge point in the first coordinate system. For example, x1 is 2, y1 is 7, x3 is 0.2, and y4 is 0.5.

[0086] Optionally, the number of selected edge points can be controlled within a preset number. For example, the preset number is 100. For example, if the number of pixel points on the confidence map with a confidence greater than the preset confidence is greater than 100, then the top 100 pixel points with the highest confidence are selected from these pixel points as edge points.

[0087] As another way, the preset number can be 2. The two predicted images include: a confidence map and a first direction offset map, or the two predicted images include: a confidence map and a second direction offset map.

[0088] In one implementation, the number of the edge points is multiple. According to the first positions of the edge points in the predicted image, step S140 can be as follows. Through a first coordinate conversion relationship, the first position of each edge point in the first coordinate system where the predicted image is located among the multiple edge points is converted into a third position in the second coordinate system where the driving image is located, where the first coordinate conversion relationship is the conversion relationship between the first coordinate system and the second coordinate system; according to the pixel points corresponding to the third positions on the driving image, the contour of the target to be detected is formed, and the fourth position where the center point of the contour is located is obtained; according to the fourth position where the center point of the contour is located, the second position of the target to be detected in the second coordinate system where the driving image is located is obtained.

[0089] Exemplarily, the target detection method further includes: obtaining a bounding box surrounding the contour through a preset function, where the bounding box is a rectangle. For example, the preset function can be the boundingRect function of opencv; obtaining the side lengths of the bounding box according to the fifth positions of two opposite endpoints on the bounding box in the second coordinate system; the obtaining the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located includes: obtaining the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located and the side lengths of the bounding box. For example, the second position can be represented in the form of (x, y, w, h), where w and h are the side lengths of two adjacent sides of the bounding box, and (x, y) is the fourth position where the center point is located.

[0090] Optionally, when the target to be detected is a speed bump, it can be inferred whether the speed bump is a horizontal speed bump or a vertical speed bump according to the values of w and h.

[0091] Optionally, the target detection method further includes: converting the second position of the target to be detected in the second coordinate system where the driving image is located through a second coordinate conversion relationship into a sixth position in the third coordinate system corresponding to the vehicle, where the second coordinate conversion relationship is the conversion relationship between the second coordinate system and the third coordinate system.

[0092] In this embodiment, converting the target to be detected into the third coordinate system corresponding to the vehicle facilitates subsequent vehicle control such as braking and decelerating.

[0093] The present disclosure provides a target detection method. Refer to Figure 8 , and this method includes:

[0094] Step S310: Input training data.

[0095] Step S320: Region clipping.

[0096] Step S330: Make labels.

[0097] Step S340: Train the model.

[0098] Among them, for the detailed descriptions of steps S310 to S340, please refer to the above steps and will not be elaborated here.

[0099] Taking the target to be detected as a speed bump as an example, this embodiment provides a target detection method. Refer to Figure 9 , and this method includes:

[0100] Step S410: Input the pictures collected by the camera.

[0101] Step S420: Region clipping.

[0102] Step S430: Call the prediction model and output a prediction image containing the positions of the edge points.

[0103] Step S440: Obtain the minimum circumscribed polygon contour.

[0104] Step S450: Convert it into the form of a standard detection box.

[0105] Step S460: Output the detection result of the speed bump.

[0106] Among them, for the detailed descriptions of steps S410 to S460, please refer to the above steps and will not be elaborated here.

[0107] In practical applications, the recall rate of the target detection method provided in this embodiment is between 0.68 and 0.74, greatly improving the detection effect of the speed bump.

[0108] To implement the above method embodiments, the present disclosure provides an object detection device. Please refer to Figure 10 , the object detection device 500 includes:

[0109] A processing module 510, configured to acquire a driving image collected during the driving of a vehicle, and process the driving image to obtain a to-be-detected image, where the to-be-detected image is an image including a to-be-detected target;

[0110] A prediction module 520, configured to input the to-be-detected image into a prediction model to obtain a prediction image for predicting edge points of the to-be-detected target in the to-be-detected image;

[0111] A first position acquisition module 530, configured to acquire a first position of the edge point in a first coordinate system where the to-be-detected image is located according to the prediction image;

[0112] A second position acquisition module 540, configured to obtain a second position of the to-be-detected target in a second coordinate system where the driving image is located according to the first position where the edge point in the prediction image is located.

[0113] Optionally, the number of the prediction images is a preset number, and the preset number of prediction images includes position information of edge points. The first position acquisition module 530 includes:

[0114] A first position acquisition sub-module, configured to obtain the first position of the edge point in the first coordinate system where the to-be-detected image is located according to the position information of the edge point in the preset number of prediction images.

[0115] Optionally, the preset number of prediction images includes: a confidence map, a first direction offset map, and a second direction offset map; wherein, the size of the confidence map, the size of the first direction offset map, and the size of the second direction offset map are all the same as the size of the to-be-detected image. The confidence map includes the confidence corresponding to each pixel point among multiple pixel points on the to-be-detected image and the first coordinate position of each pixel point in the first coordinate system where the to-be-detected image is located. The confidence corresponding to each pixel point is used to characterize the confidence of the pixel point belonging to the edge pixel points of the to-be-detected target. The first coordinate position includes a first coordinate on a first coordinate axis in the first coordinate system and a second coordinate on a second coordinate axis in the first coordinate system. The first direction offset map includes a third coordinate of each pixel point on the first coordinate axis, and the second direction offset map includes a fourth coordinate of each pixel point on the second coordinate axis. The first position acquisition sub-module includes:

[0116] An edge point determination module, configured to use the pixel points on the confidence map with a confidence greater than a preset confidence as edge points;

[0117] A third position acquisition module, configured to acquire the first coordinate position of the edge point in the first coordinate system from the confidence map, acquire the third coordinate of the edge point on the first coordinate axis from the first direction offset map, and acquire the fourth coordinate of the edge point on the second coordinate axis from the second direction offset map, where the first coordinate position includes the first coordinate of the edge point on the first coordinate axis and the second coordinate of the edge point on the second coordinate axis;

[0118] A fourth position acquisition module, configured to obtain the fifth coordinate on the first coordinate axis according to the first coordinate and the third coordinate of the edge point, and obtain the sixth coordinate on the second coordinate axis according to the second coordinate and the fourth coordinate of the edge point;

[0119] A first position determination module, configured to determine the first position of the edge point in the first coordinate system according to the fifth coordinate and the sixth coordinate of the edge point in the first coordinate system.

[0120] Optionally, the number of the edge points is multiple, and the second position acquisition module 540 includes:

[0121] A first conversion module, configured to convert the first position of each edge point among the multiple edge points in the first coordinate system where the prediction image is located to the third position in the second coordinate system where the driving image is located through a first coordinate conversion relationship, where the first coordinate conversion relationship is the conversion relationship between the first coordinate system and the second coordinate system;

[0122] A contour acquisition module, configured to form the contour of the target to be detected according to the pixel points corresponding to the third position on the driving image, and acquire the fourth position where the center point of the contour is located;

[0123] A second position acquisition sub-module, configured to obtain the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located.

[0124] The target detection device 500 further includes:

[0125] An enclosing box acquisition module, configured to obtain an enclosing box enclosing the contour through a preset function, where the enclosing box is rectangular;

[0126] A side length acquisition module, configured to obtain the side length of the bounding box according to the fifth positions of two opposite endpoints on the bounding box in the second coordinate system;

[0127] A second position acquisition sub-module, specifically configured to obtain the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located and the side length of the bounding box.

[0128] Optionally, the target detection device 500 further includes:

[0129] A sample acquisition module, configured to acquire a sample image to be recognized, where the sample image to be recognized is an image including a sample target;

[0130] A sample construction module, configured to construct a sample prediction image according to the sample image to be recognized, where the sample prediction image includes position information of pixel points on the edge of the sample target;

[0131] A training module, configured to train an initial model according to the sample image to be recognized to obtain the prediction model.

[0132] Optionally, the sample construction module includes:

[0133] A sample construction sub-module, configured to construct a confidence map label, a first direction offset map label, and a second direction offset map label according to the sample image to be recognized, where the sizes of the confidence map label, the first direction offset map label, and the second direction offset map label are the same as the size of the sample image to be recognized, the confidence map label includes the confidence corresponding to each pixel point among multiple pixel points on the sample image to be recognized and the coordinate position of each pixel point in the first coordinate system where the sample image to be recognized is located, the confidence corresponding to each pixel point is used to represent the confidence of the pixel point belonging to the pixel points on the edge of the sample target, the first direction offset map label includes the coordinate position of each pixel point on the first coordinate axis, and the second direction offset map label includes the coordinate position of each pixel point on the second coordinate axis.

[0134] Optionally, the training module includes:

[0135] An input module, configured to input the sample image to be recognized into the initial model to obtain an intermediate confidence map, a first intermediate offset map label, and a second intermediate offset map label;

[0136] A loss function acquisition module, configured to obtain a first loss function according to the confidence map label and the intermediate confidence map, and obtain a second loss function according to the first intermediate offset map label, the second intermediate offset map label, the first direction offset map label, and the second direction offset map label;

[0137] A model training module, configured to train the initial model according to the first loss function and the second loss function to obtain the prediction model.

[0138] Optionally, the sample acquisition module includes:

[0139] A driving sample image acquisition module, configured to acquire driving sample images collected during the driving of the vehicle;

[0140] A sample processing module, configured to process the driving sample images to obtain the sample images to be recognized.

[0141] Optionally, the size of the sample image to be recognized is smaller than the size of the driving sample image, and the object detection device 500 further includes:

[0142] An initial confidence map label construction module, configured to construct an initial confidence map label, wherein the pixel values of all pixel points on the initial confidence map label are 0;

[0143] A first annotation module, configured to annotate the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points;

[0144] A first division module, configured to divide the driving sample image into a preset number of pixel regions, wherein each pixel region in the preset number of pixel regions corresponds to each pixel point in the initial confidence map label;

[0145] A first target pixel region determination module, configured to use the pixel region where any contour point on the contour line is located as the first target pixel region;

[0146] A confidence map label acquisition module, configured to set the pixel value of the pixel on the initial confidence map label corresponding to the first target pixel region to 1 to obtain the confidence map label.

[0147] Optionally, the size of the sample image to be recognized is smaller than the size of the driving sample image, and the object detection device 500 further includes:

[0148] A first initial offset map label construction module, configured to construct a first initial offset map label, wherein the pixel values of all pixel points on the first initial offset map label are 0;

[0149] The second annotation module is used to annotate the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points;

[0150] The second division module is used to divide the driving sample image into a preset number of pixel regions, where each pixel region in the preset number of pixel regions corresponds to each pixel point in the first initial offset map label;

[0151] The second target pixel region determination module is used to use the pixel region where any contour point on the contour line is located as the second target pixel region;

[0152] The first direction offset map label acquisition module is used to set the pixel value of the pixel on the first initial offset map label corresponding to the second target pixel region to the first target value to obtain the first direction offset map label.

[0153] Optionally, the target detection device 500 further includes:

[0154] The first distance acquisition module is used to acquire the first distance of each contour point in the second target region from the center point of the second target region to obtain a plurality of first distances;

[0155] The first average distance acquisition module is used to calculate the average value of the plurality of first distances to obtain the first average distance;

[0156] The first target value acquisition module is used to use the first average distance as the first target value.

[0157] Optionally, the size of the sample image to be recognized is smaller than the size of the driving sample image, and the target detection device 500 further includes:

[0158] The second initial offset map label construction module is used to construct a second initial offset map label, where the pixel values of all pixel points on the second initial offset map label are 0;

[0159] The third annotation module is used to annotate the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points;

[0160] The preset number of pixel region division module is used to divide the driving sample image into a preset number of pixel regions, where each pixel region in the preset number of pixel regions corresponds to each pixel point in the second initial offset map label;

[0161] The third target pixel region acquisition module is used to use the pixel region where any contour point on the contour line is located as the third target pixel region;

[0162] The second target value obtaining module is configured to set the pixel value of the pixel point corresponding to the third target pixel region on the second initial offset map label as the second target value, so as to obtain the second direction offset map label.

[0163] Optionally, the target detection device 500 further includes:

[0164] The second distance obtaining module is configured to obtain the second distance between each contour point in the third target region and the center point of the third target region, so as to obtain a plurality of second distances;

[0165] The second average distance obtaining module is configured to calculate the average value of the plurality of second distances, so as to obtain the second average distance;

[0166] The second target value obtaining module is configured to use the second average distance as the second target value.

[0167] Optionally, the processing module 510 includes:

[0168] The first clipping module is configured to perform clipping processing on the driving image to obtain the image to be detected.

[0169] Optionally, the processing module 510 includes:

[0170] The first downsampling module is configured to perform downsampling processing on the driving image to obtain the image to be detected.

[0171] Optionally, the processing module 510 includes:

[0172] The second clipping module is configured to perform clipping processing on the driving image to obtain a clipped image;

[0173] The second downsampling module is configured to perform downsampling processing on the clipped image to obtain the image to be detected.

[0174] Optionally, the target detection device 500 further includes:

[0175] The second conversion module is configured to perform coordinate conversion on the second position of the target to be detected in the second coordinate system where the driving image is located through a second coordinate conversion relationship, and convert it into the sixth position in the third coordinate system corresponding to the vehicle, where the second coordinate conversion relationship is the conversion relationship between the second coordinate system and the third coordinate system.

[0176] Regarding the target detection device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0177] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of the object detection method provided by the present disclosure are implemented.

[0178] Figure 11 FIG. is a block diagram of a vehicle applied to an object detection method shown according to an exemplary embodiment. For example, vehicle 600 may be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. Vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.

[0179] Referring to Figure 11 , vehicle 600 may include various subsystems. For example, the infotainment system 610, the perception system 620, the decision control system 630, the drive system 640, and the computing platform 650. Among them, vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of vehicle 600 may be interconnected by wired or wireless means.

[0180] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, and a navigation system, etc.

[0181] The perception system 620 may include several sensors for sensing information about the environment around vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), a lidar, a millimeter wave radar, an ultrasonic radar, and a camera device.

[0182] The decision control system 630 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.

[0183] The drive system 640 may include components that provide motive power for vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a powertrain, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.

[0184] Some or all functions of vehicle 600 are controlled by the computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652, and the processor 651 may execute instructions 653 stored in the memory 652.

[0185] The processor 651 can be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0186] The memory 652 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0187] In addition to the instructions 653, the memory 652 can also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 652 can be used by the computing platform 650.

[0188] In the embodiments of the present disclosure, the processor 651 can execute the instructions 653 to complete all or part of the steps of the above-mentioned object detection method.

[0189] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-mentioned object detection method when executed by the programmable device.

[0190] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0191] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A target detection method, characterized in that, The method includes: Obtain a driving image collected during the driving of a vehicle, and process the driving image to obtain a to-be-detected image, where the to-be-detected image is an image including a to-be-detected target; Input the to-be-detected image into a prediction model to obtain a prediction image for predicting edge points of the to-be-detected target in the to-be-detected image; According to the prediction image, obtain a first position of the edge points in a first coordinate system where the to-be-detected image is located; According to the first position where the edge points in the prediction image are located, obtain a second position of the to-be-detected target in a second coordinate system where the driving image is located.

2. The method according to claim 1, wherein The number of the prediction images is a preset number, and the preset number of prediction images includes position information of the edge points. The step of obtaining the first position of the edge points in the first coordinate system where the to-be-detected image is located according to the prediction image includes: Obtain the first position of the edge points in the first coordinate system where the to-be-detected image is located according to the position information of the edge points in the preset number of prediction images.

3. The method according to claim 2, wherein The preset number of prediction images includes: a confidence map, a first direction offset map, and a second direction offset map; wherein, the size of the confidence map, the size of the first direction offset map, and the size of the second direction offset map are all the same as the size of the to-be-detected image. The confidence map includes the confidence corresponding to each pixel point among a plurality of pixel points on the to-be-detected image and the first coordinate position of each pixel point in the first coordinate system where the to-be-detected image is located. The confidence corresponding to each pixel point is used to represent the confidence of the pixel point belonging to the edge pixel points of the to-be-detected target. The first coordinate position includes the first coordinate on the first coordinate axis in the first coordinate system and the second coordinate on the second coordinate axis in the first coordinate system. The first direction offset map includes the third coordinate of each pixel point on the first coordinate axis, and the second direction offset map includes the fourth coordinate of each pixel point on the second coordinate axis; The step of obtaining the first position of the edge points in the first coordinate system where the to-be-detected image is located according to the position information of the edge points in the preset number of prediction images includes: Take the pixel points on the confidence map with a confidence greater than a preset confidence as edge points; Obtain the first coordinate position of the edge points in the first coordinate system from the confidence map, obtain the third coordinate of the edge points on the first coordinate axis from the first direction offset map, and obtain the fourth coordinate of the edge points on the second coordinate axis from the second direction offset map, where the first coordinate position includes the first coordinate of the edge points on the first coordinate axis and the second coordinate of the edge points on the second coordinate axis; Obtain a fifth coordinate on the first coordinate axis according to the first coordinate and the third coordinate of the edge points, and obtain a sixth coordinate on the second coordinate axis according to the second coordinate and the fourth coordinate of the edge points; Determine the first position of the edge point in the first coordinate system according to the fifth coordinate and the sixth coordinate of the edge point in the first coordinate system.

4. The method according to claim 1, characterized in that The number of the edge points is multiple. Obtaining the second position of the target to be detected in the second coordinate system where the driving image is located according to the first position of the edge points in the prediction image includes: Convert the first position of each edge point among the multiple edge points in the first coordinate system where the prediction image is located into a third position in the second coordinate system where the driving image is located through a first coordinate conversion relationship, where the first coordinate conversion relationship is the conversion relationship between the first coordinate system and the second coordinate system; According to the pixel points corresponding to the third position on the driving image, form the contour of the target to be detected, and obtain the fourth position where the center point of the contour is located; Obtain the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located.

5. The method according to claim 4, characterized in that, It further includes: Obtain a bounding box enclosing the contour through a preset function, where the bounding box is rectangular; Obtain the side length of the bounding box according to the fifth positions of two opposite end points on the bounding box in the second coordinate system; The obtaining the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located includes: Obtain the second position of the target to be detected in the second coordinate system where the driving image is located according to the fourth position where the center point of the contour is located and the side length of the bounding box.

6. The method according to any one of claims 1 to 5, characterized in that, The prediction model is obtained through the following method: Obtain a sample image to be recognized, where the sample image to be recognized is an image containing a sample target; Construct a sample prediction image according to the sample image to be recognized, where the sample prediction image includes the position information of the pixel points on the edge of the sample target; Train an initial model according to the sample image to be recognized to obtain the prediction model.

7. The method according to claim 6, wherein The constructing a sample prediction image according to the sample image to be recognized includes: Construct a confidence map label, a first direction offset map label, and a second direction offset map label according to the sample image to be recognized, where the sizes of the confidence map label, the first direction offset map label, and the second direction offset map label are all the same as the size of the sample image to be recognized. The confidence map label includes the confidence corresponding to each pixel point among the multiple pixel points on the sample image to be recognized and the coordinate position of each pixel point in the first coordinate system where the sample image to be recognized is located. The confidence corresponding to each pixel point is used to represent the confidence that the pixel point belongs to the pixel point on the edge of the sample target. The first direction offset map label includes the coordinate position of each pixel point on the first coordinate axis, and the second direction offset map label includes the coordinate position of each pixel point on the second coordinate axis.

8. The method according to claim 7, wherein Training the initial model according to the to-be-recognized sample image to obtain the prediction model includes: Inputting the to-be-recognized sample image into the initial model to obtain an intermediate confidence map, a first intermediate offset map label, and a second intermediate offset map label; Obtaining a first loss function according to the confidence map label and the intermediate confidence map, and obtaining a second loss function according to the first intermediate offset map label, the second intermediate offset map label, the first direction offset map label, and the second direction offset map label; Training the initial model according to the first loss function and the second loss function to obtain the prediction model.

9. The method according to claim 7, wherein Obtaining the to-be-recognized sample image includes: Obtaining a driving sample image collected during the vehicle's driving; Processing the driving sample image to obtain the to-be-recognized sample image.

10. The method according to claim 9, wherein The size of the to-be-recognized sample image is smaller than that of the driving sample image, and the confidence map label is constructed in the following manner: Constructing an initial confidence map label, where the pixel values of all pixel points on the initial confidence map label are 0; Labeling the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; Dividing the driving sample image into a preset number of pixel regions, where each pixel region in the preset number of pixel regions corresponds to each pixel point on the initial confidence map label; Taking the pixel region where any contour point on the contour line is located as the first target pixel region; Setting the pixel value of the pixel on the initial confidence map label corresponding to the first target pixel region to 1 to obtain the confidence map label.

11. The method according to claim 9, wherein The size of the to-be-recognized sample image is smaller than that of the driving sample image, and the first direction offset map label is constructed in the following manner: Constructing a first initial offset map label, where the pixel values of all pixel points on the first initial offset map label are 0; Labeling the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; Dividing the driving sample image into a preset number of pixel regions, where each pixel region in the preset number of pixel regions corresponds to each pixel point on the first initial offset map label; Taking the pixel region where any contour point on the contour line is located as the second target pixel region; Setting the pixel value of the pixel on the first initial offset map label corresponding to the second target pixel region to a first target value to obtain the first direction offset map label.

12. The method according to claim 11, wherein The first target value is obtained in the following manner: Obtaining a first distance from each contour point in the second target region to the center point of the second target region to obtain a plurality of first distances; Calculating the average value of the plurality of first distances to obtain a first average distance; Taking the first average distance as the first target value.

13. The method according to claim 9, wherein The size of the to-be-recognized sample image is smaller than that of the driving sample image, and the second direction offset map label is constructed in the following manner: Constructing a second initial offset map label, where the pixel values of all pixel points on the second initial offset map label are 0; Annotate the contour points of the sample target on the driving sample image to obtain a contour line composed of the contour points; Divide the driving sample image into a preset number of pixel regions, where each pixel region in the preset number of pixel regions corresponds to each pixel point in the second initial offset map label; Take the pixel region where any contour point on the contour line is located as the third target pixel region; Set the pixel value of the pixel point corresponding to the third target pixel region on the second initial offset map label to a second target value to obtain the second direction offset map label.

14. The method according to claim 13, wherein The second target value is obtained by the following method: Obtain the second distance from each contour point in the third target region to the center point of the third target region to obtain a plurality of second distances; Calculate the average value of the plurality of second distances to obtain a second average distance; Take the second average distance as the second target value.

15. The method according to any one of claims 1 to 5, characterized in that, The processing of the driving image to obtain a to-be-detected image includes: Perform a cropping process on the driving image to obtain the to-be-detected image.

16. The method according to any one of claims 1 to 5, characterized in that The processing of the driving image to obtain a to-be-detected image includes: Perform a downsampling process on the driving image to obtain the to-be-detected image.

17. The method according to any one of claims 1 to 5, characterized in that The processing of the driving image to obtain a to-be-detected image includes: Perform a trimming process on the driving image to obtain a trimmed image; Perform a downsampling process on the trimmed image to obtain the to-be-detected image.

18. The method according to any one of claims 1 to 5, characterized in that It further includes: Through a second coordinate conversion relationship, convert the second position of the to-be-detected target in the second coordinate system where the driving image is located into a sixth position in the third coordinate system corresponding to the vehicle, where the second coordinate conversion relationship is the conversion relationship between the second coordinate system and the third coordinate system.

19. A target detection device, characterized in that, It includes: A processing module, configured to acquire a driving image collected during the driving of the vehicle and process the driving image to obtain a to-be-detected image, where the to-be-detected image is an image containing a to-be-detected target; A prediction module, configured to input the to-be-detected image into a prediction model to obtain a prediction image for predicting the edge points of the to-be-detected target in the to-be-detected image; A first position acquisition module, configured to acquire the first position of the edge points in the first coordinate system where the to-be-detected image is located according to the prediction image; A second position acquisition module, configured to obtain the second position of the to-be-detected target in the second coordinate system where the driving image is located according to the first position where the edge points in the prediction image are located.

20. A vehicle, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, when the processor is configured to execute the instructions, it implements the steps of the method according to any one of claims 1 to 18.

21. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, it implements the steps of the method according to any one of claims 1 to 18.