Target detection method, target detection model training method and related product

By integrating multimedia data and radar data into a target detection method and combining it with style transfer technology to train the model, the problem of low accuracy in traditional target detection under complex environments is solved, thereby improving detection accuracy and the safety of autonomous driving.

CN121330641APending Publication Date: 2026-01-13BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410930991.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Traditional target detection methods have low accuracy in complex environments, are easily affected by multimedia data, and have poor detection capabilities for some objects using radar data.

Method used

By fusing multimedia and radar data from multiple time points, target detection is performed using the unaffected features of radar data in complex environments. Detection is further enhanced by combining multimedia data features with style transfer techniques to train the samples, thereby improving detection accuracy.

Benefits of technology

It improves the accuracy of target detection in complex environments, avoids the inaccuracies of multimedia data detection and the insufficient perception of some objects by radar data, and enhances the accuracy of target detection and the safety of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330641A_ABST
    Figure CN121330641A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a target detection method, a target detection model training method and a related product. The target detection method comprises the following steps: acquiring target multimedia data and target radar data at multiple moments; fusing the target multimedia data and the target radar data at the multiple moments to obtain target data; and performing target detection based on the target data. According to the embodiment of the invention, the method can achieve the fusion of the multimedia data and the radar data at a plurality of moments, carries out the target detection, and improves the target detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an object detection method, an object detection model training method, and related products. Background Technology

[0002] With the development of autonomous driving technology, the safety of autonomous vehicles has received widespread attention. Object detection is a crucial component of autonomous driving. In real-world parking environments, various obstacles such as pedestrians, vehicles, and fences may be encountered. Therefore, effective obstacle identification can provide sufficient guidance for subsequent decision-making in autonomous driving.

[0003] Traditional target detection methods use visual sensors to acquire images of the environment around the vehicle. Then, the acquired environmental images are segmented to obtain obstacles in the environmental images and the relative position information of the obstacles and the vehicle.

[0004] However, in complex environments, target detection is easily affected, resulting in lower target detection accuracy. Summary of the Invention

[0005] This application provides a target detection method, a target detection model training method, and related products. By fusing multimedia data and radar data from multiple time points, target detection is not affected by radar data in complex environments. The fused features are used for target detection, and target detection can be performed by combining multimedia data features without being affected by complex environments, thus improving the accuracy of target detection.

[0006] In a first aspect, embodiments of this application provide a target detection method, including:

[0007] Acquire target multimedia data and target radar data at multiple time points;

[0008] The target multimedia data and target radar data at the multiple time points are fused to obtain the target data;

[0009] Target detection is performed based on the target data.

[0010] Secondly, embodiments of this application provide a method for training an object detection model, including:

[0011] Obtain multiple original samples and multiple style samples;

[0012] Based on the multiple style samples, style transfer is performed on the multiple original samples to obtain multiple transferred samples;

[0013] Based on the multiple original samples and the multiple transfer samples, multiple target samples are obtained;

[0014] The target detection model is obtained by training the model based on the multiple target samples.

[0015] Thirdly, embodiments of this application provide a target detection device, which includes an acquisition unit and a detection unit;

[0016] The acquisition unit is used to acquire the first multimedia data and the first radar data at the current moment;

[0017] The detection unit is used to perform target detection based on the first multimedia data and the first radar data.

[0018] Fourthly, embodiments of this application provide a target detection model training device, which includes an acquisition unit and a processing unit;

[0019] The acquisition unit is used to acquire multiple original samples and multiple style samples;

[0020] The processing unit is used to perform style transfer on the multiple original samples based on the multiple style samples to obtain multiple transferred samples;

[0021] Based on the multiple original samples and the multiple transfer samples, multiple target samples are obtained;

[0022] The target detection model is obtained by training the model based on the multiple target samples.

[0023] Fifthly, embodiments of this application provide an electronic device, including: a processor and a memory, the processor being connected to the memory, the memory being used to store a computer program, and the processor being used to execute the computer program stored in the memory, so that the electronic device performs the method as described in the first or second aspect.

[0024] In a sixth aspect, embodiments of this application provide a computer-readable storage medium storing a computer program that causes a computer to perform the method described in the first or second aspect.

[0025] In a seventh aspect, embodiments of this application provide a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program, the computer being operable to perform the method as described in the first or second aspect.

[0026] Eighthly, embodiments of this application provide a vehicle, the vehicle including the target detection device as described in the third aspect and the target detection model training device as described in the fourth aspect.

[0027] Implementing the embodiments of this application has the following beneficial effects:

[0028] As can be seen, in this embodiment, target multimedia data and target radar data collected at multiple times can be acquired, fused, and then used for target detection. This avoids the problem of inaccurate target detection caused by complex environments when using multimedia data, thus improving the accuracy of target detection. Furthermore, it avoids the problem of poor object perception when using radar data, further improving the accuracy of target detection. By fusing target multimedia data and target radar data from multiple times, the target data can include more data features, further improving the accuracy of target detection. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of a vehicle target detection system provided in an embodiment of this application;

[0031] Figure 2 This is a schematic diagram illustrating an application scenario of a target detection method provided in an embodiment of this application;

[0032] Figure 3 A flowchart illustrating a target detection model training method provided in an embodiment of this application;

[0033] Figure 4 A schematic diagram of a target detection model provided in an embodiment of this application;

[0034] Figure 5 A schematic diagram of a feature fusion model provided in an embodiment of this application;

[0035] Figure 6 A schematic diagram illustrating a style transfer effect provided in an embodiment of this application;

[0036] Figure 7 This is a schematic diagram of the structure of a target detection model provided in an embodiment of this application;

[0037] Figure 8 A schematic flowchart of a target detection method provided in an embodiment of this application;

[0038] Figure 9 A flowchart of a target detection model provided in an embodiment of this application;

[0039] Figure 10 A schematic diagram of a target detection device provided in an embodiment of this application;

[0040] Figure 11 This is a schematic diagram of a target detection model training device provided in an embodiment of this application;

[0041] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0043] The terms "first," "second," "third," and "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0044] In this document, the term "embodiment" means that a particular feature, result, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0045] First, refer to Figure 1 , Figure 1 This is a schematic diagram of a vehicle target detection system provided in an embodiment of this application. Figure 1 As shown, the target detection system includes: a target detection device 100, a camera 102, a radar 103, and a display device 104.

[0046] The camera 102 is a multimedia data acquisition device for the vehicle, used to collect multimedia data, such as image data or video data, within the vehicle's field of view. Optionally, the camera 102 may include, for example, a forward-facing ranging camera, a 360-degree surround-view camera, a rear-view reversing camera, a low-light night vision camera, a thermal imaging camera, an infrared camera, a monocular camera, a multi-view camera, etc., without limitation. The radar 103 is a safety assistance and positioning device for the vehicle, which helps the vehicle achieve environmental detection, precise positioning, and adaptive cruise control by sending detection waves and receiving reflected waves. Optionally, the radar 103 may include, for example, millimeter-wave radar, ultrasonic radar, lidar, pulse radar, frequency-modulated continuous wave (FMCW) radar, microwave shock radar, etc., without limitation.

[0047] In this embodiment, the multimedia data collected by camera 102 and the radar data collected by radar 103 can be transmitted to the target detection device 100 in real time. The target detection device 100 performs target detection on the multimedia data collected by camera 102 and the radar data collected by radar 103 to obtain target information within the vehicle's field of vision, so as to plan a driving route for the vehicle based on the target information within the vehicle's field of vision. Optionally, the target detection device 100 can send the multimedia data collected by camera 102 and the radar data collected by radar 103 to a display device 104, and display the multimedia data and radar data on the display device 104. Optionally, the display device 104 can be a vehicle tablet, vehicle computer, etc. Responding to user operation, the display device 104 can obtain the multimedia data collected by camera 102 and the real-time radar data collected by radar 103 from the target detection device 100, and display the multimedia data and the real-time radar data on the display device 104 to instruct the user to perform vehicle driving operations based on the multimedia data and the real-time radar data.

[0048] Optionally, the target detection model used by the target detection device 100 for target detection is trained by the target detection model training device 101. It should be noted that the target detection device 100 can interact with the target detection model training device 101 in real time to perform target detection on multimedia data collected by the camera 102 and radar data collected by the radar 103. The target detection model training device 101 is used to train the target detection model and provides the trained target detection model to the target detection device 100, enabling the target detection device 100 to perform target detection on the multimedia data collected by the camera 102 and the radar data collected by the radar 103 based on the target detection model. The training samples for the target detection model training device 101 can be obtained by real-time acquisition from the camera and radar, and after appropriate preprocessing. Optionally, the target detection device 100 and the target detection model training device 101 can be high-performance computing units such as the vehicle's central processing unit, domain controller, electronic control unit (ECU), in-car application service (ICAS), and main control unit (MCU).

[0049] For example, the target detection device 100 detects target information based on multimedia data collected by the camera 102 and radar data collected by the radar 103. For example, after detecting obstacle information, the detected target information can be sent to the display device 104, and the display device 104 can display the detected target information to indicate the presence of target objects in the user's field of vision during vehicle operation.

[0050] In the embodiments of this application, the target detection method is mainly applied to the autonomous driving technology of vehicles. The vehicle detects targets within its field of vision and plans a driving route based on the target detection results, thereby improving the safety of autonomous driving.

[0051] like Figure 2 As shown, Figure 2 This is a schematic diagram illustrating an application scenario of a target detection method provided in an embodiment of this application. When vehicle 201 is traveling on a road, its camera can capture images or videos within its field of view 202 in real time. Furthermore, vehicle 201's radar can send detection waves 203 towards the target area and receive the returned reflected waves to obtain corresponding radar data, thereby determining the object category in the target area based on the radar data. For example... Figure 2As shown, vehicle 201 sends a detection wave 203 to obstacle 204 via radar. After receiving the returned reflected wave, the target detection device can analyze the corresponding radar data and determine the category of obstacle 204 based on the analyzed radar data. Thus, vehicle 201 can perform route planning based on the category of obstacle 204, thereby controlling vehicle 201 to avoid obstacle 204.

[0052] For example, when the camera of vehicle 201 captures images of traffic light 205 and the zebra crossing, or obtains images of traffic light 205 and the zebra crossing from the video captured by the camera, and determines the categories of traffic light 205 and the zebra crossing by combining them with radar data, the vehicle can stop and wait at the zebra crossing when traffic light 205 is red. When the traffic light turns green, vehicle 201 can also determine whether there are pedestrians 206 on the zebra crossing by real-time collected multimedia data and radar data, thereby determining whether to cross the zebra crossing.

[0053] It should be noted that when vehicle 201 is driving in complex environments, such as in low light or in adverse weather, the detection effect of using multimedia data collected by cameras for target detection is easily affected, resulting in lower detection accuracy. Furthermore, when using radar data for target detection, the detection results are affected by the radar's detection waves and the type of obstacle, leading to lower target detection accuracy.

[0054] Therefore, in the target detection method of this application embodiment, which is applied to the above-mentioned target detection system and the above-mentioned application scenario, the target detection device 100 acquires the first multimedia data and the first radar data at the current moment;

[0055] The target detection device 100 performs target detection based on first multimedia data and first radar data.

[0056] In the target detection model training method of this application embodiment, the target detection model training device 101 acquires multiple original samples and multiple style samples;

[0057] The object detection model training device 101 performs style transfer on multiple original samples based on multiple style samples to obtain multiple transferred samples.

[0058] The target detection model training device 101 obtains multiple target samples based on multiple original samples and multiple transfer samples;

[0059] The target detection model training device 101 trains the model based on multiple target samples to obtain the target detection model.

[0060] As can be seen, in the target detection system and application scenarios described above, vehicles can acquire multimedia data within their field of view via cameras and detect the surrounding environment using radar to obtain radar data. The target detection device can then perform target detection based on this multimedia and radar data, obtaining the target detection result. Route planning and control are then performed based on the target detection result, which is displayed on a display device. The target detection model used in the target detection process is obtained through sample training by a target detection model training device. Therefore, the target detection device can acquire multimedia data from cameras and radar data from radar, perform target detection on this data, and plan the vehicle's route based on the target detection result, thus improving the accuracy of target detection and the safety of vehicle driving.

[0061] To facilitate understanding of the object detection method in the embodiments of this application, the training process of the object detection model will be described below in conjunction with the object detection model. For example... Figure 4 As shown, the target detection model includes a sample construction model, a target detection model, and a target prediction model.

[0062] See Figure 3 , Figure 3 This is a flowchart illustrating a method for training an object detection model according to an embodiment of this application. The method is applied to the aforementioned object detection model training apparatus and includes, but is not limited to, the following steps:

[0063] 301: Obtain multiple original samples and multiple style samples.

[0064] In this embodiment, the vehicle's visual sensors and radar, such as a monocular camera and millimeter-wave radar, can collect multiple multimedia data and multiple radar data within the vehicle's field of view. These multimedia data and radar data are then fused to obtain multiple sample data. Preprocessing of these sample data yields multiple raw samples. Furthermore, the monocular camera and millimeter-wave radar can collect multiple style sample data under various weather conditions, lighting conditions, etc., and preprocessing of these style sample data yields multiple style samples.

[0065] Optionally, multiple multimedia and radar data sets can be collected using a monocular camera and millimeter-wave radar. Data collection can be conducted multiple times on both closed and open test roads within the park. Multiple test vehicles can be deployed on closed test sections within the park to simulate various traffic environments, enriching the diversity of the collected multimedia and radar data. Furthermore, specific scenarios incorporating various obstacles can be set up to include more relevant information in the collected multimedia and radar data.

[0066] For example, a monocular camera can continuously collect multimedia data from multiple moments at a time, and a millimeter-wave radar can continuously collect radar data from multiple moments at a time. The following description will take the collected multimedia data as an image as an example. The processing method when the collected multimedia data is video is similar to that in the embodiments of this application, and will not be repeated here.

[0067] In one feasible embodiment, such as Figure 5 As shown, this embodiment of the application will be described using data collected continuously at three time points by a camera and radar. The method of collecting data at multiple time points is similar to this embodiment and will not be repeated here. The image collected by the monocular camera at each time point includes an RGB three-channel image, and the radar data collected by the radar includes two-channel radar data: the distance (range) between the target and the millimeter-wave radar, and the radial velocity (doppler) of the target relative to the millimeter-wave radar. Figure 5 As shown, the images acquired consecutively at three time points are Image 1, Image 2, and Image 3, each image including R, G, and B channels. The radar data acquired consecutively at three time points are Radar 1, Radar 2, and Radar 3, each radar data point including range and doppler data. During image data preprocessing, the RGB three-channel images of each time point are first fused with the two-channel radar data of each time point, that is, the three-channel images are fused with the two-channel radar data to obtain each 5-channel feature map. For example, a stitching method can be used for fusion, stitching together the R, G, and B channels of Image 1 with the range and doppler data of Radar 1 to obtain a 5-channel feature map. Figure 1 .

[0068] It should be noted that the embodiments of this application only illustrate the image at each moment, including the RGB three-channel image of the image, and each radar data, including range and doppler data. Those skilled in the art can also fuse other features of the image and radar data. The method of fusing other features of the image and radar data is similar to that of this application and will not be described here. Similarly, the multi-channel feature map obtained by fusing other features of the image and radar data is similar to the 5-channel feature map and 15-channel feature map of the embodiments of this application and will not be described here.

[0069] Using the above method, three time-time corresponding 5-channel feature maps can be obtained each time, i.e., 5-channel features. Figure 1 , 5 Channel characteristics Figure 2 and 5 Channel characteristics Figure 3 Furthermore, by concatenating the 5-channel feature maps obtained at each of the three time points, we can obtain, as shown below. Figure 5The 15-channel feature map is shown. Then, it is checked whether the 5-channel feature maps corresponding to the three time points obtained each time include obstacle images. The 15-channel feature maps that do not contain obstacle images are deleted, thus obtaining multiple original samples. It should be noted that both the 5-channel and 15-channel feature maps obtained by stitching are stitched together vertically, that is, the feature maps of each channel are superimposed. Optionally, the point cloud image of the radar data can also be projected onto the image data to fuse the image data and the radar data; this is not limited here.

[0070] It should be noted that the collected style samples should include various lighting and weather conditions, and the nature of the target should remain unchanged under these conditions. Therefore, by using multiple style samples to perform style transfer on multiple original samples, target detection scenarios under various lighting and weather conditions can be simulated, improving the accuracy of target detection.

[0071] 302: Based on multiple style samples, perform style transfer on multiple original samples to obtain multiple transferred samples.

[0072] In this embodiment, multiple style samples are used to perform style transfer on multiple original samples, which can simulate target detection scenarios under various lighting and climatic conditions.

[0073] For example, such as Figure 4 As shown, multiple original samples and multiple style samples are input into the sample construction model to construct the dataset samples. Specifically, the multiple original samples and multiple style samples are first input into the style transfer model, which uses the multiple style samples to perform style transfer on the multiple original samples. Optionally, the style transfer model can be the AnimeGANv2 model. The AnimeGANv2 model is a deep learning model for image or video stylization and animation, which combines neural style transfer and Generative Adversarial Networks (GANs) to achieve image or video stylization and animation.

[0074] It should be noted that the embodiments in this application are only illustrated using the AnimeGANv2 model as the style transfer model. The methods of using other style transfer models are similar to this method and will not be described here.

[0075] In this embodiment, the AnimeGANv2 model is used to transfer style samples, including weather conditions such as rain and snow, as well as various lighting conditions, to multiple original samples, resulting in multiple transferred samples. It should be noted that since the original samples are preprocessed and each original sample contains a target, the resulting multiple transferred samples are images or videos containing targets under various weather and lighting conditions. Figure 6 As shown, Figure 6 The image above is the original sample collected. Figure 6 The image below shows the style-transfer sample. Figure 6 The migrated sample is based on the original sample, with brighter lighting and snow added to the vehicle driving road and surrounding environment of the original sample to simulate the environmental conditions for vehicle target detection on a snowy day.

[0076] It should be noted that each original sample is formed by fusing third-level multimedia data and third-level radar data from multiple time points, and each transfer sample is formed by fusing fourth-level multimedia data and fourth-level radar data from multiple time points. Specifically, the vehicle collects third-level multimedia data from multiple time points via cameras and radar data from multiple time points each time, and then fuses the third-level multimedia data and third-level radar data from each time point to obtain multiple original samples. Thus, each original sample includes 15 channels of features. Similarly, each transfer sample is formed by fusing fourth-level multimedia data and fourth-level radar data from multiple time points, and each transfer sample includes 15 channels of features.

[0077] 303: Based on multiple original samples and multiple transfer samples, multiple target samples are obtained.

[0078] In this embodiment of the application, a target detection dataset can be constructed using the multiple original samples and multiple transfer samples obtained above, wherein the samples in the target detection dataset are the multiple target samples mentioned above.

[0079] like Figure 4 As shown, multiple target samples are obtained based on multiple original samples and multiple transfer samples. For example, this may include the following steps:

[0080] Multiple original samples are fused with multiple transferred samples to obtain multiple synthetic samples;

[0081] Multiple original samples, multiple transfer samples, and multiple synthetic samples are used as multiple target samples.

[0082] In one feasible embodiment, multiple original samples and multiple migrated samples can be fused by splicing.

[0083] Specifically, such as Figure 4 As shown, by concatenating multiple original samples with multiple transfer samples, such as... Figure 4The C operation shown yields multiple synthetic samples. It should be noted that in this embodiment, the third multimedia data of the original samples at multiple time points is stitched left-right to the fourth multimedia data of the migrated samples. Taking a frame from an acquired image or video as an example, the RGB images of three consecutive time points of each original sample can be stitched left-right to the RGB images of three consecutive time points of the migrated samples. Similarly, the third radar data of the original samples at multiple time points is stitched left-right to the fourth radar data of the migrated samples, meaning that the two-channel radar data of three consecutive time points of each original sample is stitched left-right to the two-channel radar data of three consecutive time points of the migrated samples. Furthermore, by using left-right stitching to stitch each third multimedia data point with its corresponding fourth multimedia data, and by stitching the third radar data with its corresponding fourth radar data, the number of feature channels remains unchanged; thus, the resulting multiple synthetic samples include 15 feature channels.

[0084] For example, before fusing multiple original samples with multiple transferred samples to obtain multiple synthetic samples, the following steps may also be included:

[0085] Multiple sets of third-party multimedia data at the current moment are used as third-party multimedia data at multiple moments of the target's original sample; and multiple sets of third-party radar data at the current moment are used as third-party radar data at multiple moments of the target's original sample.

[0086] In this embodiment, the target original sample is any one of the plurality of original samples. The following explanation uses an example where the original sample includes multimedia data and radar data from three different time points.

[0087] Specifically, the target's original sample collected at the current moment does not have third multimedia data and third radar data at the three time points. That is, when the current moment is the first time the third multimedia data is collected, the third multimedia data at the current moment is copied twice to obtain three copies of the third multimedia data. These three copies of the third multimedia data are used as the third multimedia data of the target's original sample at the three time points.

[0088] Optionally, when the current moment is the second time to collect the third multimedia data, a copy of the third multimedia data at the current moment is made, and the current moment, the previous moment, and the copied third multimedia data are used as the third multimedia data at the three moments of the target original sample.

[0089] Optionally, when the current moment is the first time third radar data is acquired, the third radar data at the current moment is copied twice, resulting in three copies of the third radar data. These three copies of the third radar data are used as the third radar data for the target's original sample at three different times. When the current moment is the second time third radar data is acquired, the third radar data at the current moment is copied once, and the third radar data at the current moment, the previous moment, and the copied data are used as the third radar data for the target's original sample at three different times.

[0090] As can be seen, when multiple times of the third-channel multimedia data and third-channel radar data of the original target sample are missing, multiple copies of the third-channel multimedia data of the original target sample at the current time can be copied as the third-channel multimedia data of the original target sample at multiple times. Similarly, multiple copies of the third-channel radar data of the original target sample at the current time can be copied as the third-channel radar data of the original target sample at multiple times. Thus, each original sample includes third-channel multimedia data and third-channel radar data from multiple times. After fusing the third-channel multimedia data and third-channel radar data, each multi-channel feature, i.e., the synthetic sample, is obtained, thereby constructing a target detection dataset and improving the training accuracy of the target detection model.

[0091] Furthermore, multiple original samples, multiple transferred samples, and multiple synthetic samples are used together as the target detection dataset, thus obtaining multiple target samples. It should be noted that since the multiple original samples, multiple transferred samples, and multiple synthetic samples obtained above all have 15-channel features, they can be used as multiple target samples to train the target detection model, that is, each training sample has 15-channel features.

[0092] It can be seen that by fusing the third multimedia data of each original sample with the fourth multimedia data of each transferred sample in a one-to-one correspondence, and by fusing the third radar data of each original sample with the fourth radar data of each transferred sample in a one-to-one correspondence, multiple synthetic samples can be obtained. Thus, multiple original samples, multiple transferred samples, and multiple synthetic samples are used together as a target detection dataset, that is, as multiple target samples, to train the target detection model, thereby increasing the richness and accuracy of the training samples, improving the training effect of the target detection model, and thus improving the accuracy of target detection using the target detection model.

[0093] 304: The target detection model is obtained by training the model based on multiple target samples.

[0094] In this embodiment of the application, the above-mentioned multiple target samples can be input into... Figure 4 The target detection model shown is trained to obtain the target detection model.

[0095] For example, each target sample includes second location information and a second category for the target, where the second location information and the second category are pre-labeled location information and category. Training the model based on multiple target samples yields a target detection model, which may include, for example, the following steps:

[0096] Upsample each target sample to obtain the fourth feature corresponding to each target sample;

[0097] The fourth feature corresponding to each target sample is fused to obtain the second fused feature corresponding to each target sample.

[0098] Feature extraction is performed on the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample;

[0099] Based on the second target features corresponding to each target sample, predict the third location information and third category of the target in each target sample;

[0100] The first loss is obtained based on the second and third location information of the target in each target sample;

[0101] The second loss is derived based on the second and third categories of the target in each target sample;

[0102] The target detection model is obtained by training the model based on the first loss and the second loss.

[0103] In the embodiments of this application, such as Figure 4 As shown, by inputting multiple target samples into the target detection model, each target sample is first upsampled. It should be noted that deep learning models typically receive input feature maps of 224×224 pixels. In traditional techniques, resizing target samples to 224×224 pixels to obtain feature maps easily leads to loss of feature map information. Therefore, in this embodiment, each target sample is upsampled, for example, resized to a resolution of 672×672 pixels, to obtain the fourth feature. When the target sample is an image, this fourth feature can be a feature map.

[0104] It should be noted that in this embodiment, only the pixel resolution of the fourth feature is described as 672×672. The processing method for the fourth feature with other resolutions is similar to this method and will not be described here.

[0105] In one feasible embodiment, such as Figure 7As shown, feature fusion is performed on the fourth feature corresponding to each target sample. For example, the 672×672×15 fourth feature can be input into a convolutional head for feature fusion, fusing the features included in the fourth feature with radar data features to obtain the second fused feature. The convolutional head sequentially includes: a convolutional layer with a kernel of 7 and a stride of 2, a Parametric Rectified Linear Unit (PReLU) nonlinear activation layer, a convolutional layer with a kernel of 5 and a dilation rate of 2, a Layer Normalization (LN) layer, and a Global Average Pooling (GAP) layer. By processing the fourth feature of each target sample through the above convolutional head, a second fused feature of 224×224 pixels corresponding to each target sample can be obtained.

[0106] It is understood that this embodiment only illustrates the example of feature fusion of the fourth feature of each target sample using a convolutional head. Other methods for feature fusion of the fourth feature of each target sample are similar and are not limited here. Through feature fusion of the fourth feature, a second fused feature of 224×224 pixels can be obtained, which can then be input into the deep learning model.

[0107] Furthermore, such as Figure 7 As shown, feature extraction is performed on the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample. The feature extraction of the second fusion feature includes: fourth attention processing, fifth attention processing, and sixth attention processing.

[0108] In one feasible embodiment, such as Figure 7 As shown, the ConvNeXt model is used as the basic backbone network in the feature extraction process. The basic backbone network consists of multiple backbone network blocks, and each backbone network block has a corresponding attention mechanism added, such as... Figure 7 The fourth, fifth, and sixth attention mechanisms are shown. The ConvNeXt model is an improvement on the ResNet network, achieving better classification accuracy on the ImageNet dataset compared to the original ResNet model. It should be noted that this application's embodiments only illustrate the use of the ConvNeXt model as the basic backbone network for feature extraction. Those skilled in the art can use other backbone networks for feature extraction, and the methods for feature extraction using other backbone networks are similar to this method and are not limited here.

[0109] For example, feature extraction is performed on the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample. This may include the following steps:

[0110] A fourth attention process is applied to the second fusion feature corresponding to each target sample to obtain the fifth feature corresponding to each target sample;

[0111] Perform fifth attention processing on the fifth feature corresponding to each target sample to obtain the sixth feature corresponding to each target sample;

[0112] The sixth attention process is applied to the sixth feature corresponding to each target sample to obtain the second target feature corresponding to each target sample.

[0113] In this embodiment of the application, firstly, a fourth attention process is applied to the second fusion feature corresponding to each target sample. The fourth attention mechanism can be, for example, a Coordinate attention mechanism. For example,... Figure 7 As shown, the fourth attention processing includes: N1 backbone network blocks with added Coordinate attention mechanisms, i.e., ConvNeXt blocks, each with a dimension (dim) of 96. The second fusion feature corresponding to each target sample is passed through the connection layers of N1 ConvNeXt blocks and the Coordinate attention mechanism, and then downsampled twice. The first downsampling yields a 56×56×96 feature, and the second downsampling yields a 28×28×192 feature, which is the fifth feature corresponding to each target sample. The fourth attention processing on the second fusion feature increases the number of channels in the resulting fifth feature, meaning more features are extracted, which is beneficial for feature analysis in the target detection model. Furthermore, the fourth attention processing on the second fusion feature reduces the size of the resulting fifth feature, concentrating the extracted features in valuable regions, i.e., regions containing targets, thereby improving the accuracy of the target detection model training.

[0114] It should be noted that the embodiments in this application are only used as an example of using the connection between ConvNeXt Block and Coordinate attention mechanism to perform fourth attention processing. Those skilled in the art can also use other attention mechanisms to perform fourth attention processing or increase the number of connection layers between ConvNeXt Block and Coordinate attention mechanism to perform fourth attention processing, which is not limited here.

[0115] Furthermore, a fifth attention process is applied to the fifth feature corresponding to each target sample. This fifth attention mechanism can, for example, employ the ECANet attention mechanism. For instance, ... Figure 7 As shown, the fifth attention process includes: N2 backbone network blocks connected to the ECANet attention mechanism, i.e., ConvNeXt Blocks, each with a depth of 96. The fifth feature corresponding to each target sample is concentrated in the region where the target is located through the connection layer of N2 ConvNeXt Blocks and the ECANet attention mechanism. Then, a downsampling is performed to obtain a 14×14×384 feature, which is the sixth feature corresponding to each target sample.

[0116] It should be noted that the embodiments in this application are only used as an example of using the connection between ConvNeXt Block and ECANet attention mechanism to perform the fifth attention processing. Those skilled in the art can also use other attention mechanisms to perform the fifth attention processing or increase the number of connection layers between ConvNeXt Block and ECANet attention mechanism to perform the fifth attention processing, which is not limited here.

[0117] Furthermore, a sixth attention process is applied to the sixth feature corresponding to each target sample. This sixth attention mechanism can be, for example, the SimAM attention mechanism. For instance,... Figure 7 As shown, the sixth attention processing includes: N3 backbone network blocks connected to the SimAM attention mechanism, i.e., ConvNeXt Blocks, each with a dim value of 96. The sixth feature corresponding to each target sample is concentrated in the region where the target is located through the connection layers of the N3 ConvNeXt Blocks and the SimAM attention mechanism. Then, the feature data processed by the SimAM attention mechanism is subjected to global average pooling to obtain a 7×7×768 feature. This 7×7×768 feature is then normalized and passed through a fully connected layer to obtain the second target feature corresponding to each target sample.

[0118] It should be noted that the embodiments in this application are only used as an example of using the connection between ConvNeXt Block and SimAM attention mechanism to perform sixth attention processing. Those skilled in the art can also use other attention mechanisms to perform sixth attention processing or increase the number of connection layers between ConvNeXt Block and SimAM attention mechanism to perform sixth attention processing, which is not limited here.

[0119] It is understandable that using a combination of multiple attention mechanisms for feature extraction results in higher accuracy compared to using a single attention mechanism, thereby improving the accuracy of target detection.

[0120] As can be seen, in this embodiment, by performing a fourth attention process on the second fusion feature corresponding to each target sample, a fifth feature corresponding to each target sample can be obtained. Then, a fifth attention process is performed on the fifth feature corresponding to each target sample to obtain a sixth feature corresponding to each target sample. Finally, a sixth attention process is performed on the sixth feature corresponding to each target sample to obtain a second target feature corresponding to each target sample. Thus, after the fourth, fifth, and sixth attention processes, feature extraction of the second fusion feature corresponding to each target sample can be completed, resulting in a second target feature corresponding to each target sample. Target detection is then performed on the second target feature corresponding to each target sample. The extracted second target feature contains a large amount of feature data, improving the accuracy of the target detection model training.

[0121] Then, the second target feature corresponding to each target sample is input into the target prediction model. For example, the target prediction model can be a YOLO detector head. By predicting the second target feature corresponding to each target sample using this YOLO detector head, the third location information and third category of the target in each target sample can be obtained. It should be noted that this embodiment only uses a YOLO detector head as an example for illustration; those skilled in the art can also use other target prediction models to predict the second target feature corresponding to each target sample, and this is not limited thereto.

[0122] For example, based on the second target features corresponding to each target sample, predicting the third location information and third category of the target in each target sample may include the following steps:

[0123] Based on the second target feature corresponding to each target sample, multiple second prediction boxes are obtained for each target sample;

[0124] The prediction for each second prediction box includes the confidence level of the target;

[0125] The second predicted box with a confidence level greater than the second threshold is used as the second target predicted box for each target sample.

[0126] The location information of the second target prediction box corresponding to each target sample is used as the third location information of the target in each target sample;

[0127] Predict the probability of each category for the second target bounding box corresponding to each target sample;

[0128] The category with the highest probability is used as the third category of the target in each target sample.

[0129] In this embodiment, the second target feature corresponding to each target sample is input into the YOLO detection head. In the YOLO detection head, the second target feature corresponding to each target sample is first divided into blocks, creating multiple grids. Then, multiple consecutive grids are combined to form a bounding box, resulting in multiple bounding boxes, i.e., multiple second prediction boxes corresponding to each target sample.

[0130] Furthermore, the grid where the target center point is located and the grid where the target contour is located are predicted, and the confidence level of including the target in each second prediction box is predicted, that is, the confidence level of the grid where the target contour is located. The second prediction boxes with a confidence level greater than a second threshold are used as the second target prediction boxes corresponding to each target sample.

[0131] Then, based on the center point of the second target prediction bounding box and the predicted center point of the target, the center point of the second target prediction bounding box is adjusted to the predicted center point of the target to obtain the position information of the second target prediction bounding box. For example, the position information of the second target prediction bounding box can be (x, y, w, h), where (x, y) represents the coordinates of the center point of the second target prediction bounding box, w represents the width of the second target prediction bounding box, and h represents the height of the second target prediction bounding box. The position information of the second target prediction bounding box corresponding to each target sample is used as the third position information of the target in each target sample. Optionally, the second target prediction bounding box can be a three-dimensional cylindrical box to select the target.

[0132] Next, the second target prediction box corresponding to each target sample is classified, and the probability of each second target prediction box for each category is predicted. The category with the highest probability is taken as the third category of the target in each target sample.

[0133] As can be seen, segmenting the second target feature of each target sample yields multiple second predicted bounding boxes for each target sample. By predicting the confidence level of the grid where the target contour is located, the second predicted bounding boxes with a confidence level greater than a second threshold are used as the second target predicted bounding boxes for each target sample. Based on the positional information of the second target predicted bounding boxes, the third positional information of the target in each target sample is obtained. Furthermore, by predicting the probability of each target sample's corresponding second target predicted bounding box for each category, the category with the highest probability is used as the third category of the target in each target sample. Therefore, the target prediction model can predict the third positional information and third category of the target in each target sample, enabling the training of the target detection model and improving the accuracy of target detection.

[0134] Furthermore, based on the second location information and the predicted third location information of the target in each target sample, a first loss is calculated, and the training parameters of the model are adjusted based on the first loss to optimize the model training. Additionally, based on the second category and the predicted third category of the target in each target sample, a second loss is calculated, and the training parameters of the model are adjusted based on the second loss to optimize the model training and improve the training accuracy of the target detection model. After the above model training, the target detection model can be obtained.

[0135] Therefore, by upsampling each target sample, a fourth feature corresponding to each target sample is obtained. Feature fusion is then performed on each fourth feature to obtain a second fused feature. Feature extraction is then performed on each second fused feature to obtain a second target feature corresponding to each target sample. Based on the second target feature corresponding to each target sample, the third location information and third category of the target in each target sample are predicted. A first loss is obtained based on each second and third location information, and the training parameters of the model are adjusted based on the first loss to optimize the model's training. A second loss is obtained based on each second and third category, and the training parameters of the model are adjusted based on the second loss to optimize the model's training, thereby improving the training accuracy of the target detection model and obtaining a high-performance target detection model.

[0136] As can be seen, in the embodiments of this application, when training the target detection model, the original samples, transfer samples and synthetic samples are used together as target samples for model training. Furthermore, each target sample is composed of multimedia data and radar data fused from multiple time points, which improves the diversity of training samples. Based on the above target samples, a high-performance target detection model can be obtained, thereby improving the accuracy of target detection.

[0137] See Figure 8 , Figure 8 This is a flowchart illustrating a target detection method provided in an embodiment of this application. The method is applied to the aforementioned target detection device, and employs the target detection model trained as described above. The method includes, but is not limited to, the following steps:

[0138] 801: Acquire target multimedia data and target radar data at multiple time points.

[0139] In this embodiment of the application, the target multimedia data at multiple times includes the first multimedia data at the current time and the second multimedia data at multiple historical times, and the target radar data at multiple times includes the first radar data at the current time and the second radar data at multiple historical times.

[0140] The camera can collect real-time multimedia data and send it to the target detection device. Similarly, the radar can collect real-time radar data and send it to the target detection device. Thus, the target detection device can obtain the current multimedia data and radar data. The first multimedia data may include, for example, an image or a video frame from the current moment. Furthermore, the target detection device can acquire second multimedia data and second radar data from multiple historical moments to obtain target multimedia data and target radar data from multiple points in time.

[0141] Optionally, the target detection device can synchronize the first multimedia data and the first radar data to the display device in the vehicle to indicate to the user the vehicle's surrounding environment and driving route at the current moment.

[0142] In this embodiment, the target detection device acquires first multimedia data and first radar data at the current moment, and acquires second multimedia data and second radar data from multiple historical moments. For example, the target detection device can acquire the first multimedia data and first radar data at the current moment, and acquire the second multimedia data and second radar data from the two moments preceding the current moment, thereby obtaining multimedia data and radar data from three consecutive moments.

[0143] It should be noted that the embodiments of this application are only illustrated by the example of a camera acquiring multimedia data at three consecutive moments and a radar acquiring radar data at three consecutive moments. The method for acquiring multimedia data and radar data at multiple moments is similar to this method and will not be described here.

[0144] 802: Target data is obtained by fusing target multimedia data and target radar data from multiple time points.

[0145] In this embodiment of the application, the target detection device can identify target information within the vehicle's field of vision by fusing target multimedia data and target radar data from multiple moments and performing target detection on the fused target data.

[0146] For example, fusing target multimedia data and target radar data from multiple time points to obtain target data may include the following steps:

[0147] The first multimedia data and the first radar data at the current moment are fused to obtain the first fused data;

[0148] By fusing second multimedia data and second radar data from multiple historical moments, multiple second fused data are obtained;

[0149] The first fused data is fused with multiple second fused data to obtain the target data.

[0150] In this embodiment, the target detection device obtains first fused data by fusing the first multimedia data and the first radar data at the current moment. Optionally, the fusion can be performed by splicing, which can be a concatenation operation.

[0151] In one feasible embodiment, multimedia data as an image is used as an example for illustration, such as... Figure 5 As shown, the first multimedia data at each time step includes feature maps of three channels (RGB), and the first radar data at each time step includes radar data of two channels (range and doppler). First, the feature maps of the three channels and the radar data of the two channels are concatenated to obtain the 5-channel feature map corresponding to each time step. For example, [the following is a concatenation process]. Figure 5 The feature maps of the RGB three channels of Image 1 and the radar data of the range and doppler two channels of Radar 1 are concatenated to obtain a 5-channel feature map. Figure 1 .

[0152] Furthermore, the second multimedia data and second radar data from multiple historical moments are stitched together to obtain multiple second fused data sets, where the second multimedia data and second radar data from each historical moment are fused accordingly. For example... Figure 5 As shown, the feature maps of the three RGB channels of image 2 are concatenated with the radar data of the range and doppler channels of radar 2 to obtain a 5-channel feature map. Figure 2 The feature maps of the RGB three channels of image 3 are concatenated with the radar data of the range and doppler channels of radar 3 to obtain a 5-channel feature map. Figure 3 Therefore, we can obtain 5-channel feature maps corresponding to multiple historical moments, which is the second fused data, such as... Figure 5 The 5-channel features shown Figure 2 and 5 Channel characteristics Figure 3 .

[0153] Then, the first fused data is fused with multiple second fused data sets to obtain the target data. For example... Figure 5 As shown, the 5-channel features Figure 1 , 5 Channel characteristics Figure 2 and 5 Channel characteristics Figure 3 The top and bottom layers are stitched together to obtain a 15-channel feature map, which is the target data.

[0154] It should be noted that the embodiments of this application only use the first multimedia data at each time moment, which includes the feature map of the three RGB channels, and the first radar data at each time moment, which includes the radar data of the range and doppler channels, as examples for illustration. Those skilled in the art can also stitch together other features of the multimedia data and radar data. The method of stitching together other features of the multimedia data and radar data is similar to this method. Furthermore, the method of stitching together other features of the multimedia data and radar data to obtain a multi-channel feature map is similar to this method and will not be described here.

[0155] It can be seen that by fusing the first multimedia data and the first radar data at the current moment to obtain the first fused data, and fusing the second multimedia data and the second radar data at multiple historical moments to obtain multiple second fused data, and then fusing the first fused data with the multiple second fused data, the target data can be obtained. Thus, the target data obtained includes multi-channel feature data at multiple moments, which increases the features of the target data. Target detection based on this target data can improve the accuracy of target detection.

[0156] For example, before fusing second multimedia data and second radar data from multiple historical moments to obtain multiple second fused data, the following steps may also be included:

[0157] Multiple sets of first multimedia data at the current moment are used as second multimedia data at multiple historical moments; and multiple sets of first radar data at the current moment are used as second radar data at multiple historical moments.

[0158] In this embodiment of the application, when the vehicle system is first started, the camera and radar are just starting to operate. The target detection device can only acquire the first multimedia data and first radar data collected by the camera at the current moment, but cannot acquire the second multimedia data and second radar data at multiple historical moments.

[0159] At this time, since it is impossible to obtain the second multimedia data and second radar data at multiple historical moments, the target detection device will copy the first multimedia data at the current moment multiple times as the second multimedia data at multiple historical moments, and copy the first radar data at the current moment multiple times as the second radar data at multiple historical moments.

[0160] For example, when the target detection device needs to acquire second multimedia data and second radar data from two historical moments, if only the first multimedia data and first radar data from the current moment exist, then two copies of the first multimedia data from the current moment are copied as the second multimedia data from the two historical moments, and two copies of the first radar data from the current moment are copied as the second radar data from the two historical moments. If only the second multimedia data and second radar data from one historical moment exist, then one copy of the first multimedia data from the current moment is copied as the second multimedia data from the other historical moment, and one copy of the first radar data from the current moment is copied as the second radar data from the other historical moment.

[0161] It can be seen that when it is impossible to obtain second multimedia data and second radar data from multiple historical moments, the target data can be obtained by copying the first multimedia data and first radar data from the current moment multiple times and fusing the multiple copies of the first multimedia data and first radar data, thereby performing target detection.

[0162] 803: Target detection based on target data.

[0163] In this embodiment, the target data includes features of target multimedia data at multiple time points and all features of target radar data. The target detection device can extract a first target feature from the target data by performing operations such as upsampling, feature fusion, and feature extraction, and then perform target detection based on the first target feature.

[0164] For example, target detection based on target data may include the following steps:

[0165] Upsample the target data to obtain the first feature;

[0166] The first feature is fused to obtain the first fused feature;

[0167] The first fused feature is used to extract features to obtain the first target feature;

[0168] Based on the first target features, the first location information and first category of the target are obtained.

[0169] In the embodiments of this application, such as Figure 9As shown, the target data is first upsampled, for example, resized to a 672×672 pixel resolution with 15 channels, which is the first feature. The first feature is then input into a convolutional head for feature fusion. The first feature sequentially passes through a convolutional layer with a kernel of 7 and a stride of 2, a PReLU nonlinear activation layer, a convolutional layer with a kernel of 5 and a dilation rate of 2, a Layer Norm layer, and a global average pooling layer to obtain the first fused feature. This first fused feature incorporates multimedia data and radar data from multiple time points.

[0170] It should be noted that the embodiments of this application are only illustrated using the feature with a first feature of 672×672 pixel resolution and 15 channels as an example. The processing method for other pixel resolutions and channels of the first feature is similar to this method and will not be described here.

[0171] Furthermore, the first fused feature is input into the backbone network for feature extraction.

[0172] Specifically, feature extraction is performed on the first fused feature to obtain the first target feature, which may include the following steps:

[0173] The first fusion feature is subjected to a first attention process to obtain the second feature;

[0174] The second feature is processed with a second attention process to obtain the third feature;

[0175] The third feature is processed by a third attention process to obtain the first target feature.

[0176] For example, this application uses the ConvNeXt model as the basic backbone network for illustration. The method for processing using other basic backbone networks is similar and will not be described further here. The basic backbone network includes multiple backbone network blocks, each of which has a corresponding attention mechanism added, such as... Figure 9 The diagram illustrates the first, second, and third attention mechanisms. The first attention process for the first fused feature is completed by inputting the first fused feature into N1 backbone network blocks (ConvNeXt Blocks) connected to the Coordinate attention mechanism, followed by two downsampling operations. For example, the dim value of each ConvNeXt Block can be 96. The first downsampling yields a 56×56×96 feature, and the second downsampling yields a 28×28×192 feature, which is the second feature.

[0177] It should be noted that the embodiments of this application are only described using the Coordinate attention mechanism as the first attention mechanism. The processing method for other attention mechanisms as the first attention mechanism is similar to this method and will not be described here.

[0178] Then, the second feature is input into N2 backbone network blocks connected to the ECANet attention mechanism, i.e., ConvNeXt Blocks, and downsampling is performed to complete the second attention processing of the second feature. Each ConvNeXt Block has a dim value of 96, and after downsampling, a 14×14×384 feature is obtained, which is the third feature.

[0179] It should be noted that the embodiments in this application are only illustrated using ECANet attention mechanism as the second attention mechanism. The processing method for other attention mechanisms as the second attention mechanism is similar to this method and will not be described here.

[0180] Next, the third feature is input into N3 ConvNeXt Blocks connected to the SimAM attention mechanism, and then subjected to full-episode average pooling to complete the third attention processing of the third feature. Each ConvNeXt Block has a dim value of 96, and after full-episode average pooling, a 7×7×768 feature is obtained. This 7×7×768 feature is then passed through layer normalization and a fully connected layer to obtain the first target feature.

[0181] It should be noted that the embodiments of this application are only illustrated using SimAM attention mechanism as the third attention mechanism. The processing method for other attention mechanisms as the third attention mechanism is similar to this method and will not be described here.

[0182] Therefore, by performing a first attention process on the first fused feature, a second feature can be obtained; by performing a second attention process on the second feature, a third feature can be obtained; and by performing a third attention process on the third feature, a first target feature can be obtained, thereby completing the feature extraction of the first fused feature and improving the accuracy of target detection.

[0183] Furthermore, the first target feature is input into the target prediction model, such as... Figure 9 As shown, the target prediction model can be, for example, a YOLO detection head. By detecting the first target features using the YOLO detection head, the first location information and first category of the target can be obtained. Thus, target detection based on target data is achieved, obtaining the first location information and first category of the target, thereby improving the accuracy of target detection.

[0184] For example, obtaining the first location information and first category of the target based on the first target features may include the following steps:

[0185] Based on the features of the first target, multiple first prediction boxes are obtained;

[0186] Obtain the confidence level of the target in each first prediction box;

[0187] The first predicted bounding box with a confidence level greater than the first threshold is used as the first target predicted bounding box;

[0188] Use the location information of the first target prediction box as the first location information of the target;

[0189] Obtain the probability of the first target prediction box for each category;

[0190] The category with the highest probability is used as the first category of the target.

[0191] In this embodiment, the first target feature is divided into multiple grids by segmenting it into blocks. Then, each grid extends outwards by an equal number of bounding boxes, resulting in multiple first predicted bounding boxes. Next, the grid containing the target center point and the grid containing the target contour are predicted, and the confidence level of each first predicted bounding box including the target, i.e., the confidence level of the grid containing the target contour, is predicted. First predicted bounding boxes with a confidence level greater than a first threshold are designated as first target predicted bounding boxes.

[0192] Further, based on the center point of the first target prediction bounding box and the predicted center point of the target, the center point of the first target prediction bounding box is adjusted to the predicted center point of the target to obtain the position information of the first target prediction bounding box. For example, the position information of the first target prediction bounding box can be (x, y, w, h), where (x, y) represents the coordinates of the center point of the first target prediction bounding box, w represents the width of the first target prediction bounding box, and h represents the height of the first target prediction bounding box. The position information of the first target prediction bounding box is used as the first position information of the target. Optionally, the first target prediction bounding box can be a three-dimensional cylindrical box to select the target.

[0193] Next, the first target prediction box is classified, and the probability of the first target prediction box in each category is predicted. The category with the highest probability is taken as the first category of the target.

[0194] Optionally, the first target prediction box, the first location information of the target, and the first category can be displayed in real time on the vehicle's display device to indicate target information on the user's driving route.

[0195] Optionally, the first target prediction box, the first location information and the first category of the target are synchronized to the vehicle's route planning control unit. The route planning control unit can plan a route for the vehicle to travel based on the first location information and the first category of the target, and control the vehicle's route travel.

[0196] Therefore, based on the first target features, multiple first detection boxes can be segmented from the first target features, and the confidence level of each first prediction box including the target can be predicted. The first prediction box with a confidence level greater than a first threshold is taken as the first target prediction box. Thus, the position information of the first target prediction box can be used as the first position information of the target. The first target prediction box is then classified, and the probability of the first target prediction box in each category is predicted. The category with the highest probability is taken as the first category of the target, thereby obtaining the first position information and the first category of the target, realizing target detection and improving the accuracy of target detection.

[0197] As can be seen, in this embodiment, target multimedia data and target radar data collected at multiple times can be acquired, fused, and then used for target detection. This avoids the problem of inaccurate target detection caused by complex environments when using multimedia data, thus improving the accuracy of target detection. Furthermore, it avoids the problem of poor object perception when using radar data, further improving the accuracy of target detection. By fusing target multimedia data and target radar data from multiple times, the target data can include more data features, further improving the accuracy of target detection.

[0198] See Figure 10 , Figure 10 This is a schematic diagram of a target detection device provided in an embodiment of this application. Figure 10 As shown, the target detection device 1000 includes an acquisition unit 1001 and a detection unit 1002;

[0199] Among them, the acquisition unit 1001 is used to acquire target multimedia data and target radar data at multiple times;

[0200] The detection unit 1002 is used to fuse target multimedia data and target radar data at multiple times to obtain target data;

[0201] Target detection based on target data.

[0202] In one feasible embodiment, the target multimedia data at multiple times includes the first multimedia data at the current time and the second multimedia data at multiple historical times, and the target radar data at multiple times includes the first radar data at the current time and the second radar data at multiple historical times.

[0203] In fusing target multimedia data and target radar data from multiple time points to obtain target data, the detection unit 1002 is specifically used for:

[0204] The first multimedia data and the first radar data at the current moment are fused to obtain the first fused data;

[0205] The second multimedia data and second radar data from multiple historical moments are fused to obtain multiple second fused data, wherein the second multimedia data and second radar data from each historical moment are fused accordingly.

[0206] The first fused data is fused with multiple second fused data to obtain the target data.

[0207] In a feasible embodiment, before fusing the second multimedia data and second radar data from multiple historical moments to obtain multiple second fused data, the detection unit 1002 is further configured to:

[0208] Multiple sets of first multimedia data at the current moment are used as second multimedia data at multiple historical moments; and multiple sets of first radar data at the current moment are used as second radar data at multiple historical moments.

[0209] In one feasible embodiment, in terms of target detection based on target data, the detection unit 1002 is specifically used for:

[0210] Upsample the target data to obtain the first feature;

[0211] The first feature is fused to obtain the first fused feature;

[0212] The first fused feature is used to extract features to obtain the first target feature;

[0213] Based on the first target features, the first location information and first category of the target are obtained.

[0214] In a feasible embodiment, in extracting features from the first fused features to obtain the first target features, the detection unit 1002 is specifically used for:

[0215] The first fusion feature is subjected to a first attention process to obtain the second feature;

[0216] The second feature is processed with a second attention process to obtain the third feature;

[0217] The third feature is processed by third attention to obtain the first target feature.

[0218] In a feasible embodiment, the detection unit 1002, in obtaining the first location information and first category of the target based on the first target features, is specifically used for:

[0219] Based on the features of the first target, multiple first prediction boxes are obtained;

[0220] Obtain the confidence level of the target in each first prediction box;

[0221] The first predicted bounding box with a confidence level greater than the first threshold is used as the first target predicted bounding box;

[0222] Use the location information of the first target prediction box as the first location information of the target;

[0223] Obtain the probability of the first target prediction box for each category;

[0224] The category with the highest probability is used as the first category of the target.

[0225] See Figure 11 , Figure 11 This is a schematic diagram of a target detection model training device provided in an embodiment of this application. The target detection model training device 1100 includes: an acquisition unit 1101 and a processing unit 1102;

[0226] Among them, the acquisition unit 1101 is used to acquire multiple original samples and multiple style samples;

[0227] Processing unit 1102 is used to perform style transfer on multiple original samples based on multiple style samples to obtain multiple transferred samples;

[0228] Multiple target samples are obtained based on multiple original samples and multiple transfer samples;

[0229] A target detection model is obtained by training the model based on multiple target samples.

[0230] In one feasible embodiment, each original sample is formed by fusing third multimedia data and third radar data at multiple time points, and each migrated sample is formed by fusing fourth multimedia data and fourth radar data at multiple time points.

[0231] In obtaining multiple target samples based on multiple original samples and multiple transfer samples, the processing unit 1102 is specifically used for:

[0232] Multiple original samples are fused with multiple transferred samples to obtain multiple synthetic samples;

[0233] Multiple original samples, multiple transfer samples, and multiple synthetic samples are used as multiple target samples.

[0234] In a feasible embodiment, before fusing multiple original samples with multiple migrated samples to obtain multiple synthetic samples, the processing unit 1102 is further configured to:

[0235] Multiple sets of third-party multimedia data at the current moment are used as third-party multimedia data at multiple moments of the target's original sample; and multiple sets of third-party radar data at the current moment are used as third-party radar data at multiple moments of the target's original sample.

[0236] In one feasible embodiment, each target sample includes second location information and a second category of the target;

[0237] In terms of training the model based on multiple target samples to obtain the target detection model, the processing unit 1102 is specifically used for:

[0238] Upsample each target sample to obtain the fourth feature corresponding to each target sample;

[0239] The fourth feature corresponding to each target sample is fused to obtain the second fused feature corresponding to each target sample.

[0240] Feature extraction is performed on the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample;

[0241] Based on the second target features corresponding to each target sample, predict the third location information and third category of the target in each target sample;

[0242] The first loss is obtained based on the second and third location information of the target in each target sample;

[0243] The second loss is derived based on the second and third categories of the target in each target sample;

[0244] The target detection model is obtained by training the model based on the first loss and the second loss.

[0245] In a feasible embodiment, feature extraction is performed on the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample. The processing unit 1102 is specifically used for:

[0246] A fourth attention process is applied to the second fusion feature corresponding to each target sample to obtain the fifth feature corresponding to each target sample;

[0247] Perform fifth attention processing on the fifth feature corresponding to each target sample to obtain the sixth feature corresponding to each target sample;

[0248] The sixth attention process is applied to the sixth feature corresponding to each target sample to obtain the second target feature corresponding to each target sample.

[0249] In a feasible embodiment, in predicting the third location information and third category of the target in each target sample based on the second target features corresponding to each target sample, the processing unit 1102 is specifically used for:

[0250] Based on the second target feature corresponding to each target sample, multiple second prediction boxes are obtained for each target sample;

[0251] The prediction for each second prediction box includes the confidence level of the target;

[0252] The second predicted box with a confidence level greater than the second threshold is used as the second target predicted box for each target sample.

[0253] The location information of the second target prediction box corresponding to each target sample is used as the third location information of the target in each target sample;

[0254] Predict the probability of each category for the second target bounding box corresponding to each target sample;

[0255] The category with the highest probability is used as the third category of the target in each target sample.

[0256] See Figure 12 , Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 12 As shown, the electronic device 1200 includes a transceiver 1201, a processor 1202, and a memory 1203. These are connected via a bus 1204. The memory 1203 stores computer programs and data, and can transfer data stored in the memory 1203 to the processor 1202. The electronic device 1200 can be either the target detection device 1000 or the target detection model training device 1100 described above.

[0257] When the electronic device 1200 is a target detection device 1000, the processor 1202 reads the computer program in the memory 1203 and performs the following operations:

[0258] Acquire target multimedia data and target radar data at multiple time points;

[0259] Target data is obtained by fusing target multimedia data and target radar data from multiple time points;

[0260] Target detection based on target data.

[0261] In one feasible embodiment, the target multimedia data at multiple times includes the first multimedia data at the current time and the second multimedia data at multiple historical times, and the target radar data at multiple times includes the first radar data at the current time and the second radar data at multiple historical times.

[0262] In fusing target multimedia data and target radar data from multiple time points to obtain target data, processor 1202 is specifically used to perform the following operations:

[0263] The first multimedia data and the first radar data at the current moment are fused to obtain the first fused data;

[0264] The second multimedia data and second radar data from multiple historical moments are fused to obtain multiple second fused data, wherein the second multimedia data and second radar data from each historical moment are fused accordingly.

[0265] The first fused data is fused with multiple second fused data to obtain the target data.

[0266] In a feasible embodiment, before fusing the second multimedia data and second radar data from multiple historical moments to obtain multiple second fused data, the processor 1202 is further configured to perform the following operations:

[0267] Multiple sets of first multimedia data at the current moment are used as second multimedia data at multiple historical moments; and multiple sets of first radar data at the current moment are used as second radar data at multiple historical moments.

[0268] In one feasible embodiment, in terms of target detection based on target data, the processor 1202 is specifically configured to perform the following operations:

[0269] Upsample the target data to obtain the first feature;

[0270] The first feature is fused to obtain the first fused feature;

[0271] The first fused feature is used to extract features to obtain the first target feature;

[0272] Based on the first target features, the first location information and first category of the target are obtained.

[0273] In a feasible embodiment, in extracting features from the first fused features to obtain the first target features, the processor 1202 is specifically configured to perform the following operations:

[0274] The first fusion feature is subjected to a first attention process to obtain the second feature;

[0275] The second feature is processed with a second attention process to obtain the third feature;

[0276] The third feature is processed by a third attention process to obtain the first target feature.

[0277] In a feasible embodiment, in obtaining the first location information and first category of the target based on the first target features, the processor 1202 is specifically configured to perform the following operations:

[0278] Based on the features of the first target, multiple first prediction boxes are obtained;

[0279] Obtain the confidence level of the target in each first prediction box;

[0280] The first predicted bounding box with a confidence level greater than the first threshold is used as the first target predicted bounding box;

[0281] The position information of the first target prediction box is used as the first position information of the target;

[0282] Obtain the probability of the first target prediction box for each category;

[0283] The category with the highest probability is used as the first category of the target.

[0284] When the electronic device 1200 is a target detection model training device 1100, the processor 1202 reads the computer program in the memory 1203 and performs the following operations:

[0285] Obtain multiple original samples and multiple style samples;

[0286] Based on multiple style samples, style transfer is performed on multiple original samples to obtain multiple transferred samples;

[0287] Multiple target samples are obtained based on multiple original samples and multiple transfer samples;

[0288] A target detection model is obtained by training the model based on multiple target samples.

[0289] In one feasible embodiment, each original sample is formed by fusing third multimedia data and third radar data at multiple time points, and each migrated sample is formed by fusing fourth multimedia data and fourth radar data at multiple time points.

[0290] In obtaining multiple target samples based on multiple original samples and multiple transfer samples, the processor 1202 is specifically used to perform the following operations:

[0291] Multiple original samples are fused with multiple transferred samples to obtain multiple synthetic samples;

[0292] Multiple original samples, multiple transfer samples, and multiple synthetic samples are used as multiple target samples.

[0293] In a feasible embodiment, before fusing multiple original samples with multiple migrated samples to obtain multiple synthetic samples, the processor 1202 is also configured to perform the following operations:

[0294] Multiple sets of third-party multimedia data at the current moment are used as third-party multimedia data at multiple moments of the target's original sample; and multiple sets of third-party radar data at the current moment are used as third-party radar data at multiple moments of the target's original sample.

[0295] In one feasible embodiment, each target sample includes second location information and a second category of the target;

[0296] In training a model based on multiple target samples to obtain a target detection model, the processor 1202 is specifically used to perform the following operations:

[0297] Upsample each target sample to obtain the fourth feature corresponding to each target sample;

[0298] The fourth feature corresponding to each target sample is fused to obtain the second fused feature corresponding to each target sample.

[0299] Feature extraction is performed on the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample;

[0300] Based on the second target features corresponding to each target sample, predict the third location information and third category of the target in each target sample;

[0301] The first loss is obtained based on the second and third location information of the target in each target sample;

[0302] The second loss is derived based on the second and third categories of the target in each target sample;

[0303] The target detection model is obtained by training the model based on the first loss and the second loss.

[0304] In a feasible embodiment, in extracting features from the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample, the processor 1202 is specifically configured to perform the following operations:

[0305] A fourth attention process is applied to the second fusion feature corresponding to each target sample to obtain the fifth feature corresponding to each target sample;

[0306] Perform fifth attention processing on the fifth feature corresponding to each target sample to obtain the sixth feature corresponding to each target sample;

[0307] The sixth attention process is applied to the sixth feature corresponding to each target sample to obtain the second target feature corresponding to each target sample.

[0308] In a feasible embodiment, in predicting the third location information and third category of the target in each target sample based on the second target features corresponding to each target sample, the processor 1202 is specifically configured to perform the following operations:

[0309] Based on the second target feature corresponding to each target sample, multiple second prediction boxes are obtained for each target sample;

[0310] The prediction for each second prediction box includes the confidence level of the target;

[0311] The second predicted box with a confidence level greater than the second threshold is used as the second target predicted box for each target sample.

[0312] The location information of the second target prediction box corresponding to each target sample is used as the third location information of the target in each target sample;

[0313] Predict the probability of each category for the second target bounding box corresponding to each target sample;

[0314] The category with the highest probability is used as the third category of the target in each target sample.

[0315] It should be understood that the electronic devices mentioned in this application may include smartphones (such as Android phones, iOS phones, Windows Phones, etc.), tablet computers, PDAs, laptops, mobile internet devices (MIDs), or wearable devices. The above-mentioned electronic devices are merely examples and not exhaustive, and include, but are not limited to, the electronic devices described above. In practical applications, the above-mentioned electronic devices may also include: intelligent in-vehicle terminals, computer equipment, etc.

[0316] This application also provides a computer-readable storage medium storing a computer program that is executed by a processor to implement some or all of the steps of any of the methods described in the above method embodiments.

[0317] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments.

[0318] This application embodiment also provides a vehicle, the vehicle including the above-described target detection device 1000 and the above-described target detection model training device 1100.

[0319] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0320] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0321] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0322] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0323] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software program module.

[0324] If the integrated unit is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0325] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0326] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A target detection method, characterized in that, include: Acquire target multimedia data and target radar data at multiple time points; The target multimedia data and target radar data at the multiple time points are fused to obtain the target data; Target detection is performed based on the target data.

2. The method according to claim 1, characterized in that, The target multimedia data at multiple times includes the first multimedia data at the current time and the second multimedia data at multiple historical times; the target radar data at multiple times includes the first radar data at the current time and the second radar data at multiple historical times. The process of fusing target multimedia data and target radar data at multiple time points to obtain target data includes: The first multimedia data and the first radar data at the current moment are fused to obtain the first fused data; The second multimedia data and second radar data from the multiple historical moments are fused to obtain multiple second fused data, wherein the second multimedia data and second radar data from each historical moment are fused accordingly. The first fused data is fused with the plurality of second fused data to obtain the target data.

3. The method according to claim 2, characterized in that, Before fusing the second multimedia data and second radar data from the multiple historical moments to obtain multiple second fused data, the method further includes: Multiple sets of first multimedia data at the current moment are used as second multimedia data at the multiple historical moments; and multiple sets of first radar data at the current moment are used as second radar data at the multiple historical moments.

4. The method according to any one of claims 1-3, characterized in that, The target detection based on the target data includes: The target data is upsampled to obtain the first feature; The first feature is fused to obtain the first fused feature; The first fused feature is used to extract features to obtain the first target feature; Based on the first target feature, the first location information and the first category of the target are obtained.

5. The method according to claim 4, characterized in that, The step of extracting features from the first fused features to obtain the first target features includes: The first fused feature is subjected to a first attention process to obtain a second feature; The second feature is subjected to a second attention process to obtain the third feature; The third feature is subjected to third attention processing to obtain the first target feature.

6. The method according to claim 4 or 5, characterized in that, The step of obtaining the first location information and first category of the target based on the first target feature includes: Based on the first target features, multiple first prediction boxes are obtained; Obtain the confidence level of the target in each first prediction box; The first predicted bounding box with a confidence level greater than the first threshold is used as the first target predicted bounding box; The position information of the first target prediction box is used as the first position information of the target; Obtain the probability of the first target prediction box for each category; The category with the highest probability is taken as the first category of the target.

7. A method for training an object detection model, characterized in that, include: Obtain multiple original samples and multiple style samples; Based on the multiple style samples, style transfer is performed on the multiple original samples to obtain multiple transferred samples; Based on the multiple original samples and the multiple transfer samples, multiple target samples are obtained; The target detection model is obtained by training the model based on the multiple target samples.

8. The method according to claim 7, characterized in that, Each original sample is formed by fusing third multimedia data and third radar data at multiple time points, and each migrated sample is formed by fusing fourth multimedia data and fourth radar data at multiple time points. Based on the multiple original samples and the multiple transfer samples, multiple target samples are obtained, including: The multiple original samples and the multiple migrated samples are fused to obtain multiple synthetic samples; The plurality of original samples, the plurality of migrated samples, and the plurality of synthetic samples are used as the plurality of target samples.

9. The method according to claim 8, characterized in that, Before fusing the plurality of original samples with the plurality of migrated samples to obtain a plurality of synthetic samples, the method further includes: Multiple sets of third multimedia data at the current moment are used as the third multimedia data at multiple moments of the target's original sample; and multiple sets of third radar data at the current moment are used as the third radar data at multiple moments of the target's original sample, wherein the target's original sample is any one of the multiple original samples.

10. The method according to any one of claims 7-9, characterized in that, Each target sample includes the target's second location information and second category; The process of training the model based on the multiple target samples to obtain the target detection model includes: Upsample each target sample to obtain the fourth feature corresponding to each target sample; The fourth feature corresponding to each target sample is fused to obtain the second fused feature corresponding to each target sample. Feature extraction is performed on the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample; Based on the second target features corresponding to each target sample, predict the third location information and third category of the target in each target sample; The first loss is obtained based on the second and third location information of the target in each target sample; The second loss is derived based on the second and third categories of the target in each target sample; The target detection model is obtained by training the model based on the first loss and the second loss.

11. The method according to claim 10, characterized in that, The step of extracting features from the second fusion feature corresponding to each target sample to obtain the second target feature corresponding to each target sample includes: A fourth attention process is applied to the second fusion feature corresponding to each target sample to obtain the fifth feature corresponding to each target sample; Perform fifth attention processing on the fifth feature corresponding to each target sample to obtain the sixth feature corresponding to each target sample; The sixth attention process is applied to the sixth feature corresponding to each target sample to obtain the second target feature corresponding to each target sample.

12. The method according to claim 10 or 11, characterized in that, The prediction of the third location information and third category of the target in each target sample based on the second target feature corresponding to each target sample includes: Based on the second target feature corresponding to each target sample, multiple second prediction boxes are obtained for each target sample; The prediction for each second prediction box includes the confidence level of the target; The second predicted box with a confidence level greater than the second threshold is used as the second target predicted box for each target sample. The location information of the second target prediction box corresponding to each target sample is used as the third location information of the target in each target sample; Predict the probability of each category for the second target bounding box corresponding to each target sample; The category with the highest probability is used as the third category of the target in each target sample.

13. A target detection device, characterized in that, The target detection device includes an acquisition unit and a detection unit; The acquisition unit is used to acquire target multimedia data and target radar data at multiple times. The detection unit is used to fuse the target multimedia data and target radar data at the multiple time points to obtain target data; Target detection is performed based on the target data.

14. A target detection model training device, characterized in that, The target detection model training device includes an acquisition unit and a processing unit; The acquisition unit is used to acquire multiple original samples and multiple style samples; The processing unit is used to perform style transfer on the multiple original samples based on the multiple style samples to obtain multiple transferred samples; Based on the multiple original samples and the multiple transfer samples, multiple target samples are obtained; The target detection model is obtained by training the model based on the multiple target samples.

15. An electronic device, characterized in that, include: A processor and a memory, the processor being connected to the memory, the memory being used to store a computer program, the processor being used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-12.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the method as described in any one of claims 1-12.

17. A vehicle, characterized in that, The vehicle includes the target detection device as described in claim 13 and the target detection model training device as described in claim 14.