Model training method, electronic device, computer readable medium and program product
By converting point cloud data into radar images and using mature image detection models to generate pseudo-labels, the problem of manual annotation increases costs is solved, and a model that can perform radar image target detection is trained under low-cost conditions.
Patent Information
- Application Number
- CN202510123369.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the existing model training methods, manual annotation method increases the cost of model training, especially in the training of radar image target detection models.
By acquiring point cloud data and shooting images, it is converted into radar images, and using mature image detection models to detect the captured images, generate pseudo-labels, map them into the radar images, and form a radar image with pseudo-labels, which is used to train the target detection model.
With low data labeling costs, a model that can detect radar image objects is trained, reducing the need for manual labeling and improving the accuracy of training data.
Smart Images

Figure CN120047772A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly relates to a model training method, an electronic device, a computer-readable medium, and a program product. Background Art
[0002] In model training, a large amount of training data is generally obtained by manual annotation, and then the model is trained using the large amount of training data to obtain a mature model. However, the manual annotation method will increase the cost of model training. Summary of the Invention
[0003] In view of this, this application provides a model training method, an electronic device, a computer-readable medium, and a program product, which can realize the training of a radar image target detection model on the basis of low data annotation cost.
[0004] A model training method according to an embodiment of the present invention, the method includes: obtaining multiple groups of point cloud data and captured images with corresponding relationships, and converting the point cloud data into radar images, where the point cloud data and the corresponding captured images are acquisition data of the same area at the same data acquisition moment; performing target detection on the captured images based on a first image detection model to obtain image detection results of multiple targets; mapping the image detection results of the multiple targets in the radar image corresponding to the captured images to generate pseudo-labels respectively corresponding to the multiple targets in the radar image, where the pseudo-labels are used to identify the corresponding targets; using the radar image with multiple pseudo-labels as a radar image with pseudo-labels; generating a first training set based on multiple radar images with pseudo-labels; and training a second image detection model based on the first training set to train and obtain a first model capable of directly performing target detection based on radar images.
[0005] In some embodiments, after training and obtaining the first model, the method further includes: generating a sample set based on a second training set and the first model, where the second training set includes multiple radar images; and training a target tracking model based on the sample set to train and obtain a second model capable of realizing target tracking based on the output of the first model; where the sample set includes: multiple groups of cross-frame positive sample pairs, and the positive sample pairs have passed multi-frame verification, where the positive sample pairs correspond to feature pairs of the same target.
[0006] In some embodiments, the method for determining that the first positive sample pair in the multiple groups of cross-frame positive sample pairs passes the multi-frame verification is as follows: Input the multiple radar images in the second training set into the first model respectively to obtain the features of all targets in the multiple radar images; wherein, the features of all targets in the multiple radar images include a first feature, a second feature, and a third feature, where the first feature is the feature of the target in the radar image generated based on the first point cloud data, the second feature is the feature of the target in the radar image generated based on the second point cloud data, the third feature is the feature of the target in the radar image generated based on the third point cloud data, and the data acquisition times of the first point cloud data, the second point cloud data, and the third point cloud data are adjacent; Determine that the first feature and the third feature are successfully matched based on the bipartite graph matching algorithm, and determine the first feature and the third feature as the first positive sample pair; Determine that the first feature and the second feature are successfully matched based on the bipartite graph matching algorithm, and, the second feature and the third feature are successfully matched, and determine that the first positive sample pair passes the multi-frame verification.
[0007] In some embodiments, the sample set further includes: multiple groups of negative sample pairs, and the negative sample pairs correspond to feature pairs of different targets; The step of training the target tracking model based on the sample set to obtain the second model includes: training a preset multi-layer perceptron network in a contrastive learning manner based on the multiple groups of cross-frame positive sample pairs, the multiple groups of negative sample pairs, and a first loss function, so that the multi-layer perceptron network can output a second distance and a third distance, where the second distance corresponds to the distance between the features of the positive sample pair, and the third distance corresponds to the distance between the features of the negative sample pair; Based on the training result of the multi-layer perceptron network, obtain the second model.
[0008] In some embodiments, the method for determining the first negative sample pair in the multiple groups of negative sample pairs is as follows: Input the multiple radar images in the second training set into the first model respectively to obtain the features of all targets in the multiple radar images; Determine the fourth feature and the fifth feature in the same frame of radar image among the features of all targets in the multiple radar images; Determine the fourth feature and the fifth feature as the first negative sample pair.
[0009] In some embodiments, based on the multiple sets of cross-frame positive sample pairs, the multiple sets of negative sample pairs, and the first loss function, a preset multi-layer perceptron network is trained by contrastive learning, including: generating multiple positive sample differences based on the multiple sets of cross-frame positive sample pairs, and generating multiple negative sample differences based on the multiple sets of negative sample pairs; using the multiple positive sample differences and the multiple negative sample differences as the input for the first round of training of the multi-layer perceptron network to obtain multiple positive sample distances corresponding to the multiple positive sample differences output by the multi-layer perceptron network, and multiple negative sample distances corresponding to the multiple negative sample differences; determining the loss value of the first loss function based on the multiple positive sample distances and the multiple negative sample distances; in the case where the loss value of the first loss function does not meet the training termination condition, adjusting the network parameters of the multi-layer perceptron network to perform the next round of training; in the case where the loss value of the first loss function meets the training termination condition, using the trained multi-layer perceptron network as the second model.
[0010] In some embodiments, after obtaining the second model, the method further includes: setting a distance threshold corresponding to the second model based on the second distance and the third distance, where the distance threshold is greater than the second distance and less than the third distance.
[0011] In some embodiments, the first loss function includes triplet loss: triplet loss L = max(D P - D N + α, 0); where D P is the multiple positive sample distances output by the multi-layer perceptron network; D N is the multiple negative sample distances output by the multi-layer perceptron network; α is a margin parameter used to represent the minimum distance between the first sample distance and the second sample distance, where the first sample distance is any one of the multiple negative sample distances, and the second sample distance is any one of the positive sample distances; L is the loss value of the first loss function.
[0012] In some embodiments, after training the first model and the second model, the method further includes: obtaining fourth point cloud data collected by the radar at a first moment; obtaining fifth point cloud data collected by the radar at a second moment; generating a first radar image based on the fourth point cloud data, and generating a second radar image based on the fifth point cloud data; inputting the first radar image into the first model to obtain the features of each of the m targets in the first radar image; and inputting the second radar image into the first model to obtain the features of each of the n targets in the second radar image; determining q tracking targets based on the similarity between the features of each of the m targets and the features of each of the n targets, so as to implement target tracking of the q tracking targets from the first moment to the second moment.
[0013] In some embodiments, the determining q tracking targets based on the similarity between the features of each of the m targets and the features of each of the n targets includes: determining that the similarity between the feature of a first target among the m targets and the feature of a second target among the n targets satisfies a preset condition based on the second model, and determining the first target or the second target as a third target among the q tracking targets.
[0014] In some embodiments, the determining that the similarity between the feature of a first target among the m targets and the feature of a second target among the n targets satisfies a preset condition based on the second model includes: inputting the difference between the feature of the first target and the feature of the second target into the second model; obtaining a first distance between the feature of the first target and the feature of the second target output by the second model; and determining that the similarity between the feature of the first target and the feature of the second target satisfies the preset condition based on the first distance being less than a distance threshold corresponding to the second model.
[0015] An electronic device according to an embodiment of the present invention includes: one or more processors; one or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the model training method according to an embodiment of the present invention.
[0016] A computer-readable medium according to an embodiment of the present invention, wherein instructions are stored on the readable medium, and when the instructions are executed on a computer, the computer executes the model training method according to an embodiment of the present invention.
[0017] A computer program product according to an embodiment of the present invention, characterized in that it includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the model training method according to the embodiment of the present invention is implemented.
[0018] Advantages of the present invention:
[0019] In an embodiment of the present invention, the point cloud data of the radar is converted into a radar image, and a mature image detection model (basically no need to be retrained) is used to perform target detection on the captured image (corresponding to the point cloud data), and then the detected target is mapped to the radar image through a certain mapping relationship, thereby completing the annotation of the radar image. Then, based on the annotated radar image, a general image detection model is trained, and a model dedicated to radar image target detection can be obtained. In this way, the annotation of the radar image is realized automatically, significantly reducing the requirements for manual annotation. Therefore, on the basis of low data annotation cost, the training of the radar image target detection model can be completed. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 The figure shows a schematic flow chart of multi-target tracking at two adjacent data acquisition times provided by an embodiment of the present application.
[0021] Figure 2 The figure shows a schematic flow chart of multi-target tracking over a period of time provided by an embodiment of the present application.
[0022] Figure 3 The figure shows a schematic diagram of a radar image provided by an embodiment of the present application.
[0023] Figure 4 The figure shows a schematic process diagram of inputting the z-th input data into an unsupervised tracking model provided by an embodiment of the present application.
[0024] Figure 5a The figure shows a schematic diagram of image z and image z + 1 provided by an embodiment of the present application.
[0025] Figure 5b The figure shows a schematic diagram of the matching result output by an unsupervised tracking model provided by an embodiment of the present application.
[0026] Figure 6 The figure shows a schematic flow chart of training a first model provided by an embodiment of the present application.
[0027] Figure 7 The figure shows a schematic flow chart of training a radar image detection model provided by an embodiment of the present application.
[0028] Figure 8 The figure shows a schematic process diagram of generating a radar image with pseudo labels provided by an embodiment of the present application.
[0029] Figure 9 The figure shows a schematic flow chart of training an unsupervised tracking model provided by an embodiment of the present application.
[0030] Figure 10a The figure shows a schematic process diagram of constructing a negative sample set provided by an embodiment of the present application.
[0031] Figure 10b The figure shows a schematic process diagram of multi-frame verification of positive sample pairs provided by an embodiment of the present application.
[0032] Figure 11 The figure shows a schematic diagram of sample distance provided by an embodiment of the present application.
[0033] Figure 12 The figure shows a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0034] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification and specific implementation manners.
[0035] For ease of understanding, the terms related to the embodiments of the present application will be explained first.
[0036] (1) Millimeter-wave radar: A radar system operating in the millimeter-wave frequency band. Millimeter waves include electromagnetic waves with wavelengths between 1 millimeter and 10 millimeters, and the corresponding frequency band of millimeter waves can be 30 GHz to 300 GHz. In this frequency band, a millimeter-wave radar can detect the position, speed, and other characteristics of an object by transmitting and receiving electromagnetic waves. Here, a millimeter-wave radar capable of detecting the distance, direction, height, and speed characteristics of an object can be called a 4D millimeter-wave radar. All millimeter-wave radars involved in the embodiments of the present application can be 4D millimeter-wave radars. Without specific distinction hereinafter, they are all collectively referred to as "millimeter-wave radar".
[0037] (2) Point cloud data: Spatial information in the detection area detected by a millimeter-wave radar. Point cloud data consists of a large number of points, and each point represents a target in the detection area. Based on each point in the point cloud data, characteristics such as the position, speed, and reflection intensity of the target corresponding to each point can be obtained. Among them, the reflection intensity of the target can be used to represent the ability of the target to reflect the electromagnetic waves emitted by the millimeter-wave radar. The smoother the surface of the target, the larger the volume, and the greater the reflection intensity.
[0038] Here, the point cloud data can be the point cloud data collected by a millimeter-wave radar, or the point cloud data collected by other sensors such as lidar. This application does not make restrictive descriptions about the sensors for collecting point cloud data. Hereinafter, taking the sensor for collecting point cloud data as a millimeter-wave radar as an example for illustration, and the millimeter-wave radar for collecting point cloud data can be a 4D millimeter-wave radar. In the following, no distinction is made between the millimeter-wave radar and the 4D millimeter-wave radar, and they are collectively referred to as "millimeter-wave radar".
[0039] (3) Target tracking: A computer vision task whose purpose is to track the motion trajectory of a target simultaneously. The main processes of target tracking include target detection and target tracking. Taking the target tracking implemented based on a millimeter-wave radar as an example, target detection includes detecting targets in the detection area in the point cloud data collected at one time; target tracking includes matching the targets in the point cloud data collected continuously twice. In this way, target tracking can determine the motion trajectory of the target in the continuously collected point cloud data. When there are multiple targets in the detection area, target tracking can determine the motion trajectories of multiple targets in the continuously collected point cloud data. At this time, target tracking can be called multi-target tracking.
[0040] Based on the foregoing content, it can be seen that the method of manually annotating point cloud data will result in a relatively high cost in the model training process.
[0041] To solve the above technical problems, the embodiments of this application provide a model training method. Based on this model training method, an electronic device can, based on a data acquisition system with joint calibration of a millimeter-wave radar and a camera, while collecting the point cloud data of the detection area, synchronously collect the captured images of the detection area, and perform image detection on multiple targets in the captured images based on a mature image detection model to obtain the image detection results of the multiple targets. Furthermore, the electronic device can map the image detection results of the multiple targets onto the radar image generated based on the point cloud data to obtain a radar image with pseudo-labels (hereinafter referred to as a radar image with pseudo-labels). Furthermore, the electronic device can train the image detection model through the radar image with pseudo-labels to train the image detection model into a radar image detection model capable of detecting each target in the radar image.
[0042] It can be understood that based on the model training method provided by the embodiments of this application, during the process of generating the radar image with pseudo-labels, manual annotation is not required, effectively reducing the labor cost and resource cost required for generating the radar image with pseudo-labels.
[0043] Further, the electronic device can construct positive sample pairs and negative sample pairs of each target based on the trained radar image detection model. Among them, the positive sample pair of the target is used to determine different features of the target in radar images collected at different times, and the negative sample pair of the target is used to determine different features between the target and other targets in the radar image. Furthermore, the electronic device can, based on the contrast learning method, train a multi-layer perceptron network to learn the similarity between different features of the same target in different radar images, and the difference in features of different targets in the radar image, so as to train the multi-layer perceptron network into an unsupervised tracking model that can accurately match the same target in different radar images. For example, for multiple radar images corresponding to continuously collected point cloud data, the trained unsupervised tracking model can accurately match the targets in two adjacent radar images in the time order of collecting the point cloud data. Therefore, the trained unsupervised tracking model can continuously track each target in the detection area according to the continuously collected multiple point cloud data.
[0044] Among them, the electronic device for training the radar image detection model and the unsupervised tracking model can be any electronic device capable of performing deep learning model training, including but not limited to laptop computers, desktop computers, tablet computers, servers, etc., which are not limited here. The electronic devices using the trained radar image detection model and the unsupervised tracking model for target tracking include but are not limited to in-vehicle devices, drones, smart home devices, robots, security monitoring systems, traffic management systems, etc., which are electronic devices capable of obtaining point cloud data. Hereinafter, no specific distinction will be made between the electronic devices using the radar image detection model and the unsupervised tracking model, and the electronic devices for training the radar image detection model and the unsupervised tracking model. Hereinafter, the electronic device 10 will be used as an example for illustration.
[0045] Next, the process of the electronic device 10 training the radar image detection model and the unsupervised tracking model will be described.
[0046] Specifically, the electronic device 10 can first train the radar image detection model so that the radar image detection model can perform target detection on radar images.
[0047] Specifically, in the process of training the object detection ability of the radar image detection model, the electronic device 10 can project the point cloud data of the detection area collected by the millimeter-wave radar to the bird's eye view (BEV) to generate a radar image of the detection area. Moreover, the electronic device 10 can synchronously collect the captured images of the detection area based on the jointly calibrated data acquisition system. Furthermore, the electronic device 10 can map the image detection results of multiple objects in the captured images onto the radar image to generate a radar image with pseudo-labels. Therefore, the electronic device 10 can train a mature image detection model based on the radar image with pseudo-labels to train the image detection model into a radar image detection model. The radar image detection model learns the features of each object in the radar image. Furthermore, it can accurately detect each object in the radar image based on the learned features.
[0048] Obviously, in the process of training the radar image detection model, there is no need for manual annotation, effectively reducing the labor cost and resource cost required to generate the radar image with pseudo-labels.
[0049] In addition, compared with the method of directly manually annotating point cloud data, since the features exhibited by objects of different categories in the point cloud data may be similar. For example, the positions are similar, the speeds are similar, the reflection intensities are similar, etc. Therefore, in the process of manually annotating based on the features of objects of different categories in the point cloud data, the category of the object may be mislabeled when the features of objects of different categories are similar, resulting in low accuracy of the training data of the model. Obviously, inaccurate training data may lead to incorrect object detection by the trained model. Furthermore, the object tracking based on the object detection result will also be incorrect.
[0050] However, based on the embodiments of the present application, the training data of the radar image detection model is a radar image with pseudo-labels, and the pseudo-labels are mapped based on the object detection results of the image detection model for the captured images. Since the image detection model based on the captured images is already relatively mature and can achieve accurate object detection, the mapped pseudo-labels also have relatively high accuracy. This method ensures the accuracy of the training data of the radar image detection model, enabling the radar image detection model to achieve accurate object detection and avoiding object tracking errors caused by incorrect object detection.
[0051] Furthermore, the electronic device 10 can train an unsupervised tracking model based on the trained radar image detection model, so that the unsupervised tracking model can track each object in the detection area according to the results output by the radar image detection model.
[0052] Specifically, during the process of training the object detection ability of the unsupervised tracking model, the electronic device 10 can input multiple consecutive radar images into the trained radar image detection model to obtain the features of all objects in the multiple radar images. Furthermore, the electronic device 10 can determine multiple positive sample pairs based on the bipartite graph matching algorithm among the features of all objects, where a positive sample pair is a feature pair composed of two features of the same object. The electronic device 10 can determine multiple negative sample pairs among the features of all objects, where a negative sample pair is a feature pair composed of the features of two different objects. Furthermore, the electronic device 10 can generate a positive sample set containing multiple positive sample pairs and a negative sample set containing multiple negative sample pairs based on the multiple positive sample pairs and multiple negative sample pairs.
[0053] Furthermore, the electronic device 10 can train the unsupervised tracking model by means of contrastive learning (a way to achieve unsupervised learning), learn the similarity between different features of the same object based on the positive sample set, and learn the difference between different features of different objects based on the negative sample set. Therefore, the unsupervised tracking model can accurately match the same object in different radar images based on the learned feature similarity of the same object and the feature difference of different objects. Equivalently, the unsupervised tracking model has the object tracking ability. It can be understood that during the process of training the unsupervised tracking model by the electronic device 10, there is no need for manual annotation of training data, effectively reducing the time cost and resource cost required for training the unsupervised tracking model.
[0054] Based on the above process, the electronic device 10 completes the training of the radar image detection model and the unsupervised tracking model. Therefore, the electronic device 10 can integrate the trained radar image detection model and the unsupervised tracking model to achieve the tracking of each object in the detection area during the application process.
[0055] Specifically, after the electronic device 10 obtains the point cloud data at multiple consecutive moments collected by the millimeter-wave radar, it can convert the point cloud data at multiple consecutive moments into multiple consecutive frames of radar images. The electronic device 10 can input each frame of the radar image into the radar image detection model to obtain the features of each object in each frame of the radar image. Furthermore, the electronic device 10 can input the features of each object in two adjacent frames of radar images into the unsupervised tracking model to obtain the object matching result of the two adjacent frames of radar images, that is, match the same object in the two adjacent frames of radar images. Here, matching the same object in two adjacent frames of radar images is equivalent to completing the object tracking of each object in the detection area at two adjacent moments.
[0056] Next, in conjunction with the accompanying drawings and specific embodiments, the model training method provided by the embodiments of the present application will be described in detail. It should also be noted that the numbering of the steps in the method and process in the embodiments of the present application is for the convenience of reference, rather than limiting the order. If there is an order between the steps, it shall be subject to the written description.
[0057] First, the process of multi-target tracking of the electronic device 10 based on the trained radar image detection model and the unsupervised tracking model will be described in conjunction with the accompanying drawings.
[0058] Exemplarily, Figure 1 Taking two adjacent data acquisition moments within a period of time as an example, a schematic flow diagram of the electronic device 10 performing multi-target tracking at two adjacent data acquisition moments is shown. It can be understood that the continuous multi-target tracking of the electronic device 10 within a period of time is equivalent to successively executing the following Figure 1 shown process for each adjacent data acquisition moment within a period of time.
[0059] Here, Figure 1 The execution subject of each step in the shown process is the electronic device 10, and the execution subject of each step will not be repeatedly described below.
[0060] Referring to Figure 1 , the process of the electronic device 10 achieving multi-target tracking at two adjacent data acquisition moments may include:
[0061] S100: Obtain the point cloud data collected by the radar at the first moment.
[0062] Exemplarily, the electronic device 10 may obtain the point cloud data of the detection area collected by the radar at the first moment (i.e., the aforementioned fourth point cloud data). Since the detection area of the radar includes m targets at the first moment, the point cloud data at the first moment includes m targets. Here, the radar includes a millimeter-wave radar.
[0063] S101: Obtain the point cloud data collected by the radar at the second moment.
[0064] Exemplarily, the electronic device 10 may obtain the point cloud data of the detection area collected by the radar at the second moment (i.e., the aforementioned fifth point cloud data). Since the detection area of the radar includes n targets at the second moment, the point cloud data at the second moment includes n targets. It should be clear that the detection area at the first moment and the detection area at the second moment may be all or partially overlapping areas, or completely different areas.
[0065] Here, the first moment and the second moment may be any two adjacent data acquisition moments during the process of the millimeter-wave radar collecting point cloud data.
[0066] S102: Generate a first radar image based on the point cloud data at the first moment, and generate a second radar image based on the point cloud data at the second moment.
[0067] Exemplarily, the electronic device 10 may project the point cloud data at the first moment onto a bird's-eye view to obtain a first radar image, and project the point cloud data at the second moment onto a bird's-eye view to obtain a second radar image.
[0068] S103: Input the first radar image into a first model to obtain the features of each of the m targets in the first radar image.
[0069] Exemplarily, the electronic device 10 may input the first radar image into the trained first model, and the first model is used to detect multiple targets in the radar image and output the features of each target. Therefore, after the electronic device 10 inputs the first radar image into the first model, it can obtain the features of each of the m targets output by the first model.
[0070] Here, the first model includes the aforementioned radar image detection model.
[0071] S104: Input the second radar image into the first model to obtain the features of each of the n targets in the second radar image.
[0072] Exemplarily, the electronic device 10 may input the second radar image into the trained first model to obtain the features of each of the n targets output by the first model.
[0073] Here, the first model is used to detect multiple targets in the radar image and output the features of each target, and the first model is obtained by training an image detection model using radar images.
[0074] Specifically, the training data of the first model is a radar image with pseudo-labels (hereinafter referred to as "radar image with pseudo-labels") obtained based on the joint calibration of a millimeter-wave radar and a camera, and the first model is obtained based on an image detection model that has been trained and matured using the radar image with pseudo-labels. Therefore, the target detection ability of the first model for radar images is obtained by transferring the target detection ability of the image detection model for captured images. Therefore, since the image detection model can achieve accurate target detection, the radar image detection model can also achieve accurate detection of targets in radar images, avoiding target tracking errors caused by incorrect target detection.
[0075] Here, for the sake of narrative coherence, the generation process of the radar image with pseudo-labels and the training process of the first model based on the radar image with pseudo-labels will be described in detail in the following Figure 7 and will not be elaborated here.
[0076] S105: Determine q tracking targets based on the similarity between the features of each target among the m targets and the features of each target among the n targets, so as to achieve target tracking of the q tracking targets from the first moment to the second moment.
[0077] Herein, the first radar image and the second radar image are used to characterize the spatial information of the detection area at two adjacent data acquisition moments. The spatial information of the detection area changes little between two adjacent data acquisition moments. Therefore, there may be q identical targets among the m targets in the first radar image and the n targets in the second radar image. In this scenario, the electronic device 10 performs multi-target tracking on the detection area from the first moment to the second moment, which is equivalent to matching the q identical targets in the first radar image and the second radar image. Herein, the identical targets can be referred to as "tracking targets".
[0078] Exemplarily, the electronic device 10 can match the features of each target among the m targets with the features of each target among the n targets, and determine q tracking targets based on the matching results between the features.
[0079] Specifically, the electronic device 10 can match the features of the m targets with the features of each target among the n targets according to the second model obtained through training. Herein, the second model is used to determine the similarity between the features of the targets. The second model is obtained by training a multi-layer perceptron network to distinguish the features of identical targets and the features of different targets through contrastive learning. Specifically, the second model can output the distance between two features according to the difference between the two features, and this distance is used to characterize the similarity between the two features.
[0080] Taking the process of matching one tracking target as an example, the electronic device 10 can determine the difference between the features of the first target among the m targets and the features of the second target among the n targets, and input this difference into the second model. The second model can output the distance between the features of the first target and the features of the second target based on the input difference. Furthermore, the electronic device 10 can determine whether the first target and the second target are the same tracking target based on this distance.
[0081] For example, if the distance between the features of the first target and the features of the second target is less than a preset distance threshold, the electronic device 10 can determine that the features of the first target and the features of the second target are highly similar. Furthermore, it can be determined that the first target and the second target are the same tracking target.
[0082] It can be understood that for any combination of any one target among the m targets and any one target among the n targets, the above matching process is performed, and then the q tracking targets in the first radar image and the second radar image can be matched to complete multi-target tracking from the first moment to the second moment.
[0083] Here, the second model includes the aforementioned unsupervised tracking model. For the sake of narrative coherence, the training process of the unsupervised tracking model will be described in detail below Figure 9 .
[0084] Exemplarily, Figure 2 shows a schematic flow diagram of multi-object tracking by the electronic device 10 over a period of time according to the radar image detection model and the unsupervised tracking model.
[0085] Here, Figure 2 in the shown process, the execution subject of each step is the electronic device 10, and the execution subject of each step will not be repeatedly described below.
[0086] Referring to Figure 2 , the process of the electronic device 10 performing multi-object tracking over a period of time may include:
[0087] S200: Obtain multiple consecutive frames of point cloud data continuously collected by the millimeter-wave radar.
[0088] Exemplarily, during the process of multi-object tracking, the electronic device 10 may obtain the point cloud data of the detection area collected by the millimeter-wave radar at consecutive data collection moments.
[0089] Taking the automatic driving scenario of a vehicle as an example, a millimeter-wave radar may be set in the vehicle, and the electronic device 10 may be the in-vehicle device in the vehicle. Furthermore, during the automatic driving process of the vehicle, the millimeter-wave radar may collect the point cloud data of the detection area at consecutive data collection moments according to a preset collection frequency. Furthermore, the electronic device 10 may obtain in real time the consecutive multiple frames of point cloud data collected by the millimeter-wave radar at consecutive data collection moments, so as to perform multi-object tracking based on the continuously obtained consecutive multiple frames of point cloud data.
[0090] S201: Convert the consecutive multiple frames of point cloud data into consecutive multiple frames of radar images.
[0091] Exemplarily, the electronic device 10 may project the obtained consecutive multiple frames of point cloud data onto a bird's-eye view to obtain consecutive multiple frames of radar images, and each frame of radar image corresponds to the spatial information (i.e., point cloud data) of the detection area at each data collection moment.
[0092] Exemplarily, Figure 3 shows a schematic diagram of a radar image.
[0093] Referring to Figure 3, taking the autonomous driving scenario of a vehicle as an example, after the electronic device 10 obtains the point cloud data in real time, the obtained point cloud data can be projected onto a bird's-eye view to obtain a radar image 300 corresponding to the point cloud data. Among them, the X-axis of the radar image 300 points to the forward direction of the millimeter-wave radar, and the Y-axis points to the horizontal direction.
[0094] S202: Input consecutive multiple frames of radar images into a radar image detection model to obtain the features of each target in each frame of radar image.
[0095] Exemplarily, the electronic device 10 can input each frame of the obtained radar images into the radar image detection model one by one, and perform target detection on each frame of radar image based on the radar image detection model to obtain the features of each target in each frame of radar image.
[0096] Taking consecutive multiple data acquisition times as time 1 to time n, and consecutive multiple frames of radar images including image 1 to image n as an example, after the electronic device 10 inputs image 1 to image n into the radar image detection model one by one, it can respectively obtain to corresponding to ……, until to corresponding to
[0097] Among them, represents the i-th target in the j-th frame of radar image, represents the feature of the i-th target in the j-th frame of radar image, Q is the maximum number of targets in the radar image, which is equivalent to the maximum target tracking number of the unsupervised tracking model. Among them, 1 ≤ j ≤ n, 1 ≤ i ≤ Q.
[0098] It can be understood that if the number of targets in the j-th radar image is r, and r is less than Q, then to can be set to 0, correspondingly, to can be set to 0.
[0099] S203: Generate multiple consecutive input data based on the features of each target in consecutive multiple frames of radar images, where one input data corresponds to two adjacent frames of radar images.
[0100] Exemplarily, the electronic device 10 may start from the first frame of radar image and sequentially acquire two adjacent frames of radar images. For the two adjacent frames of radar images acquired, the electronic device 10 may perform a difference operation on the features of each target in the previous frame of radar image and the features of each target in the subsequent frame of radar image to obtain a plurality of feature differences. The electronic device 10 may use the obtained plurality of feature differences as an input data.
[0101] For example, for adjacent images z and z + 1 among images 1 to n, the electronic device 10 may obtain the to corresponding to in image z and, to corresponding to in image z + 1 based on the foregoing S202, where 1 ≤ z < n, image z corresponds to time z, and image z + 1 corresponds to time z + 1.
[0102] The electronic device 10 may perform a difference operation on the to features in to to obtain Q 2 feature differences, including:
[0103] Furthermore, the electronic device 10 may use the obtained Q 2 feature differences as the input data corresponding to images z and z + 1.
[0104] It can be understood that the electronic device 10 performs operations substantially the same as those for images z and z + 1 on all adjacent two - frame radar images among images 1 to n, and can obtain n - 1 consecutive input data. Among them, the input data corresponding to images z and z + 1 is the z - th input data among the n - 1 consecutive input data.
[0105] S204: Input the plurality of consecutive input data into the unsupervised tracking model one by one to obtain the target matching results of the adjacent two - frame radar images corresponding to each input data, so as to implement multi - target tracking.
[0106] Here, the first input training set corresponds to the first frame of radar image and the second frame of radar image. Therefore, when the electronic device 10 inputs the first input training set into the unsupervised tracking model, it can obtain the target matching results of the first frame of radar image and the second frame of radar image, which is equivalent to implementing multi - target tracking of the detection area from the first data acquisition time to the second data acquisition time.
[0107] Similarly, the second input training set corresponds to the second frame of radar image and the third frame of radar image. Therefore, when the electronic device 10 inputs the second input training set into the unsupervised tracking model, the target matching result of the second frame of radar image and the third frame of radar image can be obtained, which is equivalent to realizing multi-target tracking of the detection area from the second acquisition moment to the third acquisition moment.
[0108] By analogy, when the electronic device 10 inputs multiple consecutive input data into the unsupervised tracking model one by one, it is equivalent to realizing multi-target tracking of the detection area from the first acquisition moment to the last acquisition moment.
[0109] Exemplarily, Figure 4 Taking the z-th input data among n - 1 adjacent input data as an example, a schematic diagram of the process of inputting the z-th input data into the unsupervised tracking model is shown.
[0110] Referring to Figure 4 , based on the content in S203 above, it can be known that the z-th input data 400 includes:
[0111] Q 2 feature differences including. Therefore, after the electronic device 10 inputs the Q 2 feature differences in the z-th input data 400 into the unsupervised tracking model, the unsupervised tracking model can determine the sample distances corresponding to each of the Q 2 feature differences. Here, the sample distance corresponding to the feature difference refers to the distance between the features of two targets output by the unsupervised tracking model.
[0112] For example, the unsupervised tracking model can output a distance set 401 corresponding to the z-th input data 400, and the distance set 401 includes the sample distances output by the unsupervised tracking model corresponding to corresponding to corresponding to corresponding to corresponding to corresponding to corresponding to corresponding to corresponding to
[0113] Among them, represents the sample distance corresponding to the feature difference between the feature of the c-th target in the a-th frame of radar image and the feature of the d-th target in the b-th frame of radar image.
[0114] Furthermore, the electronic device 10 can post-process the Q 2 sample distances in the input data 400 according to a preset distance threshold. Specifically, for the sample distances in Q 2 that are less than the distance threshold, the matching result corresponding to the sample distance can be output as 1, that is, the features of the two targets corresponding to the sample distance are highly similar, indicating that the two targets are the same tracking target. For the sample distances in Q 2 that are greater than or equal to the distance threshold, the matching result corresponding to the sample distance can be output as 0, that is, the features of the two targets corresponding to the sample distance are not similar, indicating that the two targets are not the same tracking target.
[0115] Furthermore, the electronic device 10 can determine the two targets corresponding to each feature difference with a similarity of 1 as the same tracking target in the image z and the image z + 1. At this time, the electronic device 10 completes the target matching of the image z and the image z + 1 based on the output result of the unsupervised tracking model. Equivalently, the target tracking of the detection area from the z-th acquisition moment to the (z + 1)-th acquisition moment is completed.
[0116] It can be understood that the electronic device 10 sequentially executes the same process as described above for each of the multiple consecutive input data, and can complete the multi-target tracking of the detection area from the 1st acquisition moment to the nth acquisition moment. Figure 8
[0117] Next, the process of the electronic device 10 matching two radar images according to the radar image detection model and the unsupervised tracking model will be described in detail with reference to specific two frames of radar images.
[0118] Exemplarily, Figure 5a a schematic diagram of an image z and an image z + 1 is shown.
[0119] Referring to Figure 5a , the image z can be the radar image 500 in the figure, and the image z + 1 can be the radar image 501 in the figure. The radar image 500 includes 3 targets, and the radar image 501 also includes 3 targets. And, and are the same target, and are the same target, and are the same target.
[0120] Exemplarily, Figure 5b taking Figure 5a the scenario shown as an example, a schematic diagram of the matching result output by an unsupervised tracking model is shown.
[0121] Referring toFigure 5b The electronic device 10 can input the input data 502 corresponding to the radar image 500 and the radar image 501 (only 9 non-zero feature differences in the input data 502 are shown in the figure) into the unsupervised tracking model. Here, the non-zero feature differences of the input data 502 include:
[0122] The unsupervised tracking model can output a distance set 503 (only 9 non-zero sample distances in the distance set 503 are shown in the figure) according to the input input data 502. Here, the non-zero sample distances in the distance set 503 include:
[0123] Furthermore, the electronic device 10 can post-process each sample distance in the distance set 503 based on a distance threshold. Specifically, in the distance set 503 is greater than or equal to the distance threshold. Therefore, the electronic device 10 can post-process the corresponding matching result to 0. In the distance set 503 is less than the distance threshold. Therefore, the electronic device 10 can post-process the corresponding matching result to 1.
[0124] Therefore, the electronic device 10 can determine that and are the same target (for example, target A), and are the same target (for example, target B), and are the same target (for example, target C).
[0125] Equivalently, the electronic device 10 completes multi-target tracking from the t-th data acquisition moment to the (t + 1)-th data acquisition moment based on the radar image detection model and the unsupervised tracking model. Specifically, it tracks that target A moves from the position corresponding to to the position corresponding to , tracks that target B moves from the position corresponding to to the position corresponding to , and tracks that target C moves from the position corresponding to to the position corresponding to .
[0126] Next, the process of the electronic device 10 training the radar image detection model and the unsupervised tracking model will be described in detail with reference to the accompanying drawings.
[0127] For ease of description, in the following introduction Figure 6When referring to each step, the electronic device 10 is used as the executing entity, and the executing entity will not be elaborated hereinafter.
[0128] Herein, Figure 6 FIG. 6 shows a schematic flowchart of the electronic device 10 training the first model.
[0129] S600: Obtain multiple sets of point cloud data and captured images with corresponding relationships, and convert the point cloud data into radar images, where the point cloud data and the corresponding captured images are the acquisition data of the same area at the same data acquisition moment.
[0130] Exemplarily, the electronic device 10 can obtain multiple sets of point cloud data and captured images with corresponding relationships based on a jointly calibrated data acquisition system.
[0131] Herein, the method for obtaining multiple sets of point cloud data and captured images with corresponding relationships will be explained in detail in S700 hereinafter, and will not be elaborated herein.
[0132] S601: Perform object detection on the captured images based on the first image detection model to obtain the image detection results of multiple objects.
[0133] Exemplarily, the electronic device 10 can perform object detection on each captured image according to the first image detection model to obtain the image detection results of each object in each captured image.
[0134] Among them, the first image detection model can be the image detection model in S702 below, and specific details can be referred to the detailed description below Figure 7 and will not be elaborated herein.
[0135] S602: Map the image detection results of multiple objects in the corresponding radar images of the captured images to generate pseudo-labels corresponding to multiple objects in the radar images, where the pseudo-labels are used to identify the corresponding objects.
[0136] Exemplarily, for each captured image, the electronic device 10 can map the image detection results of each object in the corresponding radar image of the captured image to generate pseudo-labels corresponding to multiple objects in the radar image. In this way, it is not necessary to manually label the labels of each object in each radar image, effectively reducing the time cost and resource cost in the label annotation process. And, since the first image detection model can accurately detect each object, the pseudo-labels generated based on the mapped image detection results can also accurately identify each object in the radar image, ensuring the accuracy of the annotated pseudo-labels.
[0137] S603: Use the radar image with multiple pseudo-labels as the radar image with pseudo-labels.
[0138] Exemplarily, the electronic device 10 may use the radar image with pseudo-labels marked for each target as the radar image with pseudo-labels. It can be understood that based on the foregoing S602, the electronic device 10 may generate multiple radar images with pseudo-labels.
[0139] S604: Generate a first training set based on multiple radar images with pseudo-labels.
[0140] Exemplarily, the electronic device 10 may use multiple radar images with pseudo-labels as the first training set. The first training set includes the detection training set, detection validation set, and detection test set described below.
[0141] Specifically, for the process of the electronic device 10 dividing multiple radar images with pseudo-labels into a detection training set, a detection validation set, and a detection test set, refer to the detailed description in S703 below, and details will not be elaborated here.
[0142] S605: Train a second image detection model based on the first training set to obtain a first model capable of directly performing target detection based on radar images.
[0143] Exemplarily, the electronic device 10 may train a second image detection model based on the first training set to obtain a first model capable of directly performing target detection on radar images.
[0144] Among them, the second image detection model may be the same as the foregoing first image detection model. The second image detection model may be the image detection model in S702 below. For specific details, refer to the detailed description below Figure 7 and details will not be elaborated here. The first model may be the radar image detection model described below.
[0145] Here, the electronic device 10 trains the second image detection model to detect the radar images with pseudo-labels in the first training set, and transfers the detection ability of the second image detection model for captured images to the detection ability for radar images, obtaining the first model. Therefore, the first model can accurately detect radar images, avoiding target tracking errors caused by incorrect target detection.
[0146] Exemplarily, Figure 7 shows a schematic flowchart of an electronic device 10 training a radar image detection model.
[0147] Refer to Figure 7 , the detailed process of the electronic device 10 training the radar image detection model may include:
[0148] S700: Obtain multiple frames of point cloud data continuously collected by a millimeter-wave radar and multiple frames of captured images continuously collected by a camera, where the point cloud data and the captured images correspond one by one.
[0149] Exemplarily, the electronic device 10 can obtain multiple frames of point cloud data continuously collected by a millimeter-wave radar and multiple frames of captured images continuously collected by a camera. Moreover, the point cloud data and the captured images correspond to each other one by one.
[0150] Here, a data acquisition system for joint calibration of the millimeter-wave radar and the camera can be pre-established to collect point cloud data and captured images with a corresponding relationship through this data acquisition system. The aforementioned joint calibration refers to calibrating the millimeter-wave radar and the camera so that the data collected by the millimeter-wave radar and the camera can be fused in a unified coordinate system.
[0151] Specifically, the detection area of the millimeter-wave radar and the shooting area of the camera can be calibrated to be approximately the same. For example, the overlap degree between the detection area and the shooting area is greater than a preset threshold (such as 90%, 95%, 99%, etc.). Hereinafter, the detection area and the shooting area that have been calibrated to be approximately the same are collectively referred to as the detection area. Moreover, the millimeter-wave radar and the camera are adjusted to the same acquisition frequency so that the data acquisition moment when the millimeter-wave radar collects point cloud data is the same as the data acquisition moment when the camera collects captured images, that is, the timestamp of the point cloud data is synchronized with the timestamp of the captured image. Based on the foregoing operations, a data acquisition system for joint calibration of the millimeter-wave radar and the camera can be obtained.
[0152] Based on this data acquisition system, there is a corresponding relationship between the point cloud data and the captured images with the same timestamp, which are the point cloud data and the captured images of the same area collected at the same moment.
[0153] In addition, the data acquisition system can continuously collect point cloud data and captured images with a corresponding relationship in a variety of different scenarios. Taking the image detection model provided in the embodiments of the present application being applied to the driving environment as an example, the aforementioned scenarios may include: highway driving scenarios, low-speed driving scenarios, block scenarios, suburban scenarios, daytime scenarios, nighttime scenarios, etc. The data acquisition system covers a variety of scenarios, which is beneficial to improving the detection accuracy of the trained radar image detection model.
[0154] Based on the foregoing content, it can be known that the electronic device 10 can obtain multiple groups of point cloud data and captured images with a corresponding relationship based on the aforementioned data acquisition system.
[0155] S701: Generate radar images corresponding one by one to the captured images based on the point cloud data.
[0156] Exemplarily, after acquiring multiple groups of corresponding point cloud data and captured images, the electronic device 10 can pre-process the acquired continuous multi-frame point cloud data according to the performance of the millimeter-wave radar and the scenario in which the point cloud data is collected, and eliminate points located at the edge of the point cloud data (i.e., boundary points) and points in the point cloud data that deviate from the normal distribution (i.e., outliers), so as to reduce the noise of the point cloud data.
[0157] Furthermore, the electronic device 10 can project the pre-processed continuous multi-frame point cloud data to a bird's-eye view to obtain continuous multi-frame radar images.
[0158] Obviously, since the point cloud data and the photographed image have a corresponding relationship, the radar image generated based on the point cloud data has a corresponding relationship with the photographed image. A set of radar images and photographed images with a corresponding relationship are different data of the same area at the same time.
[0159] S702: Generate multiple frames of continuous radar images with pseudo labels based on multiple groups of corresponding photographed images and radar images.
[0160] Exemplarily, the electronic device 10 may generate a plurality of consecutive frames of radar images with pseudo-labels, that is, radar images with pseudo-labels, based on a plurality of groups of photographed images and radar images having a corresponding relationship.
[0161] Specifically, the electronic device 10 can perform target detection on multiple consecutive frames of captured images based on the image detection model to obtain image detection results of each target in each frame of captured images. Here, the aforementioned image detection model may include image detection models with high detection accuracy such as YOLO series models, R-CNN series models, Transformer series models, cornerNet, centerNet, etc. The image detection result of the target may include information such as a bounding box and category attributes generated based on the features of the target. Among them, the bounding box is used to determine the position of the target in the captured image, including the pixel coordinates of the target in the captured image; the category attribute is used to determine the category of the target.
[0162] Furthermore, for each frame of captured image, the electronic device 10 may map the image detection results of all targets therein to the radar image corresponding to the captured image. Specifically, the electronic device 10 may map the pixel coordinates of all targets in the captured image to the radar image corresponding to the captured image, and obtain the mapping results of each target in the radar image. Moreover, the electronic device 10 may add a pseudo-label corresponding to the target at the position of each mapping result according to the category attribute of the target, and obtain a radar image with a pseudo-label. Based on this method, the electronic device 10 can obtain multiple consecutive frames of radar images with pseudo-labels.
[0163] Here, Figure 8Shows a schematic diagram of a process for generating a radar image with pseudo-labels.
[0164] Refer to Figure 8 , a data acquisition system including a camera 80 and a millimeter-wave radar 81 can be mounted on a vehicle, and the detection areas of the camera 80 and the millimeter-wave radar 81 include a target A.
[0165] Based on this scenario, the electronic device 10 can obtain the captured image 800 captured by the camera 80 at time t0, and the point cloud data collected by the millimeter-wave radar 81 at time t0, and convert the point cloud data into a radar image 801.
[0166] The electronic device 10 can detect that the pixel coordinates of the target A in the captured image 800 are (u, v) based on an image detection model. Furthermore, the electronic device 10 can use the camera internal parameters (such as focal length, optical center position, etc.) to convert the pixel coordinates (u, v) into a point in the camera coordinate system of the camera 80, and then convert this point in the camera coordinate system into a point in the radar coordinate system of the millimeter-wave radar 81 through joint calibration parameters. Furthermore, the electronic device 10 can perform discretization processing on this point in the radar coordinate system to obtain the coordinates (u', v') in the radar image 801. Here, the coordinates (u', v') are the mapping results of the pixel coordinates (u, v). Therefore, the electronic device 10 can add pseudo-labels at the coordinates (u', v') of the radar image 801 according to the category attributes of the target A to obtain a radar image 802 with pseudo-labels including the target A.
[0167] Here, the joint calibration parameters are a set of parameters determined in advance for converting the camera coordinate system into the radar coordinate system, and can include a rotation matrix, a translation vector, etc. The rotation matrix is used to represent how the camera coordinate system is aligned with the radar coordinate system through rotation, and the translation vector is used to represent how the camera coordinate system is aligned with the radar coordinate system through translation.
[0168] Here, the X-axis of the camera coordinate system refers to the horizontal direction, corresponding to the width of the captured image; the Y-axis refers to the vertical direction, corresponding to the height of the image; the Z-axis represents the depth direction from the camera 600, pointing to the captured target. The X-axis of the radar coordinate system points in the forward direction of the millimeter-wave radar 81; the Y-axis refers to the horizontal direction; the Z-axis refers to the vertical direction, representing the height of the millimeter-wave radar 81 relative to the ground.
[0169] S703: Train an image detection model based on consecutive multi-frame radar images with pseudo-labels to obtain a radar image detection model.
[0170] Exemplarily, the electronic device 10 may divide the obtained consecutive multi-frame radar images with pseudo-labels into a detection training set, a detection validation set, and a detection test set. Here, the multi-frame radar images with pseudo-labels in the detection training set can maintain continuity, that is, the detection training set may include multi-frame radar images with pseudo-labels having a temporal relationship.
[0171] Here, the present application does not make restrictive descriptions on the specific division method. In an example method, the electronic device 10 may use 70% to 80% of the obtained consecutive multi-frame radar images with pseudo-labels as the detection training set, 10% to 15% as the detection validation set, and 10% to 15% as the detection test set.
[0172] Furthermore, the electronic device 10 may use each frame of the radar images with pseudo-labels in the detection training set as training data to train the image detection model, and adjust the parameters of the image detection model based on the loss function. After completing the training of the entire detection training set, the electronic device 10 may monitor the performance metrics of the trained image detection model on the detection validation set to adjust the hyperparameters (such as learning rate, batch size, etc.) of the trained image detection model. Repeat the foregoing process until the performance metrics of the image detection model on the detection validation set meet the preset criteria.
[0173] Furthermore, when the performance metrics of the image detection model meet the preset criteria, the electronic device 10 may test the trained image detection model based on the detection test set to perform a final performance evaluation on the trained image detection model. If the evaluation result meets the preset criteria, the trained image detection model may be used as the trained radar image detection model.
[0174] It can be understood that training the image detection model based on the radar images with pseudo-labels enables the image detection model to learn the mapping result of the image detection result of the target in the radar image, and furthermore, to learn the features of the target in the radar image. Therefore, an image detection model can be trained to be a radar image detection model capable of performing target detection on radar images. Here, since the image detection result of the image detection model has a high accuracy, the pseudo-label, which is the mapping result of the image detection result, can also accurately label the target. Furthermore, the radar image detection model trained, validated, and tested based on the radar images with pseudo-labels can also perform target detection on radar images with high accuracy, providing a good upstream result for the target tracking task.
[0175] After the electronic device 10 trains the radar image detection model, it may further train the unsupervised tracking model according to the radar image and the radar image detection model.
[0176] Here, the training data of the unsupervised tracking model can be generated based on the aforementioned detection training set. Taking the generation of the training data of the unsupervised tracking model based on the aforementioned detection training set as an example, the training process of the unsupervised tracking model will be described in detail below. For ease of description, hereinafter, the radar images with pseudo-labels in the detection training set will be collectively referred to as radar images.
[0177] Exemplarily, Figure 9 Fig. 5 shows a schematic flowchart of an electronic device 10 for training an unsupervised tracking model.
[0178] For ease of description, hereinafter, when introducing Figure 9 each step, the electronic device 10 is used as the execution subject, and the execution subject will not be elaborated hereinafter.
[0179] Referring to Figure 9 , the process of the electronic device 10 training the unsupervised tracking model may include:
[0180] S900: Perform object detection on the detection training set based on the radar image detection model to obtain the features of each object in each frame of radar image.
[0181] Exemplarily, the electronic device 10 may input the detection training set into the radar image detection model to perform object detection on multiple consecutive frames of radar images in the detection training set, and obtain the features of each object in each frame of radar image.
[0182] Taking the detection training set including P consecutive radar images and the maximum number of objects in the radar image being Q as an example, after the electronic device 10 inputs the detection training set into the radar image detection model, it can obtain to corresponding to wherein, represents the i-th object in the j-th radar image, represents the feature of the i-th object in the j-th radar image, where 1 ≤ j ≤ P and 1 ≤ i ≤ Q.
[0183] S901: Generate a negative sample set based on the features of each object in each frame of radar image.
[0184] Exemplarily, the electronic device 10 may generate a negative sample set based on the features of each object in the detection training set obtained in S900. The negative sample set includes multiple negative sample pairs, and a negative sample pair may include the features of two different objects.
[0185] Here, Figure 10a taking two radar images in the detection training set as an example, Fig. 6 shows a schematic diagram of a process for constructing a negative sample set.
[0186] Referring toFigure 10a Exemplarily, the detection training set includes the T-th radar image (abbreviated as "image T") and the (T + 1)-th radar image (abbreviated as "image T+1"). Among them, image T includes Image T+1 includes And and are the same target, and are the same target, and are the same target.
[0187] In the process of constructing the negative sample set, since the targets in the same frame of radar image must be different targets, for example, the categories they belong to and / or their positions in the radar image are different. Therefore, the following conclusion can be drawn: the targets in the same frame of radar image must have different features.
[0188] Furthermore, the electronic device 10 can use this conclusion as a prior condition and use the features of any two targets in a frame of radar image as a negative sample pair in the negative sample set.
[0189] For example, the electronic device 10 can obtain based on the aforementioned S900 the feature of the feature of the feature of the feature of the feature of the feature of
[0190] Specifically, the electronic device 10 can construct a negative sample pair based on the the feature of the feature of the feature of in image T: and can construct a negative sample pair based on the the feature of the feature of the feature of in image T+1: and
[0191] The electronic device 10 can add the constructed negative sample pairs to the negative sample set.
[0192] S902: Generate a plurality of positive sample pairs based on the features of the targets in each frame of radar image.
[0193] Exemplarily, the electronic device 10 may generate a positive sample set based on the features of each target in the detection training set obtained by S900. The positive sample set includes a plurality of positive sample pairs, and one positive sample pair includes two features that match the same target.
[0194] In the process of constructing the positive sample set, since multiple frames of radar images in the detection training set have temporal continuity, two adjacent frames of radar images can represent the spatial information of the detection area at two consecutive acquisition moments. Since the spatial information of the detection area changes little at two consecutive acquisition moments, correspondingly, the difference between two adjacent frames of radar images is small. Therefore, the following conclusion can be obtained: the features of the same target in adjacent radar images are similar.
[0195] Furthermore, the electronic device 10 may use this conclusion as a prior condition to determine two targets with similar features in adjacent radar images as the same target, that is, use two similar features in adjacent radar images as a positive sample pair in the positive sample set.
[0196] Continuing to refer to the foregoing Figure 10a , Figure 10a Taking two radar images in the detection training set as an example, an example process of constructing a positive sample pair is also shown.
[0197] For example, the electronic device 10 may use the bipartite graph matching algorithm to to respectively perform similarity matching with to to obtain the successfully matched and and and
[0198] Therefore, the electronic device 10 may, based on the and with a similarity greater than the preset threshold, and match and with a similarity greater than the preset threshold, and match and with a similarity greater than the preset threshold, and match as the same target.
[0199] The electronic device 10 may construct positive sample pairs based on the foregoing matching results:
[0200] In the above manner, the electronic device 10 can construct multiple positive sample pairs according to the matching results of multiple adjacent radar images.
[0201] Herein, the aforementioned bipartite graph matching algorithm includes the Hungarian algorithm, the KM algorithm (Kuhn-Munkres algorithm), etc., and the present application does not make restrictive descriptions on the specific matching algorithm.
[0202] S903: Construct a positive sample set based on the generated multiple positive sample pairs.
[0203] Exemplarily, after obtaining multiple positive sample pairs, the electronic device 10 can generate a positive sample set based on the multiple positive sample pairs.
[0204] Based on the content in the aforementioned S902, whether two features can be successfully matched into a positive sample pair will be affected by the matching accuracy of the bipartite graph matching algorithm, and the obtained positive sample pairs may be incorrect due to the low matching accuracy of the bipartite graph matching algorithm. Therefore, to avoid the adverse effects of incorrect matching of positive sample pairs on subsequent training of the unsupervised tracking model, the electronic device 10 can perform multi-frame verification on the determined positive sample pairs, and only add the successfully verified positive sample pairs to the positive sample set.
[0205] Specifically, Figure 10b Taking three consecutive radar images in the detection training set as an example, a schematic diagram of the process of performing multi-frame verification on positive sample pairs is shown.
[0206] See Figure 10b , it is known that the detection training set includes the aforementioned image T, the aforementioned image T + 1, and image T + 2 (i.e., the (T + 2)-th radar image in the detection training set). Among them, image T includes Image T + 1 includes Image T + 2 includes And, and are the same target, and are the same target, and are the same target.
[0207] Here, since the images T, T+1, and T+2 are consecutive radar images, the acquisition time of the point cloud data corresponding to the image T is close to the acquisition time of the point cloud data corresponding to the image T+2. Therefore, the differences between the images T, T+1, and T+2 are small. For example, the images T, T+1, and T+2 may include the same target, and the feature similarity of the same target in the images T, T+1, and T+2 is high. Therefore, when the matching results are correct, based on the matching result between the image T and the image T+1, and the matching result between the image T+1 and the image T+2, the matching result between the image T and the image T+2 should be deduced.
[0208] For example, matching the image T and the image T+1 can obtain a positive sample pair Matching the image T+1 and the image T+2 can obtain a positive sample pair Therefore, a positive sample pair can be deduced
[0209] Correspondingly, directly matching the image T and the image T+2 can directly obtain a positive sample pair
[0210] Obviously, the matching result deduced based on the images T, T+1, and T+2 is the same as the matching result directly obtained by matching the image T and the image T+2, indicating that the positive sample pair obtained by matching the image T and the image T+2 is correct, that is, it has passed the multi-frame verification. Therefore, the positive sample pair obtained by matching the image T and the image T+2 can be added to the positive sample set.
[0211] Conversely, if the deduced matching result is different from the directly obtained matching result, it means that at least one of the matching results between the image T and the image T+2 is incorrect. Therefore, the positive sample pair obtained by matching the image T and the image T+2 may not be added to the positive sample set.
[0212] Based on a multi-frame verification process that is Figure 5b substantially the same, the electronic device 10 only adds the positive sample pairs that have successfully passed the verification to the positive sample set. Based on this method, the accuracy of the constructed positive sample set can be effectively improved, and further, the tracking accuracy of the unsupervised tracking model can be improved.
[0213] It can be understood that the foregoing process of multi-frame verification of positive sample pairs based on three consecutive radar images is only an example, and the electronic device 10 can also perform multi-frame verification of positive sample pairs based on more than three consecutive radar images. For example, multi-frame verification of positive sample pairs can be performed based on 4 to 6 consecutive radar images.
[0214] S904: Generate a tracking training set, a tracking validation set, and a tracking test set based on the positive sample set and the negative sample set.
[0215] Exemplarily, the electronic device 10 can generate a tracking training set for training an unsupervised tracking model based on the positive sample set and the negative sample set. Specifically, the electronic device 10 can add some positive sample pairs in the positive sample set and some negative sample pairs in the negative sample set to the tracking training set.
[0216] This application does not make any restrictive statements about the number of positive sample pairs and negative sample pairs added to the tracking training set. In one example, the electronic device 10 can add 70% to 80% of the positive sample pairs in the positive sample set and 70% to 80% of the negative sample pairs in the negative sample set to the training tracking set.
[0217] In addition, the electronic device 10 can also generate a tracking validation set and a tracking test set for validating and testing the unsupervised tracking model according to the positive sample set and the negative sample set. Specifically, the electronic device 10 can add some positive sample pairs in the positive sample set and some negative sample pairs in the negative sample set to the tracking validation set, and add some positive sample pairs in the positive sample set and some negative sample pairs in the negative sample set to the tracking test set.
[0218] Similarly, this application does not make any restrictive statements about the number of positive sample pairs and negative sample pairs added to the tracking validation set and the tracking test set. In one example, the electronic device 10 can add 15% to 10% of the positive sample pairs in the positive sample set and 15% to 10% of the negative samples in the negative sample set to the tracking validation set. The electronic device 10 can add 15% to 10% of the positive sample pairs in the positive sample set and 15% to 10% of the negative samples in the negative sample set to the tracking test set.
[0219] S905: Determine the MLP network and loss function required to generate an unsupervised tracking model.
[0220] Exemplarily, the electronic device 10 can pre-design or select a multilayer perceptron (MLP) network and train the MLP network into an unsupervised tracking model based on the tracking training set. Among them, during the process of training the unsupervised tracking model, the loss function used can be the triplet loss, and the content of the triplet loss is as follows:
[0221] triplet loss L = max(D P - D N + α, 0)
[0222] where D Pis the distance between two features in a positive sample pair, D N is the distance between two features in a negative sample pair, D P and D N can be determined from the output data of the MLP network. Here, for the sake of narrative coherence, the determination of D based on the output data of the MLP network will be specifically described below P and D N The process will not be elaborated here.
[0223] where α is the margin parameter, D P -D N should be greater than α to prevent the MLP network from learning a local optimal solution where D P is slightly less than D N This ensures that the MLP network can learn discriminative features.
[0224] S906: Set the hyperparameters of the MLP network and its loss function according to the training validation set.
[0225] Here, the electronic device 10 can first set hyperparameters for the MLP network and the triplet loss based on the tracking validation set. The hyperparameters can include the number of network layers of the MLP network, activation function, learning rate, batch size (batch_size), margin parameter α in the triplet loss, etc.
[0226] Specifically, the electronic device 10 can adjust each hyperparameter during multiple rounds of validation of the MLP network based on the validation set to determine a combination of hyperparameters that can make the loss function value L of the triplet loss converge. The electronic device 10 can determine this combination of hyperparameters as the initial hyperparameters of the MLP network and the triplet loss, and further adjust the parameters of each initial hyperparameter based on the training process in S907 below.
[0227] Here, the multiple rounds of validation process of the MLP network based on the tracking validation set is essentially the same as the multiple rounds of training process of the MLP network based on the tracking training set. For specific details, refer to the description in S907 below. This will not be elaborated here.
[0228] S907: Train the MLP network based on the tracking training set.
[0229] Furthermore, the electronic device 10 can select approximately the same number of positive sample pairs and negative sample pairs from the tracking training set to generate the training data for one iteration of the unsupervised tracking model. Here, the selected positive sample pairs and negative sample pairs include multiple groups of corresponding positive and negative samples. Among them, a group of corresponding positive and negative samples includes the positive sample pair and negative sample pair of the same target. For example, the positive sample pair includes two different features of target A, and the negative sample pair includes the feature of target A and the features of other targets.
[0230] Furthermore, the electronic device 10 can calculate the difference between the two features in each selected positive sample pair to obtain each positive sample difference; calculate the difference between the two features in each selected negative sample pair to obtain each negative sample difference. The electronic device 10 can input all the obtained positive sample differences and negative sample differences into the MLP network as training data to perform one round of training on the MLP network. Here, the number of all feature differences input into the MLP network can be determined according to the batch size.
[0231] Based on all the input positive sample differences and negative sample differences, the MLP network can output the sample distances corresponding to each positive sample difference and negative sample difference. Among them, the sample distance refers to the distance between two different features in a positive sample pair / negative sample pair. The sample distance can be used to measure the similarity between two features. The larger the sample distance, the smaller the similarity between the two features, and the lower the possibility that they are the same target; the smaller the sample distance, the greater the similarity between the two features, and the higher the possibility that they are the same target.
[0232] Furthermore, the electronic device 10 can classify all the output sample distances to obtain the sample distances of all positive sample pairs (hereinafter referred to as the "positive sample distance set"), and the sample distances of all negative sample pairs (hereinafter referred to as the "negative sample distance set"). And the electronic device 10 can determine the loss function value L obtained from the first training according to the positive sample distance set, the negative sample distance set, and the triplet loss.
[0233] Specifically, the electronic device 10 can use the positive sample distance set as D in the triplet loss P input the triplet loss, use the negative sample distance set as D in the triplet loss N input the triplet loss, and based on the determined margin parameter α, obtain the loss function value L. So far, the electronic device 10 has completed the first training of the MLP.
[0234] Furthermore, the electronic device 10 can repeat the process that is essentially the same as the first training to determine the loss function value L obtained from the second training. And the electronic device 10 can adjust the parameters of the MLP network based on the convergence situation of L during the two training processes. Repeat the training process multiple times until L converges stably to the set threshold, and then the electronic device 10 can stop training the MLP network. Here, stable convergence means that L does not make obvious updates during multiple rounds of training. The specific number of training times can be adaptively determined according to the actual training situation, and this application does not limit the number of training times for stable convergence.
[0235] It should be clear that when L converges stably, it means that the trained MLP network can obtain stable D P and D N and, DP -D N Greater than α. That is, based on the training process of the MLP network, the MLP network learns the feature similarity of the same target (equivalent to the sample distance of the positive sample pair), and the feature difference of different targets (equivalent to the sample distance of the negative sample pair). Among them, there is an obvious difference between the sample distance of the positive sample pair and the sample distance of the negative sample pair.
[0236] For example, Figure 11 Shows a schematic diagram of a sample distance.
[0237] See Figure 11 , taking target A as an example, when the trained MLP network can obtain stable D P and D N After that, the targets corresponding to the features with a sample distance less than D from the feature A1 of target A are all positive samples of target A (that is, the same target as target A); the targets corresponding to the features with a sample distance greater than D from the feature A1 of target A are all negative samples of target A (that is, different targets from target A). P N P N
[0238] For example, target B is the same target as target A, target B is a positive sample of target A, and the feature B1 of target B and the feature A1 of target A are a pair of positive sample pairs. Therefore, the sample distance between feature B1 and feature A1 is less than D P .
[0239] For another example, target C is a different target from target A, target C is a negative sample of target A, and the feature C1 of target C and the feature A1 of target A are a pair of negative sample pairs. Therefore, the sample distance between feature C1 and feature A1 is greater than D N .
[0240] Among them, the distance D between D P and D N is greater than α.
[0241] S908: Set a distance threshold for the trained MLP network based on the tracking validation set.
[0242] Exemplarily, after the electronic device 10 finishes training the MLP network, it can input multiple positive sample differences and multiple negative sample differences generated based on the tracking validation set into the MLP network to obtain the sample distances corresponding to each positive sample difference output by the MLP network, and the sample distances corresponding to each negative sample difference.
[0243] Furthermore, the electronic device 10 can set a distance threshold of the MLP network according to the sample distances corresponding to the obtained positive sample differences and the sample distances corresponding to the negative sample differences. The distance threshold is greater than the sample distances corresponding to the positive sample differences and less than the sample distances corresponding to the negative sample differences. That is, the sample distances less than the distance threshold are all the sample distances corresponding to the positive sample differences, and the sample distances greater than the distance threshold are all the sample distances corresponding to the negative sample differences.
[0244] Furthermore, the MLP network can perform post-processing on the obtained sample distances based on the distance threshold and convert the sample distances into similarities for output. Specifically, for the sample distances less than the distance threshold, the MLP network can determine the output matching result as 1; for the sample distances greater than the distance threshold, the MLP network can determine the output matching result as 0.
[0245] It can be understood that based on the above method, for an input feature difference, if the output matching result is 1, it indicates that the feature difference is a positive sample difference, that is, the two features corresponding to the feature difference belong to the same target. On the contrary, for an input feature difference, if the output matching result is 0, it indicates that the feature difference is a negative sample difference, that is, the two features corresponding to the feature difference belong to different targets.
[0246] Among them, the process of the electronic device 10 generating multiple positive sample differences and multiple negative sample differences based on the tracking verification set is substantially the same as the process of obtaining multiple positive sample differences and multiple negative sample differences based on the tracking training set in S907 described above, and will not be elaborated here.
[0247] S909: Test the MLP network with the set distance threshold based on the tracking test set to obtain an unsupervised tracking model.
[0248] Exemplarily, after setting the distance threshold of the MLP network, the electronic device 10 can input the multiple positive sample differences and multiple negative sample differences generated based on the tracking test set into the MLP network to obtain the similarities corresponding to the positive sample differences output by the MLP network and the similarities corresponding to the negative sample differences.
[0249] If, among the matching results corresponding to all the positive sample differences output by the MLP network, the proportion of matching results being 1 is greater than a preset threshold (e.g., 99%, 99.9%, etc.), and, among the matching results corresponding to all the negative sample differences, the proportion of matching results being 0 is greater than a preset threshold (e.g., 99%, 99.9%, etc.), then the MLP network passes the test, and the MLP network is an unsupervised tracking model obtained through training. Otherwise, it indicates that the set distance threshold is inaccurate, and the electronic device 10 can repeat the above S708 to readjust the distance threshold until the MLP network passes the test.
[0250] It can be understood that based on S900 to S909, the electronic device 10 trains the MLP network into an unsupervised tracking model through contrastive learning. This unsupervised tracking model can determine whether two features are features of the same target based on the feature difference between the two input features. Moreover, during the process of training the unsupervised tracking model by the electronic device 10, there is no need for manual annotation of training data, effectively reducing the time cost and resource cost required for training the unsupervised tracking model.
[0251] Figure 12 According to an embodiment of the present application, a schematic structural diagram of an electronic device is shown. The electronic device may include the aforementioned electronic device 10.
[0252] As Figure 12 shown, in some embodiments, the electronic device may include one or more processors 1200, a system control logic 1201 connected to at least one of the processors 1200, a system memory 1202 connected to the system control logic 1201, a non-volatile memory (NVM) 1203 connected to the system control logic 1201, and a network interface 1204 connected to the system control logic 1201.
[0253] In some embodiments, the processor 1200 may include one or more single-core or multi-core processors. In some embodiments, the processor 1200 may include any combination of a general-purpose processor and a dedicated processor (e.g., a graphics processor, an application processor, a baseband processor, etc.). For example, the processor 1200 may include any combination of a general-purpose processor and a dedicated processor required to implement the aforementioned channel package generation method.
[0254] In some embodiments, the system control logic 1201 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1200 and / or any suitable device or component communicating with the system control logic 1201.
[0255] In some embodiments, the system control logic 1201 may include one or more memory controllers to provide an interface to the system memory 1202. The system memory 1202 may be used to load and store data and / or instructions. In some embodiments, the system memory 1202 of the electronic device may include any suitable volatile memory, such as a suitable dynamic random access memory (DRAM).
[0256] The non-volatile memory 1203 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 1203 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of a hard disk drive (HDD), a compact disc (CD) drive, and a digital versatile disc (DVD) drive.
[0257] The non-volatile memory 1203 may include a portion of the storage resources mounted on the device of the electronic device, or it may be accessible by the device but not necessarily part of the device. For example, the non-volatile memory 1203 may be accessed via the network interface 1204 over a network.
[0258] Specifically, the system memory 1202 and the non-volatile memory 1203 / memory 516 may include temporary and permanent copies of the instructions 1205, respectively. The instructions 1205 may include instructions that, when executed by at least one of the processors 1200, may cause the electronic device to implement the processes as Figures 3 to 4 shown. In some embodiments, the instructions 1205, hardware, firmware, and / or its software components may alternatively / additionally be located in the system control logic 1201, the network interface 1204, and / or the processor 1200.
[0259] The network interface 1204 may include a transceiver for providing a radio interface for the electronic device to communicate with any other suitable device (such as a front-end module, an antenna, etc.) over one or more networks. In some embodiments, the network interface 1204 may be integrated with other components of the electronic device. For example, the network interface 1204 may be integrated with at least one of the processor 1200, the system memory 1202, the non-volatile memory 1203, and a firmware device (not shown) having instructions.
[0260] The network interface 1204 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 1204 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.
[0261] In one embodiment, at least one of the processors 1200 may be logically packaged with one or more controllers for the system control logic 1201 to form a system in package (SiP). In one embodiment, at least one of the processors 1200 may be integrated with the logic of one or more controllers for the system control logic 1201 on the same die to form a system on chip (SoC).
[0262] The electronic device may further include: an input / output (I / O) device 1206. The I / O device 1206 may include a user interface that enables a user to interact with the electronic device; the design of the peripheral component interface enables peripheral components to also interact with the electronic device.
[0263] In some embodiments, the user interface may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a sensor, a speaker, a microphone, a light emitting diode, a physical button.
[0264] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.
[0265] In some embodiments, the sensor may include, but is not limited to, a gyroscope sensor, an acceleration sensor, a GNSS antenna, and a positioning unit. The positioning unit may also be part of or interact with the network interface 1204 to communicate with components of a positioning network (e.g., GNSS satellites).
[0266] The embodiments of the present application also provide a computer program product for implementing the channel package generation method provided in the above embodiments.
[0267] The embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as computer program modules or module codes executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0268] A computer program module or module code can be applied to input instructions to perform the various functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0269] The module code can be implemented in a high-level modular language or an object-oriented programming language to communicate with the processing system. When needed, the module code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any particular programming language. In any case, the language can be a compiled language or an interpreted language.
[0270] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, a floppy disk, a compact disc, a CD-ROM, a magneto-optical disc, a read only memory (ROM), a random access memory (RAM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, a flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in the form of electrical, optical, acoustic, or other propagated signals using the Internet. Thus, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0271] References in the specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one exemplary implementation or technique disclosed in embodiments of the present application. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.
[0272] The disclosure of embodiments of the present application also relates to an apparatus for performing the operations in the text. The apparatus may be specially constructed for the required purposes or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, application specific integrated circuits (ASICs), or any type of medium suitable for storing electronic instructions, and each may be coupled to a computer system bus. In addition, the computers referred to in the specification may include a single processor or may be architectures involving multiple processors for increased computing power.
[0273] In addition, the language used in this specification has been principally selected for readability and instructional purposes and may not have been selected to delineate or circumscribe the disclosed subject matter. Accordingly, the disclosure of embodiments of the present application is intended to illustrate rather than limit the scope of the concepts discussed herein.
Claims
1. A model training method, characterized in that: The method comprises: Acquire multiple sets of point cloud data and photographed images having corresponding relationships, and convert the point cloud data into radar images, wherein the point cloud data and the corresponding photographed images are collected data of the same area at the same data collection time; Performing target detection on the captured image based on a first image detection model to obtain image detection results of multiple targets; Mapping the image detection results of the multiple targets in the radar image corresponding to the captured image, and generating pseudo labels corresponding to the multiple targets in the radar image, wherein the pseudo labels are used to identify corresponding targets; Using the radar image with the plurality of pseudo labels as a pseudo-labeled radar image; Generate a first training set based on the plurality of pseudo-labeled radar images; The second image detection model is trained based on the first training set to obtain a first model that can directly perform target detection based on radar images.
2. The method according to claim 1, characterized in that After the first model is obtained through training, the method further includes: generating a sample set based on a second training set and the first model, wherein the second training set includes a plurality of radar images; and Training the target tracking model based on the sample set to obtain a second model capable of implementing target tracking based on an output of the first model; The sample set includes: multiple groups of positive sample pairs across frames, and the positive sample pairs have passed multi-frame verification, wherein the positive sample pairs correspond to feature pairs of the same target.
3. The method according to claim 2, characterized in that The method of determining whether the first positive sample pair in the plurality of groups of positive sample pairs across frames passes the multi-frame verification is as follows: Inputting the plurality of radar images in the second training set into the first model respectively to obtain features of all targets in the plurality of radar images; The features of all targets in the multiple radar images include a first feature, a second feature and a third feature, wherein the first feature is a feature of a target in a radar image generated based on the first point cloud data, the second feature is a feature of a target in a radar image generated based on the second point cloud data, and the third feature is a feature of a target in a radar image generated based on the third point cloud data, and the data collection times of the first point cloud data, the second point cloud data and the third point cloud data are adjacent; Determining that the first feature and the third feature are successfully matched based on a bipartite graph matching algorithm, and determining the first feature and the third feature as the first positive sample pair; Based on the bipartite graph matching algorithm, it is determined that the first feature matches the second feature successfully, and the second feature matches the third feature successfully, and it is determined that the first positive sample pair passes the multi-frame verification.
4. The method according to claim 2, characterized in that: The sample set also includes: a plurality of groups of negative sample pairs, wherein the negative sample pairs correspond to feature pairs of different targets; The step of training the target tracking model based on the sample set to obtain the second model includes: Based on the multiple groups of positive sample pairs across frames, the multiple groups of negative sample pairs and the first loss function, a preset multi-layer perceptual network is trained by contrastive learning, so that the multi-layer perceptual network can output a second distance and a third distance, wherein the second distance corresponds to the distance between the features of the positive sample pair, and the third distance corresponds to the distance between the features of the negative sample pair; Based on the training results of the multi-layer perception network, a second model is obtained.
5. The method according to claim 4, characterized in that The method of determining the first negative sample pair in the plurality of groups of negative sample pairs is as follows: Inputting the plurality of radar images in the second training set into the first model respectively to obtain features of all targets in the plurality of radar images; Determine, among the features of all the targets in the plurality of radar images, a fourth feature and a fifth feature located in the same frame of radar image; The fourth feature and the fifth feature are determined as a first negative sample pair.
6. The method according to claim 5, characterized in that The method of training a preset multi-layer perception network by contrastive learning based on the multiple groups of positive sample pairs across frames, the multiple groups of negative sample pairs and the first loss function includes: Generate a plurality of positive sample differences based on the plurality of groups of positive sample pairs across frames, and generate a plurality of negative sample differences based on the plurality of groups of negative sample pairs; Using the multiple positive sample differences and the multiple negative sample differences as inputs for a first round of training of the multilayer perception network, obtaining multiple positive sample distances corresponding to the multiple positive sample differences output by the multilayer perception network, and multiple negative sample distances corresponding to the multiple negative sample differences; Determine the loss value of the first loss function based on the multiple positive sample distances and the multiple negative sample distances; if the loss value of the first loss function does not meet the training termination condition, adjust the network parameters of the multi-layer perception network and perform the next round of training; When the loss value of the first loss function meets the training termination condition, the trained multi-layer perception network is used as the second model.
7. The method according to claim 6, characterized in that After obtaining the second model, the method further includes: A distance threshold corresponding to the second model is set based on the second distance and the third distance, wherein the distance threshold is greater than the second distance and less than the third distance.
8. The method according to claim 7, characterized in that The first loss function includes triplet loss: triplet loss L=max(D P -D N +α,0) Among them, D P The multiple positive sample distances output by the multi-layer perception network; D N The multiple negative sample distances output by the multi-layer perception network; α is a margin parameter, which is used to represent the minimum distance between a first sample distance and a second sample distance, wherein the first sample distance is any one of the multiple negative sample distances, and the second sample distance is any one of the positive sample distances; L is the loss value of the first loss function.
9. The method according to claim 2, characterized in that: After the first model and the second model are obtained through training, the method further includes: Acquire the fourth point cloud data collected by the radar at the first moment; Acquire fifth point cloud data collected by the radar at the second moment; generating a first radar image based on the fourth point cloud data, and generating a second radar image based on the fifth point cloud data; Inputting the first radar image into the first model to obtain features of each of the m targets in the first radar image; and Inputting the second radar image into the first model to obtain features of each of the n targets in the second radar image; Based on the similarity between the features of each target in the m targets and the features of each target in the n targets, q tracking targets are determined to achieve target tracking of the q tracking targets from the first moment to the second moment.
10. The method according to claim 9, characterized in that The determining q tracking targets based on the similarity between the feature of each target in the m targets and the feature of each target in the n targets includes: Determine based on the second model that the similarity between the feature of the first target among the m targets and the feature of the second target among the n targets meets a preset condition, and determine the first target or the second target as the third target among the q tracking targets.
11. The method according to claim 10, characterized in that The determining, based on the second model, that the similarity between the feature of the first target among the m targets and the feature of the second target among the n targets meets a preset condition includes: inputting a difference between a feature of the first target and a feature of the second target into the second model; Acquire a first distance between a feature of the first target and a feature of the second target output by the second model; Based on the first distance being less than a distance threshold corresponding to the second model, it is determined that the similarity between the feature of the first target and the feature of the second target meets a preset condition.
12. An electronic device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the model training method described in any one of claims 1 to 11.
13. A computer readable medium, characterized in that The readable medium stores instructions, which, when executed on a computer, enable the computer to execute the model training method described in any one of claims 1 to 11.
14. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements the model training method described in any one of claims 1 to 11.