A multi-sensor data fusion method and device
By using a pre-trained visual and radar feature fusion model, the problem of improper parameter design in multi-sensor data fusion is solved, achieving accurate target matching and fusion, and improving the perception capability of the autonomous driving system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MOMENTA (SUZHOU) TECHNOLOGY CO LTD
- Filing Date
- 2021-05-21
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, improper parameter design and selection during multi-sensor data fusion can lead to inaccurate matching and fusion results, and make the data difficult to maintain and expand.
A pre-trained target visual feature fusion model and radar visual feature fusion model are adopted. The matching and fusion features of visual and radar perceived targets are determined by machine learning methods, and the training model is used to achieve accurate fusion of multi-sensor data.
It achieves accurate fusion of multi-sensor data, reduces target omissions and mismatches, improves the stability and accuracy of fusion results, simplifies parameter adjustment, and facilitates subsequent maintenance and expansion.
Smart Images

Figure CN115457353B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and more specifically, to a method and apparatus for fusing multi-sensor data. Background Technology
[0002] In autonomous driving solutions, to ensure accurate decision-making in open driving environments, vehicles need to perceive surrounding targets and their size, position, speed, and other characteristics to adjust their behavior in real time. Accordingly, to achieve more comprehensive and accurate perception of the surrounding environment, different types of sensors are typically deployed to fully leverage their strengths and obtain more accurate results. Furthermore, to increase the vehicle's perception range, multiple of the same type of sensor are often deployed, such as image acquisition devices in the front, rear, left, and right directions. With this deployment scheme, the perception ranges of the sensors overlap, meaning the same target can be perceived by different sensors, resulting in multiple perception results for the same target at the same time. Ideally, only one perception result should exist for each target at any given moment. Therefore, fusing the perception results from different sensors into a single target perception result is crucial.
[0003] Currently, the process of fusing the perception results of different sensors into a single target perception result involves: obtaining visual perception results and millimeter-wave radar perception results; determining mutually matched visual perception targets and radar perception targets based on the velocity and position corresponding to each visually perceived target in the visual perception results, the velocity and position corresponding to each radar perception target in the millimeter-wave radar perception results, and a pre-set matching threshold; and determining the fused velocity and position based on a preset fusion rule and the velocity and position corresponding to the mutually matched visually perceived targets and radar perception targets, thereby realizing the fusion of the perception results of different sensors into a single target perception result.
[0004] In the above process, the matching process between visually perceived targets and radar-perceived targets, as well as the corresponding fusion process of velocity and position, are all determined by manually set thresholds and rules. This process involves the design and selection of many parameters, which is not convenient for subsequent maintenance and expansion. Furthermore, if the design and selection of parameters are inappropriate, it is easy to encounter situations where there is no matching or the fusion results are inaccurate. Summary of the Invention
[0005] This invention provides a method and apparatus for fusing multi-sensor data to achieve accurate fusion of multi-sensor data. The specific technical solution is as follows:
[0006] In a first aspect, embodiments of the present invention provide a method for fusing multi-sensor data, the method comprising:
[0007] Obtain the current visual perception result and the current first radar perception result of the target object at the current moment;
[0008] Based on the current visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model, the fused visual features corresponding to each current visual perception target are determined. The target visual feature fusion model is a model trained based on the label perception features corresponding to each sample object at each sample time and the sample visual perception features corresponding to each sample visual perception target.
[0009] Based on the fused visual features corresponding to each current visual perception target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and the pre-established radar visual feature fusion model, the matching current visual perception target and the current first radar perception target, as well as the corresponding current fused perception features, are determined. The pre-established radar visual feature fusion model is a model trained based on the sample visual perception features, the label perception features, and the sample first radar perception results corresponding to each sample first radar perception target at each sample time.
[0010] Secondly, embodiments of the present invention provide a multi-sensor data fusion apparatus, the apparatus comprising:
[0011] The first acquisition module is configured to acquire the current visual perception result and the current first radar perception result of the target object at the current moment.
[0012] The first determining module is configured to determine the fused visual features corresponding to each current visual perception target based on the current visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model. The target visual feature fusion model is a model trained based on the label perception features corresponding to each sample object at each sample time and the sample visual perception features corresponding to each sample visual perception target.
[0013] The second determining module is configured to determine the matching current visual perception target and the current first radar perception target, as well as the corresponding current fused perception features, based on the fused visual features corresponding to each current visual perception target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and a pre-established radar visual feature fusion model. The pre-established radar visual feature fusion model is a model trained based on the sample visual perception features, the label perception features, and the sample first radar perception results corresponding to each sample first radar perception target at each sample time.
[0014] As described above, the multi-sensor data fusion method and apparatus provided in this embodiment of the invention obtains the current visual perception result and the current first radar perception result corresponding to the target object at the current time; based on the visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model, the fused visual features corresponding to each current visual perception target are determined. The target visual feature fusion model is a model trained based on the label perception features corresponding to each sample object at each sample time and the sample visual perception features corresponding to each sample visual perception target; based on the fused visual features corresponding to each current visual perception target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and the pre-established radar visual feature fusion model, the matching current visual perception target and the current first radar perception target, and the corresponding current fused perception features are determined. The pre-established radar visual feature fusion model is a model trained based on the sample visual perception features, label perception features, and the sample first radar perception results corresponding to each sample first radar perception target corresponding to each sample object at each sample time.
[0015] By applying embodiments of the present invention, a pre-trained target visual feature fusion model can be used to fuse the current visual perception features corresponding to each current visual perception target in the current visual perception result, obtaining fused visual features corresponding to each current visual perception target. Then, using a pre-established radar visual feature fusion model, the fused visual features corresponding to matching current visual perception targets (i.e., those corresponding to the same physical target) and the current first radar perception features corresponding to the current first radar perception target can be fused, resulting in relatively accurate perception features for each physical target. By utilizing the target visual feature fusion model and the pre-established radar visual feature fusion model, accurate fusion of multi-sensor data can be achieved. Of course, implementing any product or method of the present invention does not necessarily require simultaneously achieving all the advantages described above.
[0016] The innovative aspects of this invention include:
[0017] 1. Based on the pre-trained target visual feature fusion model, the visual perception features corresponding to each current visual perception target in the current visual perception result can be fused to obtain the fused visual features corresponding to each current visual perception target. Then, using the pre-established radar visual feature fusion model, the fused visual features corresponding to the current visual perception target that matches, i.e., the same physical target, and the current first radar perception feature corresponding to the current first radar perception target can be fused to obtain relatively accurate perception features of each physical target. By using the target visual feature fusion model and the pre-established radar visual feature fusion model, accurate fusion of multi-sensor data can be achieved.
[0018] 2. To avoid the omission of targets in the current visual perception results, the radar target recognition model is established in advance, and the current radar perception features corresponding to the current radar perception target are used to determine the current radar real target and its corresponding current radar perception features, which can avoid the omission of targets to a certain extent.
[0019] 3. To avoid inaccurate matching in a single frame, which could lead to unstable multi-sensor data fusion results, a pre-established tracking model is trained. Based on the historical features of each historical target and the current sensing features of the current sensing target, the matching historical target and the current sensing target are determined, so as to determine the current sensing target with more accurate feature fusion of the object at the current moment.
[0020] 4. To avoid missing any targets, a pre-established prediction model can be trained in advance. This model can then predict the characteristics of each historical target at the current moment based on its historical features and the pre-established prediction model. This will allow for the supplementation of any missed targets by using the predicted characteristics of each historical target at the current moment.
[0021] 5. Provide the training process for each model to obtain accurate models, laying the foundation for subsequent sensor data fusion.
[0022] 6. In the process of establishing the pre-established tracking model, after training the intermediate tracking model that meets the fourth convergence condition, a preset smoothness loss function is added to avoid mismatch between the current visual perception target and the current first radar perception target caused by the target visual feature fusion model and the pre-established radar visual feature fusion model, as well as the error of the first radar perception feature, which would cause the position and / or velocity of the current perception target to jump relative to the corresponding historical target. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0024] Figure 1 This is a schematic flowchart of a multi-sensor data fusion method provided in an embodiment of the present invention;
[0025] Figure 2This is another flowchart illustrating the multi-sensor data fusion method provided in an embodiment of the present invention.
[0026] Figure 3 This is another flowchart illustrating the multi-sensor data fusion method provided in an embodiment of the present invention.
[0027] Figure 4 This is another flowchart illustrating the multi-sensor data fusion method provided in an embodiment of the present invention.
[0028] Figure 5A This is a schematic diagram illustrating the various structures and data flow of a single-frame fusion model provided in an embodiment of the present invention.
[0029] Figure 5B A schematic diagram illustrating the various structures and data flow of the continuous frame tracking model provided in an embodiment of the present invention;
[0030] Figure 6 This is a schematic diagram of a multi-sensor data fusion device provided in an embodiment of the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0032] It should be noted that the terms "comprising" and "having," and any variations thereof, in the embodiments and drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0033] This invention provides a method and apparatus for fusing multi-sensor data to achieve accurate fusion of multi-sensor data. The embodiments of this invention are described in detail below.
[0034] Figure 1 This is a schematic flowchart of a multi-sensor data fusion method provided in an embodiment of the present invention. The method may include the following steps S101-S103:
[0035] S101: Obtain the current visual perception result and the current first radar perception result of the target object at the current moment.
[0036] The multi-sensor data fusion method provided in this invention can be applied to any electronic device with computing capabilities, such as a terminal or a server. In one implementation, the functional software implementing this method can exist as a standalone client software or as a plugin for existing client software, for example, as a functional module of an autonomous driving system; both are possible.
[0037] The target object can be an autonomous vehicle or a robot. The target object can be equipped with various types of sensors, including but not limited to sensors for environmental perception and sensors for localization. Sensors for environmental perception can include, but are not limited to, image acquisition devices and radar. In one scenario, sensors for environmental perception can also be used to assist in target object localization. Sensors for localization can include, but are not limited to, wheel speed sensors, IMU (Inertial Measurement Unit), GPS (Global Positioning System), and GNSS (Global Navigation Satellite System).
[0038] To obtain a more comprehensive image of the surrounding environment, multiple image acquisition devices can be installed around the target object, each capturing images of its surroundings. The acquired images allow for the perception of the target object's environment, yielding visual perception results from each image acquisition device.
[0039] In one scenario, the radar may include a first radar and a sample tag radar (mentioned subsequently). The first radar may be a low-cost radar capable of obtaining information such as the position and velocity of targets in the surrounding environment, for example, a millimeter-wave radar. The sample tag radar may be a radar capable of obtaining multi-dimensional information about targets in the surrounding environment, such as a lidar. The sample tag radar can obtain information about the target's shape, size, velocity, type, and location.
[0040] In one implementation, when the target object is an autonomous vehicle, the perception target mentioned in this embodiment of the invention may include the vehicle, and may also include, but is not limited to, pedestrians. When the target object is a robot, the perception target mentioned in this embodiment of the invention may include both vehicles and pedestrians.
[0041] During the movement of the target object, its image acquisition devices can periodically acquire images of the target object's environment. These images can be perceived and identified by the image acquisition devices or other visual perception devices, yielding visual perception results from each acquisition device. Simultaneously, its first radar can periodically acquire radar data of the target object's environment, obtaining first radar perception results from this data. This provides the visual perception results and first radar perception results of the target object at each moment. Correspondingly, the electronic equipment can also obtain the visual perception results and first radar perception results of the target object at each acquisition moment, using these results as the current visual perception result and the current first radar perception result, respectively. Here, "current moment" can refer to the acquisition moment corresponding to the current sensing result requiring multi-sensor data fusion.
[0042] The current visual perception results may include: the current visual perception features of each image acquisition device of the target object at the current moment, and the current visual perception features of each visually perceived target in the environment where the target object is located. The image acquisition areas of different image acquisition devices overlap; correspondingly, the same physical target can be perceived by different image acquisition devices. Current visual perception features may include, but are not limited to: the speed, type, shape, and spatial location information of the currently visually perceived target. The current first radar perception results may include: the current first radar perception features of each visually perceived target in the environment where the target object is located, and the current first radar perception features may include, but are not limited to: the speed and spatial location information of the currently perceived target.
[0043] In one scenario, the target sensed by the first radar is generally a spatial point, and the corresponding radar sensing characteristics of the target are the velocity and spatial location information of the point.
[0044] S102: Based on the current visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model, determine the fused visual features corresponding to each current visual perception target.
[0045] Among them, the target visual feature fusion model vison_fusion MLP is a model trained based on the label perception features of each sample object at each sample time and the sample visual perception features corresponding to the visual perception target of each sample.
[0046] The electronic device locally or in a connected storage device stores a pre-trained target visual feature fusion model, which is used to fuse the current visual perception features corresponding to each current visual perception target for the same physical target, to obtain the fused visual features corresponding to each current visual perception target. This pre-trained target visual feature fusion model is a model trained based on the label perception features corresponding to each sample object at each sample time and the sample visual perception features corresponding to each sample visual perception target.
[0047] The number of visual perception results corresponding to each visual perception target can be at least one and at most the number of image acquisition devices corresponding to the corresponding sample object. For example, if there are four image acquisition devices for a sample object, then the maximum number of visual perception results corresponding to each visual perception target can be four. For clarity, the image acquisition devices for the sample object can be referred to as sample image acquisition devices.
[0048] During training, if the number of visual perception results corresponding to a sample visually perceived target is less than the preset number of perceptions (i.e., some image acquisition devices do not perceive a certain physical target, while others do), for the image acquisition devices that did not perceive the physical target, the visual perception features corresponding to that physical target in their sample visual perception results can be replaced by preset feature values. The preset number of perceptions can be the number of image acquisition devices for the sample object. The preset feature value can be 0.
[0049] Tag-aware characteristics refer to the perception features of a target object by its tag radar in the current moment, based on the surrounding environment. This tag radar can be a radar capable of acquiring multi-dimensional information about targets in the surrounding environment, such as a lidar (Light Detection and Ranging) system. Tag-aware characteristics can include information such as the shape, size, speed, type, and location of the perceived target. For clarity, the target perceived by the tag radar can be referred to as the tag radar-aware target.
[0050] Fusion can refer to the process of combining features of the same dimension from the current visual perception features corresponding to each current visual perception target into a single feature. For example, if there are three current visual perception features corresponding to current visual perception target 1 (meaning three image acquisition devices perceive current visual perception target 1), and each image acquisition device corresponds to one current visual perception feature for that target, the current visual perception features include the velocity and spatial position information of the target. Fusion of the current visual perception features corresponding to target 1 can mean: combining the velocity of the target into a single velocity from the three current visual perception features, which is then used as the fused velocity in the fused visual features corresponding to target 1; and combining the spatial position information of the target into a single spatial position information from the three current visual perception features, which is then used as the fused spatial position information in the fused visual features corresponding to target 1.
[0051] In the case where the target object is an autonomous vehicle, the sample object is an autonomous vehicle. In the case where the target object is a robot, the sample object is a robot.
[0052] In one scenario, an electronic device can determine whether a current visually perceived target corresponds to the same physical target by checking whether the spatial location information in the current visually perceived features corresponding to each current visually perceived target overlaps. Specifically, the electronic device determines that current visually perceived targets with overlapping spatial location information correspond to the same physical target; conversely, it determines that current visually perceived targets with non-overlapping spatial location information correspond to different physical targets. The aforementioned overlap of spatial location information can refer to the situation where the distance difference between the spatial location information does not exceed a preset spatial distance threshold.
[0053] S103: Based on the fused visual features corresponding to each current visually perceived target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and the pre-established radar visual feature fusion model, determine the matching current visually perceived target and the current first radar perception target, as well as the corresponding current fused perception features.
[0054] The pre-established radar visual feature fusion model, namely vision_radar_MLP (Multilayer Perceptron), is a model trained based on the visual perception features of the samples, the label perception features, and the first radar perception results of the first radar perception target corresponding to each sample object at each sample time.
[0055] In this step, the electronic device or the connected storage device stores a pre-established radar visual feature fusion model. The pre-established radar visual feature fusion model is a model trained based on the sample visual perception features of each sample object at each sample time corresponding to each sample visual perception target, the label perception features of each sample object at each sample time, and the sample first radar perception results of each sample object at each sample time corresponding to each sample first radar perception target.
[0056] The pre-established radar visual feature fusion model is used to: determine the corresponding current visual perception target and the current first radar perception target from each current visual perception and each current first radar perception target, that is, the matching current visual perception target and the current first radar perception target, and fuse the fused visual features corresponding to the matching current visual perception target and the current first radar perception features corresponding to the current first radar perception target to obtain the matching current visual perception target and the current fused perception features corresponding to the current first radar perception target.
[0057] Accordingly, the electronic device pairs each currently perceived visual target with each currently perceived first radar target in the current radar perception result. The fused visual features corresponding to each pair of currently perceived visual targets and the current radar perception features corresponding to each currently perceived first radar target are then input into a pre-established radar visual feature fusion model. Based on the fused visual features corresponding to each pair of currently perceived visual targets and the current radar perception features corresponding to each currently perceived first radar target, the affinity score (affinity_score) and the fused perception features corresponding to each pair of currently perceived visual targets and the current radar perception targets are determined using the pre-established radar visual feature fusion model. For clarity, the affinity score between each pair of currently perceived visual targets and the current radar perception targets will be referred to as the first affinity score.
[0058] The electronic device determines the matching current visual perception target and the current first radar perception target and their corresponding current fused perception features based on the first affinity score between each pair of current visual perception targets and the current first radar perception target, i.e. the corresponding matching perception targets, and outputs the result of obj_fusion.
[0059] In one scenario, the process of determining the matching current visual perception target and the current first radar perception target can be as follows: the current visual perception target and the current first radar perception target with the highest first affinity score that exceeds a preset score threshold are identified as the matching current visual perception target and the current first radar perception target.
[0060] In another scenario, due to the properties of the first radar, there may be corresponding physical targets among the targets sensed by the first radar. The first radar may sense multiple targets and their corresponding sensing features for a single physical target. To determine more accurate matching between the current visual and radar targets, the Hungarian matching algorithm and the first affinity score between each pair of paired current visual and radar targets are used to determine the optimal matching relationship between them.
[0061] After determining the optimal matching relationship between the current visually perceived target and the current first radar perceived target, if the first affinity score between the current visually perceived target and the current first radar perceived target with the optimal matching relationship exceeds a preset affinity threshold, then the current visually perceived target and the current first radar perceived target are determined to correspond to the same physical target and are considered to be matched current visually perceived targets and current first radar perceived targets. Conversely, if the first affinity score between the current visually perceived target and the current first radar perceived target with the optimal matching relationship does not exceed the preset affinity threshold, then the current visually perceived target and the current first radar perceived target are determined to correspond to different physical targets and are not considered to be matched current visually perceived targets and current first radar perceived targets.
[0062] In another implementation, for a current visually perceived target that does not match the current first radar perceived target, information representing that the current visually perceived target does not match the current first radar perceived target can be directly output, and the fused visual features corresponding to the current visually perceived target can be output. That is, for a current visually perceived target that does not match, the result of obj_vision_fusion is output.
[0063] Correspondingly, for the current first radar sensing target that does not match the current visual sensing target, information indicating that the current first radar sensing target does not match the current visual sensing target can be directly output, and its corresponding current first radar sensing feature can be output.
[0064] By applying the embodiments of the present invention, the visual perception features corresponding to each current visual perception target in the current visual perception result can be fused based on the pre-trained target visual feature fusion model to obtain the fused visual features corresponding to each current visual perception target. Then, by using the pre-established radar visual feature fusion model, the fused visual features corresponding to the current visual perception target that matches the same physical target and the current first radar perception feature corresponding to the current first radar perception target can be fused to obtain relatively accurate perception features of each physical target. By using the target visual feature fusion model and the pre-established radar visual feature fusion model, accurate fusion of multi-sensor data can be achieved.
[0065] Furthermore, in this embodiment of the invention, each model is pre-trained, and the parameters of each model are automatically adjusted and generated based on the training data. To a certain extent, there is no need for frequent manual adjustment of the corresponding parameter thresholds, which frees up manpower and facilitates subsequent maintenance and expansion.
[0066] In another embodiment of the invention, such as Figure 2 As shown, in Figure 1 Based on the process shown, the method further includes the following step S104:
[0067] S104: Based on the current radar sensing characteristics corresponding to the current radar sensing target and the pre-established radar target recognition model, determine the current radar real target and its corresponding current radar sensing characteristics from the current radar sensing targets.
[0068] The pre-established radar target recognition model is a model trained based on the radar perception features and label perception features of the first radar perception target corresponding to each sample.
[0069] Considering that the targets sensed by the first radar may include non-real targets, and considering that factors such as occlusion and harsh environmental conditions may cause occluded targets and / or unclear targets to appear in the images acquired by the image acquisition device, leading to missed visual targets, this implementation addresses these issues by pre-storing a pre-established radar target recognition model locally on the electronic device or in a connected storage device. This pre-established radar target recognition model is trained based on the first radar sensing features and label sensing features corresponding to each sample of the first radar sensed target. It can identify the first radar real targets that are real physical targets from the first radar sensed targets. Here, a real physical target refers to a target of the required sensing type; for example, if the target is an autonomous vehicle, the vehicle can be a target of the required sensing type.
[0070] Accordingly, the electronic device inputs the current radar sensing features corresponding to the current radar-sensing target into a pre-established radar target recognition model (radar_only MLP) to obtain a score for each current radar-sensing target as a real physical target. This score characterizes the probability that the corresponding current radar-sensing target is a real physical target; for example, a higher score indicates a higher probability that the corresponding current radar-sensing target is a real physical target. Furthermore, based on the scores of each current radar-sensing target as a real physical target, the current radar real target and its corresponding current radar sensing features are determined from the current radar-sensing targets. Specifically, if the score of a current radar-sensing target as a real physical target exceeds a preset real physical target score threshold, then the current radar-sensing target is determined to be a real physical target; conversely, if the score of a current radar-sensing target as a real physical target does not exceed the preset real physical target score threshold, then the current radar-sensing target is determined not to be a real physical target.
[0071] Subsequently, in one implementation, from the determined current first radar real targets, those without a matching current visual perception target are identified as supplementary first radar real targets. Based on the current first radar perception features corresponding to the supplementary first radar real targets, and the current fused perception features corresponding to the matching current visual perception target and the current first radar perception target, the perception features corresponding to the target object perceived at the current moment are determined, i.e., the current perception features corresponding to the current perception target corresponding to the target object at the current moment. Accordingly, for the supplementary first radar real targets, i.e., the unmatched current first radar real targets, their corresponding current first radar perception features are output, i.e., radar_only targets are output.
[0072] In one scenario, the current first radar sensing features corresponding to the real target of the first radar, along with the matching current visual sensing target and the current fused sensing features corresponding to the current first radar sensing target, will be supplemented and directly determined as the current sensing features corresponding to the current sensing target of the target object at the current moment.
[0073] In another embodiment of the invention, such as Figure 3 As shown, in Figure 1 Based on the process shown, the method further includes the following step S201:
[0074] S201: Obtain the historical characteristics of the target object at N historical moments before the current moment.
[0075] Where N is a positive integer, the historical features include: the fused visual features corresponding to the matching historical visual perception targets at each historical moment, the features fused and optimized with the features corresponding to the first historical radar perception targets, the optimized features corresponding to the first radar perception targets without matching visual perception targets at each historical moment, and / or the predicted features corresponding to the unperceived targets at this historical moment, which are predicted based on the perception features corresponding to the targets perceived at previous historical moments.
[0076] S202: Based on the historical characteristics of each historical target, the current perception characteristics corresponding to the current perception target, and the pre-established tracking model, determine the matching historical target and the current perception target, as well as the optimized current perception characteristics corresponding to the matching historical target and the current perception target.
[0077] The current perception features corresponding to the current perceived target include: the current fused perception features corresponding to the matching current visual perception target and the current first radar perception target, and the radar perception features corresponding to the current first radar real target without a matching current visual perception target.
[0078] In one implementation, if the current perceived target has no matching historical target, then the current perceived feature corresponding to the current perceived target that has no matching historical target is the optimized current perceived feature corresponding to the current perceived target.
[0079] The pre-established tracking model is a model trained based on the sample perception features corresponding to the sample perception targets of each sample object at each sample time, the sample history features corresponding to the sample history targets perceived by each sample object at N time points before each sample time, and the label matching information corresponding to the sample perception targets of each sample object at each sample time.
[0080] Considering the target visual feature fusion model and the pre-established radar visual feature fusion model, mismatch may occur, that is, the current visual perception target corresponding to different physical targets and the current first radar perception target are mistakenly identified as corresponding to the same physical target, and the current visual perception feature corresponding to the current visual perception target is fused together with the current first radar perception feature corresponding to the current first radar perception target, which leads to errors in the fused current perception feature.
[0081] In this embodiment of the invention, to avoid mismatches between the current visually perceived target and the current first radar-perceived target caused by the target visual feature fusion model and the pre-established radar visual feature fusion model, as well as errors in the first radar-perceived features, which could lead to abrupt changes in the position and / or velocity of the current perceived target relative to the corresponding historical target, a continuous frame tracking model, namely the pre-established tracking model TrackingNetMLP, is set up to ensure the smoothness of the target position change. This model optimizes the current perceived features corresponding to the current perceived target by using the historical features of the historical target at historical moments, resulting in optimized current perceived features corresponding to the current perceived target with relatively smoother changes.
[0082] Correspondingly, the electronic device or its connected storage device pre-stores a pre-established tracking model. This pre-established tracking model can track targets, specifically determining whether a currently perceived target and a historical target are the same physical target, and using the historical features of the historical target at a historical moment to smoothly optimize the current perceived features corresponding to the matched current perceived target. In one case, the pre-established tracking model is a convolutional network-based model with 1*N convolutional kernels. The current perceived target matching the historical target is defined as the current perceived target that corresponds to the same physical target as the historical target.
[0083] Correspondingly, the electronic device obtains the historical features of the target object perceived at N historical moments prior to the current moment, where N is a positive integer, and each of the N historical moments can correspond to at least one historical target and its historical features.
[0084] The historical features of historical targets at each historical moment can originate from three sources. The first is: for each historical moment, based on the historical visual perception features of each historical visual perception target corresponding to that historical moment and the target visual feature fusion model, the visual fusion features corresponding to each historical visual perception target at that historical moment are obtained; using the visual fusion features corresponding to each historical visual perception target at that historical moment, the first radar perception features corresponding to each historical first radar perception target at that historical moment, and a pre-established radar visual feature fusion model, the fused perception features corresponding to the matching historical visual perception target and the first radar perception target at that historical moment are determined; and using the features corresponding to targets at N moments prior to that historical moment, the fused perception features corresponding to the matching historical visual perception target and the first radar perception target at that historical moment, and a pre-established tracking model, the optimized fused perception features corresponding to the matching historical visual perception target and the first radar perception target at that historical moment are determined, i.e., the optimized features resulting from the fusion of the matching historical visual perception target's fused visual features and the historical first radar perception target's features.
[0085] The second method involves: for each historical moment, using a pre-established radar target recognition model and the first radar perception features corresponding to each first radar-sensing target at that historical moment, determining the first radar perception features of the first radar real target that did not match the historical target at that historical moment. Then, using a pre-established tracking model, the first radar perception features of the first radar real target corresponding to the historical target at that historical moment, and the features of the targets at N moments prior to that historical moment, determining the optimized features of the first radar-sensing target that did not match the visual perception target at that historical moment.
[0086] The third type is: the predicted features corresponding to the unperceived targets at the current historical moment, which are predicted based on the perceived features of the targets perceived at previous historical moments.
[0087] Accordingly, the electronic device pairs each historical target with the current perceived target. The historical features corresponding to each pair of historical targets and the current perceived features corresponding to the current perceived target are input into a pre-established tracking model. Based on the historical features of each pair of historical targets and the current perceived features of the current perceived target, the pre-established tracking model determines the affinity score (affinity_score) between each pair of historical targets and the current perceived target, as well as the optimized current perceived features corresponding to each pair of historical targets and the current perceived target. For clarity, the affinity score between each pair of historical targets and the current perceived target will be referred to as the second affinity score.
[0088] Among them, the optimized current perception features corresponding to the paired historical targets and the current perception targets include the optimized current perception features corresponding to the current perception targets that match historical targets, and the optimized current perception features corresponding to the current perception targets that do not match historical targets.
[0089] The electronic device determines the optimal matching relationship between historical targets and current sensed targets based on the Hungarian matching algorithm and the second affinity score between each pair of historical targets and current sensed targets. If the second affinity score between the optimal matching relationship between historical targets and current sensed targets exceeds a preset affinity threshold, then the historical targets and current sensed targets are determined to correspond to the same physical target, i.e., they are matched historical targets and current sensed targets. Conversely, if the second affinity score between the optimal matching relationship between historical targets and current sensed targets does not exceed the preset affinity threshold, then the historical targets and current sensed targets are determined to correspond to different physical targets, so as to achieve target tracking.
[0090] It is understandable that each historical target corresponds to N historical features. If the perceptual features of a historical target are not actually perceived at a certain historical moment, preset feature values can be used to replace the perceptual features of that historical target at that historical moment.
[0091] The label pairing information for each sample object at each sample time point includes: the sample perception target corresponding to the sample object at the sample time point, and information on whether it corresponds to the same physical target as the sample object's historical target perceived at N time points prior to the sample time point. For example, sample perception target i corresponding to sample object i at the sample time point corresponds to the same physical target as sample historical target j perceived by sample object i at N time points prior to the sample time point, and corresponds to a different physical target than other sample historical targets c perceived by sample object i at N time points prior to the sample time point; correspondingly, the label pairing information for sample perception target i includes: information representing that it corresponds to the same physical target as sample historical target j, and information representing that it corresponds to a different physical target than other sample historical targets c.
[0092] In one scenario, information indicating that the perceived target i of a sample corresponds to the same physical target as the historical target j of the sample can be represented by the ground truth value of the affinity score between the perceived target i and the historical target j of the sample. For example, the ground truth value of the affinity score is the first value, such as 1. Conversely, information indicating that the perceived target i of a sample corresponds to different physical targets from other historical targets c of the sample can be represented by the affinity score between the perceived target i of the sample and other historical targets c of the sample. For example, the ground truth value of the affinity score is the second value, such as 0.
[0093] There are two types of methods for obtaining the sample perception features corresponding to the sample perception target of each sample object at each sample time:
[0094] The first category involves obtaining the visual perception result and the first radar perception result of each sample object at each sample time. For each sample object at each historical time, based on the visual perception features of each visual perception target in the visual perception result and the target visual feature fusion model, the fused visual features corresponding to each visual perception target are determined. Each visual perception target and each first radar perception target in the first radar perception result are paired up to obtain paired visual perception targets and first radar perception targets. The fused visual features corresponding to each paired visual perception target and the first radar perception features corresponding to each first radar perception target are input into a pre-established radar visual feature fusion model to determine the affinity score between each paired visual perception target and the first radar perception target, which is used as the third affinity score, as well as the fused perception features corresponding to each paired visual perception target and the first radar perception target. Based on the third affinity score between each pair of visually perceived targets and the first radar-sensing target, and the Hungarian matching algorithm, the optimal matching relationship between the visually perceived targets and the first radar-sensing target is determined. Based on the third affinity score between the visually perceived targets and the first radar-sensing target with the optimal matching relationship and the preset affinity threshold, it is determined whether each pair of visually perceived targets and the first radar-sensing target corresponds to the same physical target.
[0095] Specifically, the sample visual perception target and the sample first radar perception target with the optimal matching relationship whose third affinity score exceeds the preset affinity threshold are identified as the sample visual perception target and the sample first radar perception target corresponding to the same physical target, that is, they are identified as the matching sample visual perception target and the sample first radar perception target. The fused perception features corresponding to the sample visual perception target and the sample first radar perception target with the optimal matching relationship are identified as the sample perception features corresponding to the sample perception target.
[0096] The second category: For each sample object at each historical moment, based on the historical radar perception results corresponding to the historical radar perception targets in the sample radar perception results and the pre-established radar target recognition model, the historical radar real targets and their corresponding historical radar perception features are determined from the historical radar perception targets. Furthermore, from the determined historical radar real targets, those without matching sample visual perception targets are identified as sample perception targets; the historical radar perception features corresponding to these unmatched sample visual perception targets are used as the sample perception features corresponding to the sample perception targets.
[0097] In another embodiment of the invention, such as Figure 2 As shown, in Figure 1 Based on the process shown, the method further includes the following steps S203-S204:
[0098] S203: Based on the historical characteristics of each historical target and the pre-established prediction model, determine the prediction characteristics of each historical target at the current moment.
[0099] The pre-established prediction model is a model trained based on the historical perception results of each sample object at N time points before each sample time, and the label matching information of each sample object at each sample time corresponding to each sample perception target.
[0100] S204: Based on the current perception characteristics of the current perception target and the predicted characteristics of each historical target at the current moment, determine the current target and its corresponding current characteristics at the current moment.
[0101] Considering that the target object may miss a target at a certain moment, in order to ensure the driving safety of the target object, the electronic device or the connected storage device can pre-store a pre-established prediction model PredictNet. This pre-established prediction model is used to predict the predicted features of each historical target at the current moment based on the historical features of the historical targets corresponding to historical moments before the current moment.
[0102] The electronic device acquires the historical characteristics of each historical target and inputs them into a pre-established prediction model. Through the pre-established prediction model, the corresponding predicted characteristics at the current moment are determined.
[0103] In one scenario, the pre-established prediction model is a convolutional network-based model with 1*N convolutional kernels used to extract features from time-series data, i.e., to extract historical features of historical targets.
[0104] Furthermore, based on the current perceived characteristics of the current perceived target and the predicted characteristics of each historical target at the current moment, the current target and its corresponding current characteristics are determined. That is, the current perceived target with existing current perceived characteristics is directly identified as the current target. And based on the predicted characteristics of each historical target at the current moment, currently unperceived targets are identified, and the predicted characteristics of the historical targets corresponding to these currently unperceived targets at the current moment are determined as the current target and its corresponding current characteristics.
[0105] In another implementation, if a target is not detected for several consecutive preset times, that is, it fails to match the historical target, the target can be destroyed, that is, the predicted features corresponding to the target can be discarded.
[0106] By using the pre-established prediction model and the pre-established tracking model, the disappearance of the target can be delayed, and the situation where a real target is not perceived for a few occasions can be resolved, thus affecting the driving decision of the target object. Furthermore, the continuity of target tracking can be improved.
[0107] In another embodiment of the present invention, before step S102, the method may further include the following step 01:
[0108] 01: Establish a target visual feature fusion model, which includes:
[0109] 011: Obtain the initial visual feature fusion model.
[0110] 012: Obtain the label perception features of each sample object at each sample time and the sample visual perception features corresponding to the visual perception target of each sample.
[0111] Among them, the label perception features are: the label perception features perceived by the sample label radar of the corresponding sample object at the corresponding sample time, and the sample visual perception features corresponding to the sample visual perception target are: the visual perception features perceived by the sample image acquisition device group of the corresponding sample object at the corresponding sample time.
[0112] 013: Using the visual perception features and label perception features corresponding to the visual perception target of each sample, train the initial visual feature fusion model until the initial visual feature fusion model reaches the first convergence condition, and determine the target visual feature fusion model.
[0113] This implementation provides a process for establishing a target visual feature fusion model. The electronic device can first obtain an initial visual feature fusion model, which can be a model based on a deep learning algorithm.
[0114] The electronic device acquires the tag-perceived features of each sample object at each sample time and the sample visual perception features corresponding to the visual perception targets of each sample. Specifically, the sample visual perception features corresponding to the visual perception targets are the visual perception features of the visual perception targets perceived by the sample image acquisition device of each sample object at each sample time. The tag-perceived features are the perception features of the sample tag radar of each sample object corresponding to the target perceived by the sample tag radar at each sample time.
[0115] The visual perception features corresponding to each sample visual perception target include: the visual perception features perceived by each sample image acquisition device of the corresponding sample object at the sample time. If a sample image acquisition device does not perceive a sample visual perception target at the sample time, the visual perception feature corresponding to that sample visual perception target of that sample image acquisition device at the sample time is replaced by a preset feature value.
[0116] There is a correspondence between the label perception features of a certain sample object at a certain sample time and the sample visual perception features of the visual perception targets of each sample.
[0117] The electronic device randomly inputs the tag perception features of the radar-sensing targets corresponding to each sample object at each sample time, and the sample visual perception features corresponding to each sample visual perception target, into the current visual feature fusion model. The current visual feature fusion model fuses the sample visual perception features corresponding to each sample visual perception target at each sample time to obtain the fused visual features of each sample visual perception target corresponding to each sample object at each sample time. The electronic device then uses the fused visual features of each sample visual perception target corresponding to each sample object at each sample time, the tag perception features of the radar-sensing targets corresponding to each sample object at each sample time, and the first preset loss function to calculate the corresponding loss value, which is used as the first loss value.
[0118] The current visual feature fusion model is either the initial visual feature fusion model or the initial visual feature fusion model after adjusting the model parameter values.
[0119] The corresponding sample tag radar-sensing targets and sample visual-sensing targets are: the sample tag radar-sensing targets and sample visual-sensing targets perceived by a certain sample object at a certain sample time for the same physical target.
[0120] The process involves determining whether the first loss value is less than a first loss threshold. If the first loss value is less than the first loss threshold, the current visual feature fusion model is considered to have reached the first convergence condition, and the target visual feature fusion model is determined. If the first loss value is not less than the first loss threshold, the model parameters of the current visual feature fusion model are adjusted, and the process of randomly inputting the label perception features corresponding to each sample object at each sample time and the visual perception features corresponding to each sample's visual perception target into the current visual feature fusion model is repeated until the calculated first loss value is less than the first loss threshold. At this point, the current initial visual feature fusion model is considered to have reached the first convergence condition, and the target visual feature fusion model is determined.
[0121] The first preset loss function can be the Smooth L1 loss function.
[0122] In another embodiment of the present invention, before step S103, the method may further include the following step 02:
[0123] 02: Establish a pre-built radar visual feature fusion model, which includes:
[0124] 021: Obtain the initial radar visual feature fusion model.
[0125] 022: Obtain the first radar sensing features of each sample object at each sample time corresponding to the first radar sensing target of each sample.
[0126] Among them, the first radar sensing features corresponding to the first radar sensing targets of each sample are: the first radar sensing features of the corresponding sample object at the corresponding sample time.
[0127] 023: Based on the visual fusion features of each sample object corresponding to the visual perception target of each sample at each sample time, the first radar perception features of each sample object corresponding to the first radar perception target of each sample at each sample time, the label matching information between each sample visual perception target and each sample first radar perception target, and the label perception features of each sample object corresponding to each sample time, train the initial radar visual feature fusion model until the initial radar visual feature fusion model reaches the second convergence condition, and determine the pre-established radar visual feature fusion model.
[0128] Among them, the sample visual fusion feature corresponding to the sample visual perception target is: the fusion feature obtained by fusing the sample visual perception feature corresponding to the sample visual perception target and the target visual feature fusion model.
[0129] The label matching information between each sample visually perceived target and each sample first radar perceived target includes: information characterizing whether the sample visually perceived target and the sample first radar perceived target correspond to the same physical target. For example, sample visually perceived target a and sample first radar perceived target b correspond to the same physical target, while sample visually perceived target a and sample first radar perceived target c correspond to different physical targets. In one case, the label matching information can be represented by the affinity score ground truth. If sample visually perceived target a and sample first radar perceived target b correspond to the same physical target, their corresponding affinity score ground truth can be set to a first value, such as 1. Conversely, if sample visually perceived target a and sample first radar perceived target c correspond to different physical targets, their corresponding affinity score ground truth can be set to a second value, such as 0.
[0130] This implementation provides the process for establishing and training a pre-built radar visual feature fusion model. The electronic device obtains the initial radar visual feature fusion model. In one case, this initial radar visual feature fusion model is a model based on a deep learning algorithm.
[0131] The electronic device obtains the sample first radar perception features corresponding to each sample object and each sample first radar perception target at each sample time. For each sample object and the perception target corresponding to each sample time, the sample visual perception target and the sample first radar perception target corresponding to the sample object at each sample time are paired up to obtain paired sample visual perception targets and sample first radar perception targets. The sample visual fusion features corresponding to the paired sample visual perception targets and the sample first radar perception features corresponding to the sample first radar perception targets, the label matching information between the paired sample visual perception targets and the sample first radar perception targets, and the label perception features of the sample visual perception target corresponding to the sample object at each sample time are randomly input into the current radar visual feature fusion model.
[0132] Based on the sample visual fusion features corresponding to each pair of sample visual perception targets and the sample first radar perception features corresponding to each pair of sample visual perception targets, the current radar visual feature fusion model determines the affinity score between each pair of sample visual perception targets and the sample first radar perception target as the fourth affinity score; and the fusion perception features corresponding to each pair of sample visual perception targets and the sample first radar perception target are used as training fusion perception features.
[0133] By utilizing the tag matching information between pairwise paired visual perception targets and sample first radar perception targets, the visual perception targets and sample first radar perception targets that truly correspond to the same physical target are identified from each pairwise paired visual perception target and sample first radar perception target, and these are designated as truly matched visual perception targets and sample first radar perception targets. Pairs of visual perception targets and sample first radar perception targets that do not correspond to the same physical target are designated as unmatched visual perception targets and sample first radar perception targets.
[0134] Using the first regression loss function, the first classification loss function, the fourth affinity scores corresponding to the visually perceived target of the true matching sample and the radar-sensed target of the sample, the training fusion perception features, and the fourth affinity scores corresponding to the visually perceived target of the mismatched sample and the radar-sensed target of the sample, the corresponding loss values are calculated and used as the second loss values.
[0135] In one implementation, the first regression loss function can be the Smooth L1 loss function, and the first classification loss function can be the BCE (Binary Cross Entropy) Loss loss function.
[0136] Specifically, for the visually perceived target and the radar-sensed target of the real matching sample, the corresponding loss values include: the classification loss value calculated using the corresponding fourth affinity score, label matching information and the first classification loss function, and the loss value calculated using the corresponding training fusion sensing features, label sensing features and the first regression loss function.
[0137] For mismatched visual perception targets and first radar perception targets, the corresponding loss values include: the loss values calculated using their corresponding fourth affinity scores and the first classification loss function.
[0138] Furthermore, a second loss value is determined using the loss values corresponding to the visually perceived target and the first radar-sensed target in the truly matched samples, and the loss values corresponding to the visually perceived target and the first radar-sensed target in the mismatched samples. Specifically, the second loss value can be the sum of the loss values corresponding to the visually perceived target and the first radar-sensed target in the truly matched samples and the mismatched samples.
[0139] Determine whether the second loss value is less than the second loss threshold; if the second loss value is less than the second loss threshold, the current radar visual feature fusion model can be considered to have reached the second convergence condition, and the current radar visual feature fusion model is determined as the pre-established radar visual feature fusion model.
[0140] If the second loss value is not less than the second loss threshold, adjust the model parameter values of the current radar visual feature fusion model, and return to the step of randomly inputting the sample visual fusion features corresponding to the paired sample visual perception targets, the sample first radar perception features corresponding to the sample first radar perception target, the label matching information between the paired sample visual perception targets and the sample first radar perception target, and the label perception features of the sample object corresponding to the sample visual perception target at the sample time into the current radar visual feature fusion model.
[0141] The current radar visual feature fusion model is either the initial radar visual feature fusion model or the radar visual feature fusion model after adjusting the model parameter values.
[0142] In another embodiment of the present invention, before step S104, the method may further include the following step 03:
[0143] 03: Establish a pre-built radar target recognition model, which includes:
[0144] 031: Obtain the initial radar target recognition model.
[0145] 032: Obtain the authenticity information of the tag corresponding to the first radar-sensing target of each sample object at each sample time.
[0146] The label authenticity information is information identifying whether the corresponding sample first radar-sensing target is a genuine first radar-sensing target. In one scenario, the label authenticity information can be represented by a truth value for the authenticity label. When the sample first radar-sensing target is a genuine first radar-sensing target, its corresponding target authenticity label truth value is the third value, for example, 1; when the sample first radar-sensing target is not a genuine first radar-sensing target, its corresponding target authenticity label truth value is the fourth value, for example, 0.
[0147] 033: Based on the radar sensing features of each sample object at each sample time corresponding to the first radar sensing target of each sample, and the authenticity information of its corresponding label, train the initial radar target recognition model until the initial radar target recognition model reaches the third convergence condition, and determine the pre-established radar target recognition model.
[0148] This implementation provides a process for establishing a pre-built radar target recognition model, allowing the electronic device to first obtain the initial radar target recognition model. It also obtains the authenticity information of the label corresponding to the first radar-sensing target for each sample object at each sample time. In one scenario, the initial radar target recognition model can be a model based on a deep learning algorithm.
[0149] For each sample object at each sample time corresponding to the first radar sensing target, the electronic device randomly inputs the sample first radar sensing feature corresponding to the first radar sensing target of the sample object at each sample time, and its corresponding label authenticity information into the current radar target recognition model to obtain the predicted authenticity information corresponding to the sample first radar sensing feature. If the label authenticity information is 1 or 0, the corresponding predicted authenticity information is between 0 and 1.
[0150] Using the second classification loss function, the predicted authenticity information and label authenticity information corresponding to the first radar sensing feature of the sample, the current loss value is calculated and used as the third loss value. It is then determined whether the third loss value is less than the third loss threshold. If the third loss value is less than the third loss threshold, the current radar target recognition model is determined to have reached the third convergence condition, and the current radar target recognition model is identified as the pre-established radar target recognition model. If the third loss value is not less than the third loss threshold, the model parameters of the current radar target recognition model are adjusted, and the process is repeated. For each sample object at each sample time, the first radar sensing feature corresponding to the first radar sensing target of that sample object at that sample time, along with its corresponding label authenticity information, is randomly input into the current radar target recognition model to obtain the predicted authenticity information corresponding to the first radar sensing feature of that sample. This process continues until the calculated third loss value is less than the third loss threshold, at which point the current radar target recognition model is determined to have reached the third convergence condition, and the current radar target recognition model is identified as the pre-established radar target recognition model.
[0151] The second classification loss function can be the BCELoss loss function.
[0152] In another embodiment of the present invention, before step S202, the method may further include the following step 04:
[0153] 04: The process of establishing a pre-built tracking model, which includes:
[0154] 041: Obtain the initial tracking model.
[0155] 042: For each sample object, obtain the sample history features corresponding to the sample historical target at N time points before each sample time.
[0156] The sample history features include: the visual fusion features corresponding to the matching sample visual perception targets at N time points before each sample time point, and the features corresponding to the first radar perception target of the sample after fusion and optimization; the optimized features corresponding to the first radar perception target of the sample at N time points before each sample time point without matching sample visual perception targets; and / or the predicted features corresponding to the unperceived targets at the current sample time point, predicted based on the perception features corresponding to the historical targets of the samples at N time points before each sample time point.
[0157] 043: For each sample object at each sample time, obtain the label pairing information corresponding to the sample perception target corresponding to the sample object at each sample time.
[0158] The label pairing information corresponding to the target perceived by the sample includes: information characterizing whether the target perceived by the sample and the historical target perceived by the sample object at N times before the sample time are the same physical target.
[0159] 044: Based on the sample perception features of each sample object at each sample time corresponding to the sample perception target, the sample history features of each sample object at N time steps before each sample time corresponding to the sample history target, the label pairing information of each sample object at each sample time corresponding to each sample perception target, and the label perception features of each sample object at each sample time, train the initial tracking model until the initial tracking model reaches the fourth convergence condition, and determine the pre-established tracking model.
[0160] Among them, the sample perception features include: the sample fusion perception features corresponding to the matching sample visual perception target and the sample first radar perception target at each sample time, and / or the first radar perception features corresponding to the sample first radar real target without matching sample visual perception targets, determined based on the sample radar perception features corresponding to the sample first radar perception target and the pre-established radar target recognition model.
[0161] This implementation provides a process for establishing a pre-built tracking model. The electronic device obtains an initial tracking model, which is a convolutional network-based model with 1*N convolutional kernels. These 1*N convolutional kernels are used to process the N historical features corresponding to the historical target of each sample object at each sample time step.
[0162] For each sample object, obtain the sample history features corresponding to the sample history targets at N time points prior to each sample time. These sample history features include: the sample history features corresponding to the sample history targets at time point T-1, time point T-2, and so on, as well as the sample history features corresponding to the sample history targets at each time point TN.
[0163] The sample history features include: the visual fusion features corresponding to the matching sample visual perception targets at N time points before each sample time point, and the features corresponding to the first radar perception target of the sample after fusion and optimization; the optimized features corresponding to the first radar perception target of the sample at N time points before each sample time point without matching sample visual perception targets; and / or the predicted features corresponding to the unperceived targets at the current sample time point, predicted based on the perception features corresponding to the historical targets of the samples at N time points before each sample time point.
[0164] The matching sample visual perception target and the sample first radar perception target represent: the visual perception target and the first radar perception target perceived by the image acquisition device and the first radar of the same sample object at the same time for the same physical target.
[0165] The visual fusion feature corresponding to the visual perception target of the sample is: the feature obtained by fusing the visual perception features perceived by the image acquisition device of the sample object at a certain sample time for the visual perception target of the sample.
[0166] For each sample object at each sample time, obtain the label pairing information corresponding to the sample perception target corresponding to the sample object at each sample time. The label pairing information corresponding to the sample perception target includes: information representing whether the sample perception target and the sample historical target corresponding to the sample object at N times before the sample time correspond to the same physical target.
[0167] The label pairing information corresponding to the sample perception target includes: the sample perception target corresponding to the sample object at the current sample time, and the ground truth value of the affinity score between the sample object and the sample historical targets corresponding to the sample object at N time points prior to the current sample time. When the sample perception target and the sample historical targets correspond to the same physical target, the ground truth value of the affinity score between the sample perception target and the sample historical targets is identified as a first value, such as 1. When the sample perception target and the sample historical targets correspond to different physical targets, the ground truth value of the affinity score between the sample perception target and the sample historical targets is identified as a second value, such as 0.
[0168] For each sample object at each sample time, the electronic device pairs each sample perception target corresponding to the sample object at that sample time with the sample historical targets corresponding to the sample object at N times prior to that sample time, obtaining paired sample perception targets and sample historical targets. The sample perception features corresponding to the paired sample perception targets, the sample historical features corresponding to the sample historical targets, the label pairing information corresponding to the sample perception targets, and the label perception features corresponding to the sample object at that sample time are randomly input into the current tracking model. The current tracking model is either the initial tracking model or the tracking model after adjusting the model parameter values.
[0169] The current tracking model uses the sample perception features corresponding to the paired sample perception targets and the sample history features corresponding to the sample history targets to obtain the affinity score between the paired sample perception targets and the sample history targets, which is the fifth affinity score, also known as the matching affinity mentioned later. The optimized sample perception features corresponding to the sample perception targets are obtained by using the sample history features corresponding to the paired sample history targets as the optimized sample perception features corresponding to the sample perception targets in each paired sample history target and sample perception target.
[0170] When the number of fifth affinity scores obtained reaches the preset batch size, the current loss value is calculated as the fourth loss value based on the third regression loss function, the third classification loss function, the fifth affinity scores between the perceived target and the historical target of the sample in the preset batch size, the ground truth values of the affinity scores between the perceived target and the historical target of the sample in the label pairing information, the optimized perceived features corresponding to the perceived targets in the historical targets and perceived targets of the sample in the pairing, and the label perceived features corresponding to the perceived targets in the historical targets and perceived targets of the sample in the pairing, and the label perceived features corresponding to the perceived targets in the historical targets and perceived targets of the sample in the pairing: these are the label perceived features corresponding to each sample object at each sample time.
[0171] It is understandable that there may be a pairing between the sample perception target and the sample history target corresponding to the same sample object at the same sample time. The sample perception target and sample history target corresponding to the same sample object at the same sample time are: the sample perception target corresponding to the same sample object at a sample time, and the sample history target corresponding to the same sample object at N time points before that sample time.
[0172] First, if the true value of the affinity score between the sample perceived target and the sample historical target in the label pairing information is the first value, the regression loss value corresponding to the sample historical target and the sample perceived target in the pairing is calculated based on the third regression loss function, the optimized sample perceived features corresponding to the sample perceived target in the pairing sample historical target and sample perceived target, and the label perceived features.
[0173] Based on the third classification loss function, the fifth affinity score between the paired sample perceived target and the sample historical target, and the true value of the affinity score between the paired sample perceived target and the sample historical target in the label pairing information, the classification loss value corresponding to the paired sample perceived target and the sample historical target is calculated.
[0174] Based on the regression loss value corresponding to the paired sample perception target and sample history target, and the classification loss value corresponding to the paired sample perception target and sample history target, the loss value corresponding to the paired sample perception target and sample history target is determined as part of the fourth loss value.
[0175] Second, if the true value of the affinity score between the sample perceived target and the sample historical target in the label pairing information is the second value, the classification loss value corresponding to the sample perceived target and the sample historical target in the label pairing information is calculated based on the third classification loss function, the fifth affinity score between the sample perceived target and the sample historical target in the pairing information, and the true value of the affinity score between the sample perceived target and the sample historical target in the label pairing information. This loss value is used as the loss value corresponding to the sample perceived target and the sample historical target in the pairing information, and is used as part of the fourth loss value.
[0176] In one scenario, the fourth loss value is determined by summing the loss values corresponding to the perceived target and historical target of the paired samples with the true value of the label pairing information affinity score being the first value, and the loss values corresponding to the perceived target and historical target of the paired samples with the true value of the label pairing information affinity score being the second value.
[0177] The system checks if the fourth loss value is less than the fourth loss threshold. If it is, the current tracking model is considered to have met the fourth convergence condition and is designated as the pre-established tracking model. Otherwise, if the fourth loss value is not less than the fourth loss threshold, the system adjusts the model parameters and returns to the previous step of inputting the following steps into the current tracking model: For each sample object at each sample time, the system randomly selects the sample perception features corresponding to each sample perception target at that sample time, the sample history features corresponding to the sample historical targets at N time points prior to that sample time, the label pairing information corresponding to each sample perception target at that sample time, and the label perception features corresponding to that sample object at that sample time. This process continues until the fourth loss value is less than the fourth loss threshold, at which point the current tracking model is considered to have met the fourth convergence condition and is designated as the pre-established tracking model.
[0178] The third regression loss value, also known as the preset regression loss function (mentioned later), can be the Smooth L1 loss function. The third classification loss value, also known as the preset classification loss function (mentioned later), can be the BCE Loss function.
[0179] To avoid mismatches between the aforementioned target visual feature fusion model and the pre-established radar visual feature fusion model, as well as errors in the first radar sensing features, which could cause jumps in the position or velocity of the sensing features corresponding to each sensing target, after the tracking model has been trained normally (i.e., trained to meet the fourth convergence condition), a preset smoothing loss function is added to fine-tune the model parameters of the tracking model that has met the fourth convergence condition. This makes the sensing features corresponding to the sensing target more smoothly change relative to the historical features of the corresponding historical target. In another embodiment of the present invention, step 045 may include the following steps 0451-0455:
[0180] 0451: For each sample object at each sample time, pair each sample perception target corresponding to the sample object at that sample time with the sample history targets corresponding to the sample object at N times before that sample time to obtain the pairwise paired sample perception targets and sample history targets corresponding to the sample object at that sample time.
[0181] 0452: Randomly input the sample perception features corresponding to the paired sample perception targets and the sample historical features corresponding to the paired sample historical targets at each sample time, the label pairing information and label perception features corresponding to the paired sample perception targets and sample historical targets into the current tracking model to obtain the matching affinity between the paired sample perception targets and sample historical targets and the optimized sample perception features corresponding to the sample perception targets.
[0182] Among them, the optimized sample perception feature corresponding to the sample perception target is: the sample perception feature corresponding to the sample perception target, and the optimized feature based on the sample historical feature corresponding to the sample historical target paired with the sample perception target.
[0183] The matching affinity between paired sample perceived targets and sample historical targets represents the probability that each paired sample perceived target and sample historical target is the same physical target. The current tracking model is either the initial tracking model or a tracking model with adjusted parameters. In one scenario, the higher the matching affinity between paired sample perceived targets and sample historical targets, the greater the probability that each paired sample perceived target and sample historical target is the same physical target, and consequently, the greater the likelihood that they are the same physical target.
[0184] 0453: When the number of matching affinity values corresponding to the obtained sample perception target reaches the preset batch size, the current loss value corresponding to the current tracking model is determined based on the matching affinity between the pairwise paired sample perception targets and the sample historical targets of the preset batch size, the corresponding label pairing information of the sample perception targets in the pairwise paired sample perception targets and sample historical targets, the optimized sample perception features and label perception features corresponding to the sample perception target, the preset regression loss function, and the preset classification loss function.
[0185] 0454: If the current loss value is less than the preset loss threshold, then the current tracking model is determined to have reached the fourth convergence condition, and an intermediate tracking model is obtained.
[0186] 0455: Using a preset smoothness loss function, a preset regression loss function, a preset classification loss function, the sample perception features corresponding to the paired sample perception targets and the sample history features corresponding to the sample history targets at each sample time, the label pairing information and label perception features corresponding to the sample perception targets in the paired sample perception targets and sample history targets, and the specified sample history features corresponding to the sample history targets in the paired sample perception targets and sample history targets, adjust the model parameters of the intermediate tracking model until the intermediate tracking model reaches the fifth convergence condition, and obtain the pre-established tracking model.
[0187] Among them, the specified sample history features corresponding to the sample history targets in the pairwise paired sample perception targets and sample history targets are: the sample history targets in the pairwise paired sample perception targets and sample history targets, and the sample history features of the sample history targets at the time preceding the sample time corresponding to the sample perception target.
[0188] 0456: If the current loss value is not less than the preset loss threshold, adjust the values of the model parameters of the tracking model and return to 0451 until the current tracking model reaches the fourth convergence condition to obtain the intermediate tracking model.
[0189] The preset smoothness loss function can be expressed by the following formula:
[0190]
[0191] Where m represents the preset batch size, λ is a hyperparameter used to adjust the loss magnitude, and obj_tracking represents the optimized sample-aware features corresponding to the pairwise paired sample-aware targets and sample-aware targets in the historical sample targets within the preset batch size in the output of the current intermediate tracking model; obj_fusion[j] t1This represents the specified historical features corresponding to the sample historical targets in a pairwise paired sample perception target and sample historical target. Specifically, the specified historical features corresponding to the sample historical targets in a pairwise paired sample perception target and sample historical target are: the historical features of the sample historical target at the time preceding the sample time corresponding to that sample perception target. For example, if the sample time corresponding to the sample perception target in a pairwise paired sample perception target and sample historical target is sample time E, then the specified historical features corresponding to the sample historical target in a pairwise paired sample perception target and sample historical target are: the historical features of the sample historical target at the time preceding sample time E.
[0192] In this embodiment of the invention, after obtaining the intermediate tracking model, the electronic device obtains the pairwise paired sample perception targets and sample historical targets corresponding to each sample object at each sample time. The sample perception features corresponding to the pairwise paired sample perception targets and the sample historical features corresponding to the sample historical targets, the label pairing information corresponding to the sample perception targets in the pairwise paired sample perception targets and sample historical targets, and the specified sample historical features corresponding to the sample historical targets in the pairwise paired sample perception targets and sample historical targets are randomly input into the current intermediate tracking model. The current intermediate tracking model is the initially obtained intermediate tracking model or the intermediate tracking model after adjusting the model parameter values.
[0193] Based on the sample perception features corresponding to the paired sample perception targets and the sample history features corresponding to the sample history targets, the current intermediate tracking model obtains the affinity score between the paired sample perception targets and the sample history targets, which is used as the sixth affinity score; and obtains the optimized sample perception features corresponding to the sample perception targets in the paired sample perception targets and sample history targets.
[0194] When the number of sixth affinity scores obtained reaches the preset batch size, the first part of the loss value is calculated based on the third regression loss function, the third classification loss function, the sixth affinity scores between the sample perception target and the sample history target corresponding to the pairwise paired sample objects at the preset batch size, the true values of the affinity scores between the pairwise paired sample perception targets and the sample history targets in the label pairing information, and the optimized sample perception features and label perception features corresponding to the sample perception targets in the pairwise paired sample perception targets and sample history targets.
[0195] The process of calculating the first part of the loss value can be found in the process of calculating the fourth loss value described above, and will not be repeated here.
[0196] Furthermore, when the number of obtained sixth affinity scores reaches the preset batch size, the current second part of the loss value is calculated based on the preset smoothness loss function, the optimized sample perception features corresponding to the sample perception targets and sample historical targets in the preset batch size corresponding to the sample objects at the sample time, and the specified sample historical features corresponding to the sample historical targets in the preset batch size corresponding to the sample perception targets and sample historical targets.
[0197] Using the first and second loss values, the current loss value corresponding to the current intermediate tracking model is determined and used as the fifth loss value.
[0198] The system determines whether the fifth loss value is less than the fifth loss threshold. If the fifth loss value is less than the fifth loss threshold, the current intermediate tracking model is considered to have met the fifth convergence condition and is designated as the pre-established tracking model. Conversely, if the fifth loss value is not less than the fifth loss threshold, the system adjusts the model parameters of the current intermediate tracking model and returns to the previous step of randomly inputting the sample perception features corresponding to each pair of paired sample perception targets and the sample history features corresponding to each pair of paired sample perception targets and sample history targets, along with the label pairing information and label perception features corresponding to each pair of paired sample perception targets and sample history targets, into the current intermediate tracking model. This process continues until the fifth loss value is less than the fifth loss threshold, at which point the current intermediate tracking model is considered to have met the fifth convergence condition and is designated as the pre-established tracking model.
[0199] In another embodiment of the present invention, before step S203, the method may further include the following step 05:
[0200] 05: The process of establishing a pre-built predictive model, which includes:
[0201] 051: Obtain the initial prediction model.
[0202] 052: Based on the sample history features and label perception features of each sample object at N time points before each sample time, train the initial prediction model until the initial prediction model reaches the sixth convergence condition, and determine the pre-established prediction model.
[0203] This implementation provides a process for establishing a pre-built prediction model. The electronic device obtains an initial prediction model, which is a convolutional network-based model with 1*N convolutional kernels, to process the historical features of the target sample corresponding to the sample object at N time points prior to each sample time.
[0204] The electronic device, for each sample object's perceived historical targets at N time points prior to each sample time, randomly inputs the historical features corresponding to those historical targets perceived at those N time points prior to the current sample time, along with the label-aware features corresponding to those historical targets from the label-aware features, into the current prediction model to obtain the predicted features corresponding to those historical targets. The current prediction model can be either the initial prediction model or a prediction model with adjusted parameters.
[0205] The label perception features corresponding to the historical targets of the samples are: the label perception features of the matching sample label radar perception targets corresponding to the sample objects corresponding to the historical targets of the samples at their corresponding sample times. The matching sample label radar perception targets are the perception targets that correspond to the same physical target as the historical targets of the samples.
[0206] After the electronic device obtains a second batch of predicted features, it calculates the current loss value as the fourth loss value using the fourth regression loss function, the predicted features corresponding to the historical targets of the second batch of samples, and the label-aware features. It then checks if the fourth loss value is less than the fourth loss threshold. If it is, the current prediction model is considered to have met the sixth convergence condition and is designated as the pre-established prediction model. If the fourth loss value is not less than the fourth loss threshold, the model parameters of the current prediction model are adjusted, and the process returns to randomly selecting the historical features corresponding to the historical targets of the sample object at N time points prior to the current sample time, along with the label-aware features corresponding to the historical targets of the sample object. This process is repeated until the fourth loss value is less than the fourth loss threshold, at which point the current prediction model is considered to have met the sixth convergence condition and is designated as the pre-established prediction model.
[0207] In one implementation, the aforementioned target visual feature fusion model, the pre-established visual radar feature fusion model, and the pre-established radar target recognition model can be collectively referred to as a single-frame fusion model, relative to the pre-established tracking model and the pre-established prediction model. Correspondingly, the pre-established tracking model and the pre-established prediction model can be referred to as a continuous-frame tracking model.
[0208] like Figure 5A The diagram shown illustrates the structure and data flow of a single-frame fusion model. Figure 5AAs shown, the single-frame fusion model includes a target visual feature fusion model "vision_fusion MLP", a pre-established visual radar feature fusion model "vision_radar_MLP", and a pre-established radar target recognition model "radar_only_MLP". The input of the target visual feature fusion model is the visual perception features corresponding to each visually perceived target at time t0, such as... Figure 5A The output of the target visual feature fusion model, as shown in "obj_cam1", "obj_cam2", etc., is the fused visual perceptual feature corresponding to each visually perceived target. Figure 5A The “obj_vision_fusion” shown.
[0209] The inputs to the pre-established visual radar feature fusion model "vision_radar_MLP" are: the output of the target visual feature fusion model, the fused visual perception features "obj_vision_fusion" corresponding to each visually perceived target, and the first radar perception features corresponding to each first radar perceived target at time t0, such as... Figure 5A The “radar” shown refers to the fused visual perception features corresponding to each visually perceived target and the first radar perception features corresponding to each first radar perception target, which are inputs in the form of each pair of visually perceived targets and first radar perception targets.
[0210] The output of the pre-established visual radar feature fusion model consists of the fused perception features "obj_fusion" and the affinity score "affinity_score" corresponding to each pair of visually perceived targets and the first radar perceived target. Then, based on the affinity scores "affinity_score" of each pair of visually perceived targets and the first radar perceived target, the Hungarian matching algorithm, and a preset affinity threshold, the electronic device determines the matching visually perceived target and the first radar perceived target, along with their corresponding fused perception features.
[0211] The input to the pre-established radar target recognition model "radar_only_MLP" is the first radar sensing feature corresponding to each first radar sensing target at time t0; the output is the score for each first radar sensing target as a real physical target. Figure 5A The “radar_only_score” shown in the figure refers to the subsequent determination of the real targets of the first radar and their corresponding first radar sensing features based on the scores of each first radar sensing target that are real physical targets and the scores that exceed the preset real physical target score threshold.
[0212] like Figure 5BThe diagram shown illustrates the structure and data flow of a continuous frame tracking model. Figure 5B As shown, the continuous frame tracking model includes a pre-established tracking model "Tracking Net" and a pre-established prediction model "Predict Net". The input of the pre-established tracking model "Tracking Net" is the perception feature corresponding to the current perceived target at time t0, including the fused perception feature corresponding to the matching visual perceived target and the first radar perceived target at time t0, as shown in obj_fusion[j]t0 in Figure 5, and the first radar perception feature corresponding to the first radar real target at time t0; and the historical features corresponding to the historical targets at each of the N historical times t0-1, t0-2...t0-N before time t0, as shown in obj_fusion[j]t1-obj_fusion[j]tn in Figure 5. Among them, the historical features corresponding to each historical target include: the fused visual features corresponding to the matching visual perception target at the historical moment, the features fused with the features corresponding to the radar perception target, the features corresponding to the first radar perception target that has no matching visual perception target in the historical radar perception results at each historical moment, and / or the predicted features corresponding to the unperceived target at the current historical moment, which are predicted based on the perception features corresponding to the target perceived at a previous time.
[0213] The output of the pre-built tracking model "Tracking Net" is the tracking result at time t0, as shown in Figure 5 as obj_tracking, and the affinity score "affinity_score" between the current perceived target and the historical target for each pair of matching targets. Among them, the tracking result at time t0 includes the optimized perception features corresponding to the successfully matched current perceived target.
[0214] The pre-established prediction model "Predict Net" takes as input the historical features corresponding to the historical targets at N historical times t0-1, t0-2...t0-N before time t0, as shown in Figure 5 (obj_fusion[j]t1-obj_fusion[j]tn); and outputs the predicted features of each historical target at time t0. Subsequently, from the predicted features of each historical target at time t0, the predicted features corresponding to the current perceived target without a matching historical target are determined, i.e., "obj_predict" as shown in Figure 5. Combined with the optimized perceived features of the currently perceived target that has a successful match, the current features corresponding to the current target are determined.
[0215] Corresponding to the above method embodiments, this invention provides an apparatus, such as... Figure 6As shown, the device may include:
[0216] The first acquisition module 610 is configured to acquire the current visual perception result and the current first radar perception result of the target object at the current moment.
[0217] The first determining module 620 is configured to determine the fused visual features corresponding to each current visual perception target based on the current visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model. The target visual feature fusion model is a model trained based on the label perception features corresponding to each sample object at each sample time and the sample visual perception features corresponding to each sample visual perception target.
[0218] The second determining module 630 is configured to determine the matching current visual perception target and the current first radar perception target, as well as the corresponding current fused perception features, based on the fused visual features corresponding to each current visual perception target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and a pre-established radar visual feature fusion model. The pre-established radar visual feature fusion model is a model trained based on the sample visual perception features, the label perception features, and the sample first radar perception results corresponding to each sample first radar perception target at each sample time.
[0219] By applying the embodiments of the present invention, the visual perception features corresponding to each current visual perception target in the current visual perception result can be fused based on the pre-trained target visual feature fusion model to obtain the fused visual features corresponding to each current visual perception target. Then, by using the pre-established radar visual feature fusion model, the fused visual features corresponding to the current visual perception target that matches the same physical target and the current first radar perception feature corresponding to the current first radar perception target can be fused to obtain relatively accurate perception features of each physical target. By using the target visual feature fusion model and the pre-established radar visual feature fusion model, accurate fusion of multi-sensor data can be achieved.
[0220] In another embodiment of the present invention, the device further includes:
[0221] The third determining module (not shown in the figure) is configured to determine the current first radar real target and its corresponding current first radar perception features from the current first radar perception targets based on the current first radar perception features corresponding to the current first radar perception target and a pre-established radar target recognition model. The pre-established radar target recognition model is a model trained based on the sample first radar perception features corresponding to each sample first radar perception target and the label perception features.
[0222] In another embodiment of the present invention, the device further includes:
[0223] The second acquisition module (not shown in the figure) is configured to acquire historical features of the target object corresponding to N historical moments before the current moment, where N is a positive integer. The historical features include: the fused visual features corresponding to the matching historical visual perception targets at each historical moment, and the features fused and optimized with the features corresponding to the historical first radar perception targets, the optimized features corresponding to the first radar perception targets that do not have matching visual perception targets in the historical radar perception results at each historical moment, and / or the predicted features corresponding to the targets not perceived at the current historical moment, which are predicted based on the perception features corresponding to the targets perceived at moments before each historical moment.
[0224] The fourth determining module (not shown in the figure) is configured to determine the matching historical targets and current sensing targets, as well as the optimized current sensing features corresponding to the matching historical targets, based on the historical features of each historical target, the current sensing features corresponding to the current sensing target, and the pre-established tracking model. The current sensing features corresponding to the current sensing target include: the current fused sensing features corresponding to the matching current visual sensing target and the current first radar sensing target, and / or the radar sensing features corresponding to the current first radar real target without a matching current visual sensing target. The pre-established tracking model is a model trained based on the sample sensing features corresponding to the sample sensing targets of each sample object at each sample time, the sample historical features corresponding to the sample historical targets of each sample object at N time steps before each sample time, and the label matching information corresponding to the sample sensing targets of each sample object at each sample time.
[0225] In another embodiment of the present invention, the device further includes:
[0226] The fifth determining module (not shown in the figure) is configured to determine the predicted features of each historical target at the current time based on the historical features of each historical target and the pre-established prediction model. The pre-established prediction model is a model trained based on the sample historical perception results of each sample object at N times before each sample time and the label matching information of each sample object at each sample time corresponding to each sample perception target.
[0227] The sixth determining module (not shown in the figure) is configured to determine the current target and its corresponding current features at the current time based on the current sensing features of the current sensing target and the predicted features of each historical target at the current time.
[0228] In another embodiment of the present invention, the device further includes:
[0229] The first model building module (not shown in the figure) is configured to establish the target visual feature fusion model before determining the fused visual features corresponding to each current visual perception target based on the visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model. Specifically, the first model building module is configured to...
[0230] Obtain the initial visual feature fusion model;
[0231] Obtain the label perception features of each sample object at each sample time and the sample visual perception features of each sample visual perception target. The label perception features are: the label perception features of the sample label radar of the corresponding sample object at the corresponding sample time. The sample visual perception features of each sample visual perception target are: the visual perception features of the sample image acquisition device group of the corresponding sample object at the corresponding sample time.
[0232] Using the visual perception features of each sample corresponding to the visual perception target and the label perception features, the initial visual feature fusion model is trained until the initial visual feature fusion model reaches the first convergence condition, thus determining the target visual feature fusion model.
[0233] In another embodiment of the present invention, the device further includes:
[0234] The second model building module (not shown in the figure) is configured to establish the pre-established radar visual feature fusion model before the step of determining the matching current visual perception target and the current first radar perception target, and the corresponding current fused perception feature based on the fused visual features corresponding to each current visual perception target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and the pre-established radar visual feature fusion model. The second model building module is specifically configured to obtain the initial radar visual feature fusion model.
[0235] Obtain the sample first radar perception features corresponding to the sample first radar perception targets of each sample object at each sample time, wherein the sample first radar perception features corresponding to the sample first radar perception targets are: the first radar perception features of the sample first radar of the corresponding sample object at the corresponding sample time.
[0236] Based on the sample visual fusion features corresponding to the visual perception targets of each sample object at each sample time, the sample first radar perception features corresponding to the first radar perception targets of each sample object at each sample time, the label matching information between each sample visual perception target and each sample first radar perception target, and the label perception features corresponding to each sample object at each sample time, the initial radar visual feature fusion model is trained until the initial radar visual feature fusion model reaches the second convergence condition, and the pre-established radar visual feature fusion model is determined. The sample visual fusion features corresponding to the sample visual perception targets are: fusion features obtained by fusing the sample visual perception features corresponding to the sample visual perception targets and the target visual feature fusion model.
[0237] In another embodiment of the present invention, the device further includes:
[0238] The third model building module (not shown in the figure) is configured to build the pre-established radar target recognition model before determining the current first radar real target and its corresponding current first radar perception features from the current first radar perception target based on the current first radar perception features corresponding to the current first radar perception target and the pre-established radar target recognition model. The third model building module is specifically configured to obtain the initial radar target recognition model.
[0239] Obtain the label authenticity information corresponding to the first radar sensing target of each sample object at each sample time, wherein the label authenticity information is: information that identifies whether the corresponding sample first radar sensing target is a real first radar sensing target;
[0240] Based on the first radar sensing features of each sample object corresponding to the first radar sensing target at each sample time, and the corresponding label authenticity information, the initial radar target recognition model is trained until the initial radar target recognition model reaches the third convergence condition, and the pre-established radar target recognition model is obtained.
[0241] In another embodiment of the present invention, the device further includes:
[0242] The fourth model building module (not shown in the figure) is configured to build the pre-established tracking model before determining the matching historical target and the current sensing target based on the historical features of each historical target, the current sensing features of the current sensing target, and the pre-established tracking model. The fourth model building module includes:
[0243] The first acquisition unit (not shown in the figure) is configured to acquire the initial tracking model;
[0244] The second obtaining unit (not shown in the figure) is configured to obtain, for each sample object, the sample historical features corresponding to the sample historical targets at N time points before each sample time point. The sample historical features include: the visual fusion features corresponding to the matching sample visual perception targets at N time points before each sample time point, and the features corresponding to the sample first radar perception targets after fusion and optimization; the optimized features corresponding to the sample first radar perception targets at N time points before each sample time point without matching sample visual perception targets; and / or the predicted features corresponding to the unperceived targets at the sample time point based on the perception features corresponding to the sample historical targets at N time points before each sample time point.
[0245] The third obtaining unit (not shown in the figure) is configured to obtain the label pairing information corresponding to the sample perception target corresponding to the sample object at each sample time for each sample object at each sample time. The label pairing information corresponding to the sample perception target includes: information characterizing whether the sample perception target and the sample historical target corresponding to the sample object at N times before the sample time are the same physical target.
[0246] The training unit (not shown in the figure) is configured to train the initial tracking model based on the sample perception features corresponding to the sample perception target of each sample object at each sample time, the sample history features corresponding to the sample history target of each sample object at N time steps before each sample time, the label pairing information corresponding to each sample perception target of each sample object at each sample time, and the label perception features corresponding to each sample object at each sample time. The initial tracking model is trained until the initial tracking model reaches the fourth convergence condition, and a pre-established tracking model is obtained. The sample perception features include: the sample fusion perception features corresponding to the matching sample visual perception target and the sample first radar perception target of each sample object at each sample time, and / or the first radar perception features corresponding to the sample first radar real target without matching sample visual perception targets determined by the sample radar perception features corresponding to the sample first radar perception target and the pre-established radar target recognition model.
[0247] In another embodiment of the present invention, the training unit is specifically configured to pair the sample perception targets corresponding to each sample object at each sample time with the sample history targets corresponding to the sample object at the sample time in pairs, so as to obtain the paired sample perception targets and sample history targets corresponding to the sample object at the sample time.
[0248] The sample perception features corresponding to the paired sample perception targets and the sample historical features corresponding to the paired sample historical targets at each sample time are randomly selected. The label pairing information and label perception features corresponding to the paired sample perception targets and sample historical targets are also input into the current tracking model to obtain the matching affinity between the paired sample perception targets and sample historical targets and the optimized sample perception features corresponding to the sample perception targets. Among them, the matching affinity between the paired sample perception targets and sample historical targets represents the probability value that the paired sample perception targets and sample historical targets are the same physical target. The current tracking model is the initial tracking model or a tracking model with adjusted model parameter values.
[0249] When the number of matching affinity values corresponding to the obtained sample perception target reaches the preset batch size, the current loss value corresponding to the current tracking model is determined based on the matching affinity between the sample perception target and the sample historical target in the preset batch size, the corresponding label pairing information of the sample perception target in the sample perception target and the sample historical target in the pair, the optimized sample perception feature and label perception feature corresponding to the sample perception target, the preset regression loss function and the preset classification loss function.
[0250] If the current loss value is less than the preset loss threshold, then the current tracking model is determined to have reached the fourth convergence condition, and an intermediate tracking model is obtained.
[0251] Using a preset smoothness loss function, a preset regression loss function, a preset classification loss function, the sample perception features corresponding to the paired sample perception targets and the sample history features corresponding to the paired sample historical targets at each sample time, the label pairing information and label perception features corresponding to the paired sample perception targets and sample history targets, and the specified sample history features corresponding to the paired sample historical targets and sample history targets, the model parameters of the intermediate tracking model are adjusted until the intermediate tracking model reaches the fifth convergence condition to obtain the pre-established tracking model. Among them, the specified sample history features corresponding to the paired sample historical targets and sample history targets are: the sample history features of the paired sample historical targets and sample history targets at the time before the sample perception target.
[0252] If the current loss value is not less than the preset loss threshold, adjust the values of the model parameters of the current tracking model, and return the sample perception features corresponding to the paired sample perception targets and the sample historical features corresponding to the paired sample perception targets and the sample historical targets at each sample time. Input the label pairing information and label perception features corresponding to the paired sample perception targets and sample historical targets into the current tracking model to obtain the matching affinity corresponding to the sample perception target. Continue until the current tracking model reaches the fourth convergence condition to obtain the intermediate tracking model.
[0253] In another embodiment of the present invention, the device further includes:
[0254] The fifth model building module (not shown in the figure) is configured to build the pre-established prediction model before the step of determining the prediction features of each historical target at the current time based on the historical features of each historical target and the pre-established prediction model. The fifth model building module is specifically configured to obtain the initial prediction model.
[0255] Based on the sample history features of each sample object at N time points before each sample time and the label-aware features, the initial prediction model is trained until the initial prediction model reaches the sixth convergence condition, thus obtaining the pre-established prediction model.
[0256] The above system and device embodiments correspond to the method embodiments and have the same technical effects. For detailed descriptions, please refer to the method embodiments. The device embodiments are derived based on the method embodiments; detailed descriptions can be found in the method embodiments section, and will not be repeated here. Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.
[0257] Those skilled in the art will understand that the modules in the apparatus of the embodiments can be distributed in the apparatus of the embodiments as described in the embodiments, or they can be located in one or more devices different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.
[0258] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fusing multi-sensor data, characterized in that, The method includes: Obtain the current visual perception result and the current first radar perception result of the target object at the current moment; Based on the current visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model, the fused visual features corresponding to each current visual perception target are determined. The target visual feature fusion model is a model trained based on the label perception features corresponding to each sample object at each sample time and the sample visual perception features corresponding to each sample visual perception target. Based on the fused visual features corresponding to each current visual perception target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and a pre-established radar visual feature fusion model, the matching current visual perception targets and current first radar perception targets, as well as their corresponding current fused perception features, are determined. This includes: pairing each current visual perception target and each current first radar perception target in the current first radar perception result into pairs; inputting the fused visual features corresponding to each pair of current visual perception targets and the current first radar perception features corresponding to each current first radar perception target into the pre-established radar visual feature fusion model; using the radar visual feature fusion model, based on the fused visual features corresponding to each pair of current visual perception targets and the current first radar perception features corresponding to each current first radar perception target, the affinity score between each pair of current visual perception targets and the current first radar perception targets, as well as the fused perception features corresponding to each pair of current visual perception targets and the current first radar perception targets, are determined; and based on the affinity score, the matching current visual perception targets and current first radar perception targets and their corresponding current fused perception features are determined. The pre-established radar visual feature fusion model is a model trained based on the sample visual perception features, the label perception features, and the sample first radar perception results corresponding to the sample first radar perception target at each sample time.
2. The method as described in claim 1, characterized in that, The method further includes: Based on the current first radar sensing features corresponding to the current first radar sensing target and the pre-established radar target recognition model, the current first radar real target and its corresponding current first radar sensing features are determined from the current first radar sensing targets. The pre-established radar target recognition model is a model trained based on the sample first radar sensing features corresponding to each sample first radar sensing target and the label sensing features.
3. The method as described in claim 2, characterized in that, The method further includes: Obtain the historical features of the target object corresponding to N historical moments before the current moment, where N is a positive integer. The historical features include: the fused visual features corresponding to the matching historical visual perception targets at each historical moment, and the features fused and optimized with the features corresponding to the historical first radar perception targets, the optimized features corresponding to the first radar perception targets without matching visual perception targets at each historical moment, and / or the predicted features corresponding to the unperceived targets at the current historical moment, which are predicted based on the perception features corresponding to the targets perceived at moments before each historical moment. Based on the historical characteristics of each historical target, the current perception characteristics corresponding to the current perceived target, and the pre-established tracking model, the matching historical targets and the current perceived targets, as well as the optimized current perception characteristics corresponding to the matching historical targets and the current perceived targets, are determined. The current perception characteristics corresponding to the current perceived targets include: the current fused perception characteristics corresponding to the matching current visual perception targets and the current first radar perception targets, and / or the radar perception characteristics corresponding to the current first radar real targets without matching current visual perception targets. The pre-established tracking model is a model trained based on the sample perception characteristics corresponding to the sample perception targets corresponding to each sample object at each sample time, the sample historical characteristics corresponding to the sample historical targets corresponding to each sample object at N time points prior to each sample time, and the label matching information corresponding to the sample perception targets corresponding to each sample object at each sample time.
4. The method as described in claim 3, characterized in that, The method further includes: Based on the historical characteristics of each historical target and the pre-established prediction model, the prediction characteristics of each historical target at the current time are determined. The pre-established prediction model is a model trained based on the historical perception results of each sample object corresponding to the historical target of each sample at N times before each sample time, and the label matching information of each sample object corresponding to each sample perception target at each sample time. Based on the current perception features of the current perception target and the predicted features of each historical target at the current time, the current target and its corresponding current features at the current time are determined.
5. The method according to any one of claims 1-4, characterized in that, Before the step of determining the fused visual features corresponding to each current visual perception target based on the visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model, the method further includes: The process of establishing the target visual feature fusion model includes: Obtain the initial visual feature fusion model; Obtain the label perception features of each sample object at each sample time and the sample visual perception features of each sample visual perception target. The label perception features are: the label perception features of the sample label radar of the corresponding sample object at the corresponding sample time. The sample visual perception features of each sample visual perception target are: the visual perception features of the sample image acquisition device group of the corresponding sample object at the corresponding sample time. Using the visual perception features of each sample corresponding to the visual perception target and the label perception features, the initial visual feature fusion model is trained until the initial visual feature fusion model reaches the first convergence condition, thus determining the target visual feature fusion model.
6. The method as described in claim 5, characterized in that, Before the step of determining the matching current visual perception target and the current first radar perception target, and the corresponding current fused perception feature, based on the fused visual features corresponding to each current visual perception target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and the pre-established radar visual feature fusion model, the method further includes: The process of establishing the pre-established radar visual feature fusion model includes: Obtain an initial radar visual feature fusion model; Obtain the sample first radar perception features corresponding to the sample first radar perception targets of each sample object at each sample time, wherein the sample first radar perception features corresponding to the sample first radar perception targets are: the first radar perception features of the sample first radar of the corresponding sample object at the corresponding sample time. Based on the sample visual fusion features corresponding to the visual perception targets of each sample object at each sample time, the sample first radar perception features corresponding to the first radar perception targets of each sample object at each sample time, the label matching information between each sample visual perception target and each sample first radar perception target, and the label perception features corresponding to each sample object at each sample time, the initial radar visual feature fusion model is trained until the initial radar visual feature fusion model reaches the second convergence condition, and the pre-established radar visual feature fusion model is determined. The sample visual fusion features corresponding to the sample visual perception targets are: fusion features obtained by fusing the sample visual perception features corresponding to the sample visual perception targets and the target visual feature fusion model.
7. The method as described in claim 6, characterized in that, Before the step of determining the current first radar real target and its corresponding current first radar sensing features from the current first radar sensing targets based on the current first radar sensing features corresponding to the current first radar sensing target and the pre-established radar target recognition model, the method further includes: The process of establishing the pre-established radar target recognition model includes: Obtain the initial radar target recognition model; Obtain the label authenticity information corresponding to the first radar sensing target of each sample object at each sample time, wherein the label authenticity information is: information that identifies whether the corresponding sample first radar sensing target is a real first radar sensing target; Based on the first radar sensing features of each sample object corresponding to the first radar sensing target at each sample time, and the corresponding label authenticity information, the initial radar target recognition model is trained until the initial radar target recognition model reaches the third convergence condition, and the pre-established radar target recognition model is obtained.
8. The method as described in claim 7, characterized in that, Before the step of determining the matching historical target and the current perceived target based on the historical characteristics of each historical target, the current perceived characteristics of the current perceived target, and the pre-established tracking model, the method further includes: The process of establishing the pre-established tracking model includes: Obtain the initial tracking model; For each sample object, obtain the sample history features corresponding to the sample history targets at N time points before each sample time. The sample history features include: the visual fusion features corresponding to the matching sample visual perception targets at N time points before each sample time, and the features corresponding to the sample first radar perception targets after fusion and optimization; the optimized features corresponding to the sample first radar perception targets at N time points before each sample time without matching sample visual perception targets; and / or the predicted features corresponding to the unperceived targets at the sample time, predicted based on the perception features corresponding to the sample history targets at N time points before each sample time. For each sample object at each sample time, obtain the label pairing information corresponding to the sample perception target corresponding to the sample object at each sample time. The label pairing information corresponding to the sample perception target includes: information characterizing whether the sample perception target and the sample historical target corresponding to the sample object at N times before the sample time are the same physical target. Based on the sample perception features corresponding to the sample perception targets of each sample object at each sample time, the sample history features corresponding to the sample history targets of each sample object at N time points before each sample time, the label pairing information corresponding to the sample perception targets of each sample object at each sample time, and the label perception features corresponding to each sample object at each sample time, the initial tracking model is trained until the initial tracking model reaches the fourth convergence condition, and a pre-established tracking model is obtained. The sample perception features include: the sample fusion perception features corresponding to the matching sample visual perception targets and the sample first radar perception targets of each sample object at each sample time, and / or the first radar perception features corresponding to the sample first radar real targets without matching sample visual perception targets determined by the sample radar perception features corresponding to the sample first radar perception targets and the pre-established radar target recognition model.
9. The method as described in claim 8, characterized in that, The step of training the initial tracking model based on the sample perception features corresponding to the sample perception target of each sample object at each sample time, the sample history features corresponding to the sample history targets of each sample object at N time points before each sample time, the label pairing information corresponding to each sample perception target of each sample object at each sample time, and the label perception features corresponding to each sample object at each sample time, until the initial tracking model reaches the fourth convergence condition, and determining the pre-established tracking model, includes: For each sample object at each sample time, the sample perception targets corresponding to the sample object at each sample time are paired with the sample history targets corresponding to the sample object at N times before the sample time to obtain the paired sample perception targets and sample history targets corresponding to the sample object at the sample time. The sample perception features corresponding to the paired sample perception targets and the sample historical features corresponding to the paired sample historical targets at each sample time are randomly selected. The label pairing information and label perception features corresponding to the paired sample perception targets and sample historical targets are also input into the current tracking model to obtain the matching affinity between the paired sample perception targets and sample historical targets and the optimized sample perception features corresponding to the sample perception targets. The matching affinity between the paired sample perception targets and sample historical targets represents the probability value that the paired sample perception targets and sample historical targets are the same physical target. The current tracking model is the initial tracking model or a tracking model with adjusted model parameter values. When the number of matching affinity values corresponding to the obtained sample perception target reaches the preset batch size, the current loss value corresponding to the current tracking model is determined based on the matching affinity between the sample perception target and the sample historical target in the preset batch size, the corresponding label pairing information of the sample perception target in the sample perception target and the sample historical target in the pair, the optimized sample perception feature and label perception feature corresponding to the sample perception target, the preset regression loss function and the preset classification loss function. If the current loss value is less than the preset loss threshold, then the current tracking model is determined to have reached the fourth convergence condition, and an intermediate tracking model is obtained. Using a preset smoothness loss function, a preset regression loss function, a preset classification loss function, the sample perception features corresponding to the paired sample perception targets and the sample history features corresponding to the paired sample historical targets at each sample time, the label pairing information and label perception features corresponding to the paired sample perception targets and sample history targets, and the specified sample history features corresponding to the paired sample historical targets and sample history targets, the model parameters of the intermediate tracking model are adjusted until the intermediate tracking model reaches the fifth convergence condition to obtain the pre-established tracking model. Among them, the specified sample history features corresponding to the paired sample historical targets and sample history targets are: the sample history features of the paired sample historical targets and sample history targets at the time before the sample perception target. If the current loss value is not less than the preset loss threshold, adjust the values of the model parameters of the current tracking model, and return to the step of randomly pairing the sample perception features corresponding to the sample perception targets and the sample historical features corresponding to the sample historical targets at each sample time, as well as the label pairing information and label perception features corresponding to the sample perception targets in the pairwise paired sample perception targets and sample historical targets, and inputting them into the current tracking model to obtain the matching affinity corresponding to the sample perception target, until the current tracking model reaches the fourth convergence condition to obtain the intermediate tracking model.
10. The method as described in claim 8, characterized in that, Before the step of determining the predicted features of each historical target at the current moment based on the historical features of each historical target and a pre-established prediction model, the method further includes: The process of establishing the pre-established prediction model, the process including: Obtain the initial prediction model; Based on the sample history features of each sample object at N time points before each sample time and the label-aware features, the initial prediction model is trained until the initial prediction model reaches the sixth convergence condition, thus obtaining the pre-established prediction model.
11. A multi-sensor data fusion device, characterized in that, The device includes: The first acquisition module is configured to acquire the current visual perception result and the current first radar perception result of the target object at the current moment. The first determining module is configured to determine the fused visual features corresponding to each current visual perception target based on the current visual perception features corresponding to each current visual perception target in each current visual perception result and the target visual feature fusion model. The target visual feature fusion model is a model trained based on the label perception features corresponding to each sample object at each sample time and the sample visual perception features corresponding to each sample visual perception target. The second determining module is configured to determine matching current visual perception targets and current first radar perception targets, and corresponding current fused perception features, based on the fused visual features corresponding to each current visual perception target, the current first radar perception features corresponding to each current first radar perception target in the current first radar perception result, and a pre-established radar visual feature fusion model. This includes: pairing each current visual perception target and each current first radar perception target in the current first radar perception result; inputting the fused visual features corresponding to each pair of current visual perception targets and the current first radar perception features corresponding to each current first radar perception target into the pre-established radar visual feature fusion model; determining the affinity score between each pair of current visual perception targets and the current first radar perception targets, and the fused perception features corresponding to each pair of current visual perception targets and the current first radar perception targets, based on the fused visual features corresponding to each pair of current visual perception targets and the current first radar perception features corresponding to each current first radar perception target; and determining matching current visual perception targets and current first radar perception targets and their corresponding current fused perception features based on the affinity score. The pre-established radar visual feature fusion model is a model trained based on the sample visual perception features, the label perception features, and the sample first radar perception results corresponding to the sample first radar perception target at each sample time.
Citation Information
Patent Citations
Vehicle detection method based on monocular vision and laser radar fusion
CN111291714A
Environmental perception method based on machine vision and millimeter wave radar data fusion
CN111505624A