Reighting data association precision improving method and device and computing equipment cluster

By using radar depth values ​​to update the homography matrix for adaptive mapping of vision and radar, the problems of complex deployment and high maintenance cost in radar and camera data association are solved, and high-precision and generalized data association is achieved.

CN120703754APending Publication Date: 2025-09-26HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410288103.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In existing technologies, data association methods for radar and cameras require the use of dedicated equipment to construct high-definition map mapping tables, resulting in complex deployment and high maintenance costs, and insufficient association accuracy in complex multi-target scenarios.

Method used

By obtaining the visual and radar target-level result point sets and using the radar depth value to update the homography matrix, adaptive mapping between vision and radar is achieved for data association.

Benefits of technology

It achieves high-precision data association in long-distance and complex multi-target scenarios, reduces deployment complexity and maintenance costs, and improves the generalization of association.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120703754A_ABST
    Figure CN120703754A_ABST
Patent Text Reader

Abstract

A thunder view data association precision improving method comprises the steps that a visual target level result point set and a radar target level result point set are obtained, the visual target level result point set is obtained based on data processing of a target object collected by a visual sensor, and the radar target level result point set is obtained based on data processing of the target object collected by a radar; target errors between the visual target level result point set and the radar target level result point set are determined based on radar depth values contained in the radar target level result point set, and the target errors comprise a first error related to the radar plane view and / or a second error related to the pixel plane view; a homography matrix between the vision sensor and the radar is updated based on the target error. According to the method, the homography matrix is updated based on the radar depth value, so that the accuracy of thunder view data association is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a method, device, and computing device cluster for improving the accuracy of radar data association. Background Art

[0002] When detecting and warning of traffic incidents, sampling cameras are often used to obtain traffic status information of the road, and AI algorithms are used to process the data obtained by the camera to achieve vehicle detection, vehicle speed prediction, and traffic status analysis. However, cameras are not sensitive to speed estimation and have poor longitudinal ranging capabilities. In recent years, millimeter-wave radars have gradually been used in traffic scenarios to accurately measure the speed and distance of vehicles. However, due to the limitations of antenna size and array layout, their lateral positioning accuracy is insufficient and they are easily interfered with by environmental signals. Therefore, accurate perception of traffic information based on radar-visual fusion (i.e., the combination of cameras and radars) has gradually become the mainstream trend. Among them, the premise of radar-visual fusion is the need to perform data association on the target-level results of the two sensors, so as to perform effective back-end information fusion.

[0003] In related technologies, the association of multi-target data using radar often uses a high-definition map mapping table to associate the target-level results of two heterogeneous signals. However, this method requires the use of dedicated equipment (such as lidar / real-time kinematic (RTK) equipment, etc.) to perform refined data collection of the target scene area and perform specific post-processing to obtain a high-definition map mapping table, that is, the map construction process is complex and costly. At the same time, since roads often undergo irregular maintenance or renovation, the high-definition map mapping table needs to be upgraded and maintained irregularly, resulting in high maintenance costs. Therefore, how to provide a precise association method for multi-target data using radar that is simple to deploy, has high association accuracy, and is highly generalizable is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] The present application provides a method, device, computing device cluster, computer storage medium and computer product for improving the accuracy of radar data association, which can provide a method for accurately associating radar multi-target data with simple deployment, high association accuracy and strong generalization.

[0005] In a first aspect, the present application provides a method for improving the accuracy of radar data association, including: obtaining a visual target level result point set and a radar target level result point set, the visual target level result point set is obtained based on data processing of the target object collected by the visual sensor, and the radar target level result point set is obtained based on data processing of the target object collected by the radar; based on the radar depth value contained in the radar target level result point set, determining the target error between the visual target level result point set and the radar target level result point set, wherein the target error includes: a first error related to the radar plane view and / or a second error related to the pixel plane view; based on the target error, updating the homography matrix between the visual sensor and the radar.

[0006] In this way, the radar depth value can be used to guide the adaptive update of the homography matrix between the visual sensor and the radar, so that the updated homography matrix can meet the requirements of radar data association in long-range and complex multi-target scenarios, thus providing a basis for the subsequent accurate association of radar data.

[0007] In one possible implementation, based on the radar depth values ​​contained in the radar target level result point set, the target error between the visual target level result point set and the radar target level result point set is determined. Specifically, the method includes: performing an inverse perspective transformation on the pixels in the visual target level result point set that match the radar points based on the radar depth values ​​expressed by the radar points in the radar target level result point set to obtain a first radar point set; and performing a calculation based on the first radar point set and the visual target level result point set to obtain the target error. In this way, the pseudo-true values ​​of the radar plane view (i.e., the points in the first radar point set) can be obtained through the inverse perspective transformation guided by the radar depth. Then, the error is solved in the radar plane view and / or pixel plane view based on the pseudo-true values, thereby obtaining the optimal radar-visual correlation mapping relationship matrix (i.e., homography matrix) that can meet the requirements of long-range, complex multi-target scenarios.

[0008] In one possible implementation, based on the radar depth values ​​expressed by the radar points in the radar target-level result point set, an inverse perspective transformation is performed on the pixels in the visual target-level result point set that match the radar points to obtain a first radar point set. Specifically, the method includes: using a homography matrix to project the radar points in the radar target-level result point set onto a pixel plane view to obtain a second radar point set; matching the second radar point set with the visual target-level result point set to obtain at least one first target pair, each of which includes a radar point in the second radar point set and a pixel in the visual target result point set; and based on the radar depth values ​​expressed by the radar points in each first target pair and the homography matrix, performing an inverse perspective transformation on the pixels in each first target pair to obtain a first radar point set, wherein each radar point in the first radar point set is associated with a pixel in each first target pair. In this way, an inverse perspective transformation guided by radar depth is achieved, laying a solid foundation for subsequent error calculation.

[0009] In one possible implementation, based on the radar depth values ​​expressed by the radar points in each first target pair, an inverse perspective transformation is performed on the pixels in each first target pair to obtain a first radar point set. Specifically, for any first target pair, the radar depth values ​​expressed by the radar points in any first target pair are processed based on the homogeneous constraint expressed by the homography matrix to obtain a depth regularization term; and based on the homography matrix and the depth regularization term, the pixels in any first target pair are processed to obtain a target radar point in the first radar point set. The target radar point is a point in the first radar point set that is associated with a pixel in any first target pair. Because the radar depth values ​​of different radar points vary, this method allows for a dynamically updated depth regularization term to be obtained. This allows for the calculation of precise inverse perspective positions using different depth regularization terms, thereby obtaining a pseudo-truth radar point, i.e., a point in the first radar point set.

[0010] In one possible implementation, calculating the target error based on the first radar point set and the visual target-level result point set specifically includes: using a homography matrix to project pixel points in the visual target-level result point set onto a radar plane view to obtain a first visual point set, and calculating based on the first visual point set and the first radar point set to obtain a first error; and / or using the homography matrix to project radar points in the first radar point set onto a pixel plane view to obtain a third radar point set, and calculating based on the third radar point set and the visual target-level result point set to obtain a second error. In this way, the error associated with the radar plane view and the error associated with the pixel plane view can be calculated, thereby facilitating subsequent updating of the homography matrix.

[0011] In one possible implementation, after updating the homography matrix between the visual sensor and the radar, the method further includes: correlating pixels in the visual target-level result point set with radar points in the radar target-level result point set using the updated homography matrix; and fusing the pixels in the visual target-level result point set with the radar points in the radar target-level result point set based on the correlation result. Because the updated homography matrix matches the current data acquisition environment, correlating the radar data using the updated homography matrix can produce accurate correlation data, thereby making subsequent fusion results more accurate.

[0012] In one possible implementation, the updated homography matrix is ​​used to associate pixels in the visual target-level result point set with radar points in the radar target-level result point set. Specifically, the following steps are performed: projecting the radar points in the radar target-level result point set onto a pixel plane view based on the updated homography matrix to obtain a fourth radar point set; matching the fourth radar point set with the visual target-level result point set to obtain at least one second target pair, each of which includes a radar point in the fourth radar point set and a pixel in the visual target result point set; and projecting the pixel points in each second target pair onto the radar plane view based on the radar depth values ​​expressed by the radar points in each second target pair and the updated homography matrix. In this way, association of radar visual data can be achieved.

[0013] In one possible implementation, the method further includes: projecting target pixels in the visual target-level result point set that do not match radar points in the fourth radar point set onto a radar plane view based on the updated homography matrix; and associating the target pixels with radar points in the radar target-level result point set in the radar plane view. In this way, association of radar visual data can be achieved.

[0014] In a second aspect, the present application provides a device for improving the accuracy of radar data association, including: an acquisition module and a processing module. The acquisition module is used to obtain a visual target level result point set and a radar target level result point set, the visual target level result point set is obtained based on the data processing of the target object collected by the visual sensor, and the radar target level result point set is obtained based on the data processing of the target object collected by the radar. The processing module is used to determine the target error between the visual target level result point set and the radar target level result point set based on the radar depth value contained in the radar target level result point set, wherein the target error includes: a first error related to the radar plane view and / or a second error related to the pixel plane view. The processing module is also used to update the homography matrix between the visual sensor and the radar based on the target error.

[0015] In one possible implementation, when the processing module determines the target error between the visual target level result point set and the radar target level result point set based on the radar depth value contained in the radar target level result point set, the processing module is specifically used to: perform an inverse perspective transformation on the pixel points in the visual target level result point set that match the radar points based on the radar depth value expressed by the radar points in the radar target level result point set to obtain a first radar point set; and perform calculations based on the first radar point set and the visual target level result point set to obtain the target error.

[0016] In one possible implementation, when the processing module performs an inverse perspective transformation on pixel points in the visual target level result point set that match the radar point based on the radar depth value expressed by the radar point in the radar target level result point set to obtain the first radar point set, the processing module is specifically configured to: project the radar points in the radar target level result point set onto a pixel plane view using a homography matrix to obtain a second radar point set; match the second radar point set with the visual target level result point set to obtain at least one first target pair, each first target pair including a radar point in the second radar point set and a pixel in the visual target result point set; and perform an inverse perspective transformation on the pixel points in each first target pair based on the radar depth value expressed by the radar point in each first target pair and the homography matrix to obtain the first radar point set, wherein a radar point in the first radar point set is associated with a pixel point in one first target pair.

[0017] In one possible implementation, when the processing module performs calculations based on the first radar point set and the visual target level result point set to obtain the target error, the processing module is specifically configured to: project pixel points in the visual target level result point set onto a radar plane view using a homography matrix to obtain a first visual point set, and perform calculations based on the first visual point set and the first radar point set to obtain a first error; and / or project radar points in the first radar point set onto a pixel plane view using a homography matrix to obtain a third radar point set, and perform calculations based on the third radar point set and the visual target level result point set to obtain a second error.

[0018] In one possible implementation, after updating the homography matrix between the visual sensor and the radar, the processing module is further used to: use the updated homography matrix to associate the pixel points in the visual target level result point set and the radar points in the radar target level result point set; and, based on the association result, fuse the pixel points in the visual target level result point set and the radar points in the radar target level result point set.

[0019] In one possible implementation, when the processing module uses the updated homography matrix to associate the pixel points in the visual target level result point set with the radar points in the radar target level result point set, the processing module is specifically configured to: project the radar points in the radar target level result point set to the pixel plane view based on the updated homography matrix to obtain a fourth radar point set; match the fourth radar point set with the visual target level result point set to obtain at least one second target pair, each second target pair including a radar point in the fourth radar point set and a pixel point in the visual target result point set; and project the pixel points in each second target pair to the radar plane view based on the radar depth value expressed by the radar point in each second target pair and the updated homography matrix.

[0020] In one possible implementation, the processing module is further configured to: project, based on the updated homography matrix, target pixel points in the visual target-level result point set that do not match radar points in the fourth radar point set onto the radar plane view; and associate, in the radar plane view, the target pixel points with the radar points in the radar target-level result point set.

[0021] In a third aspect, the present application provides a computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect.

[0022] In a fourth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect. For example, the computing device cluster may include one or more computing devices.

[0023] In a fifth aspect, the present application provides a computer program product comprising instructions that, when executed by a computing device cluster, cause the computing device cluster to perform the method described in the first aspect or any possible implementation of the first aspect. For example, the computing device cluster may include one or more computing devices.

[0024] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a schematic diagram of the architecture of a radar and vision fusion system provided in an embodiment of the present application;

[0026] Figure 2This is a flow chart of a method for improving the accuracy of radar data association provided by an embodiment of the present application;

[0027] Figure 3 This is a schematic diagram of the steps of performing an inverse perspective transformation on pixel points in a visual target-level result point set that match the radar point using a radar depth value expressed by the radar point in the radar target-level result point set, provided by an embodiment of the present application;

[0028] Figure 4 This is a schematic diagram of the steps of associating pixel points in a visual target-level result point set and radar points in a radar target-level result point set using an updated homography matrix provided by an embodiment of the present application;

[0029] Figure 5 This is a schematic structural diagram of a device for improving the accuracy of radar data association provided by an embodiment of the present application;

[0030] Figure 6 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0031] Figure 7 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0032] Figure 8 This is a structural diagram of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.

[0034] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.

[0035] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0036] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0037] Generally, when performing radar-based multi-target data association, a fixed homography matrix can be used to correlate the target-level results of two heterogeneous signals. This approach often uses sensor calibration to obtain the radar and camera's extrinsic mapping matrix. This static mapping matrix is ​​then used to transform the position of the "pixel view" to the "radar view" (i.e., the bird's-eye view (BEV) plane of the road). This process is often referred to as inverse perspective mapping (IPM). The transformation from the "pixel view" to the "radar view" can be achieved using the following "Formula 1."

[0038]

[0039] Among them, (x radar ,y radar ) is the coordinate of the radar point under the radar view, (x pixel ,y pixel ) are the coordinates of the pixel point in the pixel view, and H is the homography matrix. H can be, but is not limited to, derived from calibrated radar and camera extrinsic parameters, as well as camera intrinsic parameters. For example, H is primarily used to implement the conversion between the pixel view and the radar view, and can also be called a transformation matrix.

[0040] However, the homography matrix approach causes nonlinear increases in mapping error as the perception distance increases. This is especially true at distances above 300 meters. Even slight camera movement or displacement (e.g., a few pixels) can lead to mapping errors as high as ten meters along the road's depth, severely impacting association accuracy. Furthermore, a fixed homography matrix imposes a fixed mapping relationship on objects at varying distances, making it impossible to mitigate statistical mapping errors.

[0041] In light of this, embodiments of the present application provide a method for improving the accuracy of radar-visual data association. This method utilizes radar depth values ​​to guide the adaptive update of the homography matrix between the visual sensor and the radar, thereby obtaining the optimal radar-visual association mapping relationship matrix (i.e., the homography matrix) that satisfies long-range, complex multi-target scenarios, achieving accurate association of radar-visual multi-target data. Furthermore, this method eliminates the need for specialized equipment to collect high-definition maps, making it simple to deploy and highly generalizable.

[0042] For example, Figure 1 The schematic diagram of the architecture of a radar and vision fusion system provided by the embodiment of the present application is shown. Figure 1 As shown, the radar-visual fusion system may include: a visual sensor 110, a radar 120, a computing device 130, and a display device 140. The visual sensor 110 and the radar 120 may be mounted on a roadside pole (such as a traffic sign pole). Each pole may be mounted with at least one visual sensor 110 and at least one radar 120. The visual sensor 110 and the radar 120 may establish a communication connection with the computing device 130 via a wired or wireless network. The computing device 130 may also establish a communication connection with the display device 140 via a wired or wireless network.

[0043] Both the visual sensor 110 and the radar 120 may be responsible for collecting data on perceived objects such as vehicles on the road, and transmitting the collected data to the computing device 130. The visual sensor 110 may be, but is not limited to, a camera. The radar 120 may be, but is not limited to, a millimeter-wave radar.

[0044] The computing device 130 is a device that at least has computing processing capabilities, such as a server. When the computing device 130 is a server, the computing device 130 can be, but is not limited to, a cloud server. In this embodiment, the computing device 130 is primarily responsible for adaptively updating the homography matrix between the visual sensor 110 and the radar 120 based on the data collected by the visual sensor 110 and the radar 120, as well as performing radar-visual fusion based on the updated homography matrix, and displaying the fusion result through the display device 140. In some embodiments, the computing device 130 can be arranged independently from the display device 140, or can be integrated with the display device 140. The specific method can be determined according to actual conditions and is not limited here.

[0045] The display device 140 is a device that at least has display capabilities, such as Huawei Smart Screen. The display device 140 is mainly responsible for displaying the radar-visual fusion results of the computing device 130, such as the image shown in the display area 141, so that the user can understand the traffic status on the road, etc. In addition, the display device 140 can also be used to display the homography matrix between the visual sensor 110 and the radar 120. In addition, the display device 140 can also display the configuration interface of the radar-visual fusion system. On the configuration interface, the user can configure the parameters of the visual sensor 110 and / or the radar 120 (such as coordinates, model, viewing angle, installation height or installation angle, etc.), and can also initialize the homography matrix between the visual sensor 110 and the radar 120, etc. Of course, the configuration interface of the radar-visual fusion system can also be displayed through other terminal devices, which can be determined according to actual conditions and is not limited here.

[0046] The above is an introduction to a radar-visual data fusion system provided in an embodiment of the present application. Based on the above, a radar-visual data association accuracy improvement method provided in an embodiment of the present application is introduced below.

[0047] For example, Figure 2 The flowchart of a method for improving the accuracy of radar data association provided by an embodiment of the present application is shown. It is understood that the method can be executed by any device, equipment, platform, or equipment cluster with computing and processing capabilities. For example, the above Figure 1 The computing device 130 described in the above is executed. For the convenience of description, the following description is based on the computing device 130 as the execution subject. Of course, the computing device 130 can also be replaced by other execution subjects, and the replaced solution is still within the scope of protection of this application. Figure 2 As shown, the method for improving the accuracy of radar data association may include the following steps:

[0048] S201. The computing device 130 obtains a visual target level result point set and a radar target level result point set, wherein the visual target level result point set is obtained based on the data processing of the target object collected by the visual sensor 110, and the radar target level result point set is obtained based on the data processing of the target object collected by the radar 120.

[0049] In this embodiment, the visual sensor 110 transmits its captured image data of the target object to the computing device 130, and the radar 120 transmits its captured radar data of the target object to the computing device 130. The computing device 130 can then process the visual data using a visual target detection algorithm such as a single shot multiBox detector (SSD) or a mask region-based convolutional neural network (mask R-CNN) to obtain a visual target-level result point set. The computing device 130 can also process the radar data using a radar target detection algorithm such as a k-means clustering algorithm to obtain a radar target-level result point set. In some embodiments, the visual target-level result point set and the radar target-level result point set can also be obtained by processing the visual data and radar data, respectively, using other devices or apparatuses. In this case, the computing device 130 can directly obtain the visual target-level result point set and the radar target-level result point set from the other devices or apparatuses. In this case, the visual sensor 110 and the radar 120 do not need to send their captured data to the computing device 130. The specific details can be determined according to the actual situation and are not limited here.

[0050] S202. The computing device 130 determines a target error between the visual target level result point set and the radar target level result point set based on the radar depth value included in the radar target level result point set, wherein the target error includes: a first error associated with the radar plane view and / or a second error associated with the pixel plane view.

[0051] In this embodiment, the computing device 130 can calculate the error between the visual target level result point set and the radar target level result point set using the radar depth value included in the radar target level result point set to obtain a target error. The target error can include: a first error associated with the radar plane view (i.e., BEV) and / or a second error associated with the pixel plane view (i.e., perspective view (PV)).

[0052] As a possible implementation method, the radar depth value expressed by the radar point in the radar target level result point set can be used to perform an inverse perspective transformation on the pixel points that match the radar point in the visual target level result point set to obtain the first radar point set. Specifically, Figure 3 As shown, the following steps may be included: S301, using the homography matrix between the visual sensor 110 and the radar 120, projecting the radar points in the radar target level result point set onto the pixel plane view to obtain a second radar point set. The formula for projecting the radar points in the radar target level result point set may be:

[0053]

[0054] Among them, (U rad ,V rad ) is the radar PV point, that is, the point in the second radar point set, M is the homography matrix, (X rad ,Y rad ) is the radar BEV point, that is, the radar point in the radar target level result point set.

[0055] S302: Match the second radar point set and the visual target level result point set using an algorithm such as the Hungarian matching algorithm to obtain at least one first target pair, wherein each first target pair includes a radar point in the second radar point set and a pixel point in the visual target result point set.

[0056] S303: Based on the radar depth values ​​and homography matrices expressed by the radar points in each first target pair, perform an inverse perspective transformation on the pixel points in each first target pair to obtain a first radar point set. A radar point in the first radar point set is associated with a pixel point in the first target pair. Exemplarily, the formula for performing an inverse perspective transformation on the pixel points may be:

[0057]

[0058] Among them, (X r+c ,Y r+c ) is a point in the first radar point set; (U cam ,V cam ) is a pixel point in the first target pair; M -1 is the inverse of the homography matrix M; Z rad The radar point (U rad ,V rad ) can be obtained from the radar depth value expressed by the radar point (U rad ,V rad ) related radar data points (X rad ,Y rad ) is extracted from. Exemplarily, Z rad Can be with Y rad The same. In this way, the first radar point set is obtained. Among them, the points in the first radar point set can be understood as pseudo-true value radar points. In some embodiments, in order to facilitate calculation, the radar points (U rad ,V rad ) is processed to obtain the depth regularization term of the radar depth value, and then Z is used in the calculation rad Replace it with the calculated depth regularization term. Among them, the homogeneous constraint can be

[0059] After obtaining the first radar point set, the error between the first radar point set and the visual target-level result point set can be calculated to obtain the target error. Specifically, when the target error is the first error, the pixels in the visual target-level result point set can be projected onto the radar plane view using the homography matrix as described above in "Formula 2" to obtain the first visual point set. Then, the first visual point set and the first radar point set are calculated using a cross-entropy loss function, a mean square error (MSE) loss function, and the like to obtain the first error. When the target error is the second error, the radar points in the first radar point set can be projected onto the pixel plane view using the homography matrix as described above in "Formula 2" to obtain a third radar point set. Then, the third radar point set and the visual target-level result point set are calculated using a loss function to obtain the second error. When the target error includes both the first and second errors, the first and second errors can be calculated separately. In this way, the target error is calculated. After obtaining the target error, S203 can be executed.

[0060] S203 : The computing device 130 updates the homography matrix between the visual sensor 110 and the radar 120 based on the target error.

[0061] In this embodiment, after obtaining the target error, the computing device 130 can update the homography matrix between the visual sensor 110 and the radar 120 with the goal of minimizing the target error. In some embodiments, the homography matrix can be iteratively updated until a preset number of iterations is reached, or the updated homography matrix reaches the optimum. After each iteration completes the update of the homography matrix, the updated homography matrix can be used as a new round of homography matrix to be updated, and the aforementioned S202 and S203 are re-executed. In addition, when the first error and the second error are used simultaneously, considering that the scales of the two errors may be different, for ease of calculation, the two errors can be normalized first to obtain two normalized errors, and the homography matrix is ​​updated by the normalized errors.

[0062] In this way, the radar depth value is used to adaptively update the homography matrix between the visual sensor and the radar, thereby obtaining the optimal radar-visual correlation mapping relationship matrix (i.e., homography matrix) that meets the requirements of long-range and complex multi-target scenarios, providing a basis for the subsequent accurate association of radar-visual data.

[0063] After completing the update of the homography matrix, the computing device 130 may present the updated homography matrix to the user via the display device 140. In addition, the computing device 130 may also use the updated homography matrix to associate the pixel points in the visual target level result point set with the radar points in the radar target level result point set. Specifically, Figure 4 As shown, the following steps may be included: S401. Through the aforementioned "Formula 2" and based on the updated homography matrix, the radar points in the radar target level result point set are projected to the pixel plane view to obtain a fourth radar point set. S402. Through algorithms such as the Hungarian matching algorithm, the fourth radar point set and the visual target level result point set are matched to obtain at least one second target pair. Each second target pair includes a radar point in the fourth radar point set and a pixel point in the visual target result point set. S403. Through the aforementioned "Formula 3", and based on the radar depth values ​​expressed by the radar points in each second target pair and the updated homography matrix, the pixel points in each second target pair are projected to the radar plane view. In this way, the association of the radar visual data is completed.

[0064] exist Figure 4In S402, there may be pixels in the visual target-level result point set that do not match radar points in the fourth radar point set. Without associating these pixels, it would be difficult to fully associate all radar visual data. Therefore, in this embodiment, the aforementioned "Formula 2" can be used, and based on the updated homography matrix, all target pixels in the visual target-level result point set that do not match radar points in the fourth radar point set are projected onto the radar plane view. Then, in the radar plane view, an algorithm such as the Hungarian matching algorithm can be used to match the target pixels with the radar points in the radar target-level result point set to complete the association between the two.

[0065] After completing the association of the radar visual data in the radar plane view, the data in the visual target level result point set and the radar target level result point set can be fused based on the association result in the radar plane view to obtain the radar visual fusion result. For example, when performing the radar visual fusion, the radar points and pixel points with the association relationship can be weighted in the radar plane view, and the weighted result can be used as the radar visual fusion result. For example, in the radar plane view, if the radar point (X rad ,Y rad ) and pixel BEV point (X cam ,Y cam ) is associated, then ((X rad +X cam ) / 2,(Y rad +Y cam ) / 2) as the fusion result.

[0066] The above is an introduction to the method for improving the accuracy of radar data association provided by the embodiment of the present application. For easier understanding, the following examples are provided for illustration.

[0067] For example, assuming that the initial homography matrix is ​​M, the visual target level result point set includes pixel points (U cam ,X cam ), the radar target level result point set includes radar point (X rad ,Y rad ), and, through the radar point (X rad ,Y rad ) in the radar depth value Y rad , guide the update of the homography matrix M. At this time, first, the homography matrix M and the radar point (X rad ,Y) into the above “Formula 2”, and we can get the radar PV point (U rad ,V rad ). At this time, the radar point (X rad ,Y rad ) is projected onto PV. The radar PV point is a point in the aforementioned second radar point set.

[0068] Then, the Hungarian algorithm is used to collect the radar PV points (U rad ,V rad ) to match. Assume that the matching result is the radar PV point (U rad ,V rad ) and pixel (U cam ,V cam ) matches, which results in a target pair (also called “radar-pixel associated target pair”). The target pair is composed of radar PV points (U rad ,V rad ) and pixel (U cam ,V cam )composition.

[0069] Next, place the radar point (X rad ,Y rad ) rad As the radar depth value Z rad , and the radar depth value Z rad , the inverse of the homography matrix M and the radar PV point (U rad ,V rad ), and substitute it into the aforementioned “Formula 3” to obtain a pseudo-true value radar point (X r+c ,Y r+c ). The pseudo-true value radar point is a point in the aforementioned first radar point set.

[0070] Next, the homography matrix M and the pixel point (U cam ,V cam ) is substituted into the aforementioned “Formula 2” to obtain the pixel BEV point (X cam ,Y cam ). At this time, the pixel point (U cam ,V cam ) is projected to BEV. And, the homography matrix M and the pseudo-true value radar point (X r+c ,Y r+c ) is substituted into the aforementioned “Formula 2” to obtain the corrected radar PV point (U r+c ,V r+c ). At this time, the pseudo-true value radar point (X r+c ,Y r+c ) is projected onto PV. The pixel BEV point is a point in the aforementioned first visual point set, and the corrected radar PV point is a point in the aforementioned third radar point set.

[0071] Then, through the pseudo-true value radar point (X r+c ,Y r+c ) and pixel BEV point (X cam ,Y cam ), calculate the BEV error L related to BEV BEVAnd, by correcting the radar PV point (U r+c ,V r+c ) and pixel (U cam ,V cam ), calculate the PV error L related to PV 2D Among them, L BEV is the aforementioned first error, L 2D is the second error mentioned above.

[0072] Then, according to the BEV error L BEV-Euc and PV error L 2D-Euc Construct the joint loss L = L BEV +L 2D And, the homography matrix M is updated through the joint loss L.

[0073] Finally, the updated homography matrix can be used as the initial homography matrix for a new round, and the above process is repeated until the joint loss L converges or the preset number of iterations is reached.

[0074] It should be understood that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In addition, the various embodiments described above can be combined according to actual circumstances, and the combined solutions are still within the scope of protection of this application.

[0075] Based on the method in the above embodiment, the embodiment of the present application also provides a device for improving the accuracy of radar data association.

[0076] For example, Figure 5 The embodiment of the present application provides a schematic diagram of the structure of a device for improving the accuracy of radar data association. Figure 5 As shown, the radar data association accuracy improvement device 500 includes: an acquisition module 501 and a processing module 502. The acquisition module 501 is used to obtain a visual target level result point set and a radar target level result point set. The visual target level result point set is obtained based on the data processing of the target object collected by the visual sensor, and the radar target level result point set is obtained based on the data processing of the target object collected by the radar. The processing module 502 is used to determine the target error between the visual target level result point set and the radar target level result point set based on the radar depth value contained in the radar target level result point set, wherein the target error includes: a first error related to the radar plane view and / or a second error related to the pixel plane view. The processing module 502 is also used to update the homography matrix between the visual sensor and the radar based on the target error.

[0077] In some embodiments, when the processing module 502 determines the target error between the visual target level result point set and the radar target level result point set based on the radar depth value contained in the radar target level result point set, it is specifically used to: based on the radar depth value expressed by the radar point in the radar target level result point set, perform an inverse perspective transformation on the pixel points in the visual target level result point set that match the radar point to obtain a first radar point set; and perform calculations based on the first radar point set and the visual target level result point set to obtain the target error.

[0078] In some embodiments, when the processing module 502 performs an inverse perspective transformation on the pixel points in the visual target level result point set that match the radar point based on the radar depth value expressed by the radar point in the radar target level result point set to obtain the first radar point set, the processing module 502 is specifically configured to: project the radar points in the radar target level result point set onto a pixel plane view using a homography matrix to obtain a second radar point set; match the second radar point set with the visual target level result point set to obtain at least one first target pair, each first target pair including a radar point in the second radar point set and a pixel in the visual target result point set; and perform an inverse perspective transformation on the pixel points in each first target pair based on the radar depth value expressed by the radar point in each first target pair and the homography matrix to obtain the first radar point set, wherein a radar point in the first radar point set is associated with a pixel point in one first target pair.

[0079] In some embodiments, when the processing module 502 performs calculations based on the first radar point set and the visual target level result point set to obtain the target error, the processing module 502 is specifically configured to: use a homography matrix to project pixel points in the visual target level result point set to a radar plane view to obtain a first visual point set, and perform calculations based on the first visual point set and the first radar point set to obtain a first error; and / or use a homography matrix to project radar points in the first radar point set to a pixel plane view to obtain a third radar point set, and perform calculations based on the third radar point set and the visual target level result point set to obtain a second error.

[0080] In some embodiments, after updating the homography matrix between the visual sensor and the radar, the processing module 502 is further configured to: associate the pixel points in the visual target level result point set with the radar points in the radar target level result point set using the updated homography matrix.

[0081] In some embodiments, when the processing module 502 uses the updated homography matrix to associate the pixel points in the visual target level result point set and the radar points in the radar target level result point set, it is specifically used to: project the radar points in the radar target level result point set to the pixel plane view based on the updated homography matrix to obtain a fourth radar point set; match the fourth radar point set with the visual target level result point set to obtain at least one second target pair, each second target pair including a radar point in the fourth radar point set and a pixel point in the visual target result point set; and project the pixel points in each second target pair to the radar plane view based on the radar depth value expressed by the radar point in each second target pair and the updated homography matrix.

[0082] In some embodiments, the processing module 502 is further configured to: project all target pixel points in the visual target level result point set that do not match the radar points in the fourth radar point set onto the radar plane view based on the updated homography matrix; associate the target pixel points with the radar points in the radar target level result point set in the radar plane view; and, based on the association result, fuse the pixel points in the visual target level result point set with the radar points in the radar target level result point set.

[0083] In some embodiments, Figure 5 The acquisition module 501 and the processing module 502 shown in FIG can be implemented by software or hardware. For example, the implementation of the acquisition module 501 is described below using the acquisition module 501 as an example. Similarly, the implementation of the processing module 502 can refer to the implementation of the acquisition module 501.

[0084] As an example of a software functional unit, the acquisition module 501 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition module 501 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0085] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0086] As an example of a hardware functional unit, the acquisition module 501 may include at least one computing device, such as a server. Alternatively, the acquisition module 501 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0087] The multiple computing devices included in acquisition module 501 can be distributed in the same region or in different regions. The multiple computing devices included in acquisition module 501 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in acquisition module 501 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0088] It should be noted that, in other embodiments, the acquisition module 501 can be used to execute any step of the method for improving the accuracy of radar data association described in the above embodiment, and the processing module 502 can also be used to execute any step of the method for improving the accuracy of radar data association described in the above embodiment. In addition, the acquisition module 501 can also be combined with the processing module 502 to be responsible for executing any step of the method for improving the accuracy of radar data association described in the above embodiment. In addition, the steps that the acquisition module 501 and the processing module 502 are responsible for implementing can also be specified as needed, and the acquisition module 501 and the processing module 502 can respectively implement different steps of the method for improving the accuracy of radar data association described in the above embodiment to achieve the desired effect. Figure 5The entire functions of the radar data association accuracy improvement device 500 are shown.

[0089] The present application also provides a computing device 600. Figure 6 As shown, computing device 600 includes a bus 602, a processor 604, a memory 606, and a communication interface 608. Processor 604, memory 606, and communication interface 608 communicate with each other via bus 602. Computing device 600 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 600.

[0090] The bus 602 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The bus 604 may include a path for transmitting information between various components of the computing device 600 (eg, memory 606, processor 604, communication interface 608).

[0091] The processor 604 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0092] The memory 606 may include volatile memory, such as random access memory (RAM). The processor 604 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0093] The memory 606 stores executable program codes, and the processor 604 executes the executable program codes to respectively implement the aforementioned Figure 5The functions of the acquisition module 501 and the processing module 502 shown in FIG are implemented to realize the method for improving the accuracy of radar data association described in the above embodiment. That is, the memory 606 stores instructions for executing the method for improving the accuracy of radar data association described in the above embodiment.

[0094] Alternatively, the memory 606 stores executable codes, and the processor 604 executes the executable codes to respectively implement the aforementioned Figure 5 The functions of the radar data association accuracy improvement device 500 shown in FIG are implemented to realize the radar data association accuracy improvement method described in the above embodiment. That is, the memory 606 stores instructions for executing the radar data association accuracy improvement method described in the above embodiment.

[0095] The communication interface 603 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 600 and other devices or a communication network.

[0096] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0097] like Figure 7 As shown, the computing device cluster includes at least one computing device 600. The memory 606 in one or more computing devices 600 in the computing device cluster may store the same instructions for executing the radar data association accuracy improvement method described in the above embodiment.

[0098] In some possible implementations, the memory 606 of one or more computing devices 600 in the computing device cluster may also store partial instructions for executing the radar data association accuracy improvement method described in the above embodiment. In other words, the combination of one or more computing devices 600 can jointly execute instructions for executing the radar data association accuracy improvement method described in the above embodiment.

[0099] It should be noted that the memory 606 in different computing devices 600 in the computing device cluster can store different instructions, which are used to execute the above Figure 7 The illustrated embodiment shows partial functions of the apparatus 700 for improving the accuracy of radar data association. That is, the instructions stored in the memory 606 of different computing devices 600 can implement the functions of one or more modules in the acquisition module 501 and the processing module 502.

[0100] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 8 A possible implementation is shown. Figure 8 As shown, two computing devices 600A and 600B are connected via a network. Specifically, each computing device is connected to the network via a communication interface within the computing device. In this possible implementation, the memory 606 within computing device 600A stores instructions for executing the functions of acquisition module 501. Simultaneously, the memory 606 within computing device 600B stores instructions for executing the functions of processing module 502.

[0101] It should be understood that Figure 8 The functionality of computing device 600A shown in FIG. 6 may also be implemented by multiple computing devices 600. Similarly, the functionality of computing device 600B may also be implemented by multiple computing devices 600.

[0102] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similarly referred to as Figure 7 and Figure 8 The connection mode of the computing device cluster is different in that the memory 606 of one or more computing devices 600 in the computing device cluster may store the same instructions for executing the method in the above embodiment.

[0103] In some possible implementations, the memory 606 of one or more computing devices 600 in the computing device cluster may also store partial instructions for executing the aforementioned method for improving the accuracy of radar data association. In other words, the combination of one or more computing devices 600 can collectively execute instructions for executing the aforementioned method for improving the accuracy of radar data association.

[0104] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, including computer program instructions. When the computer program instructions are executed by a computing device cluster comprising at least one computing device, the computing device cluster executes the method in the above embodiment. Exemplarily, the computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center comprising one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc.

[0105] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product containing instructions. When the instructions are executed by a computing device, a computing device cluster including at least one computing device executes the method in the above embodiment.

[0106] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0107] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0108] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0109] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for improving the accuracy of radar data association, characterized in that: include: Obtaining a visual target level result point set and a radar target level result point set, wherein the visual target level result point set is obtained by processing data of the target object collected by a visual sensor, and the radar target level result point set is obtained by processing data of the target object collected by a radar; Determining a target error between the visual target level result point set and the radar target level result point set based on the radar depth value included in the radar target level result point set, wherein the target error includes: a first error associated with a radar plane view and / or a second error associated with a pixel plane view; Based on the target error, a homography matrix between the visual sensor and the radar is updated.

2. The method according to claim 1, characterized in that The determining, based on the radar depth value included in the radar target level result point set, a target error between the visual target level result point set and the radar target level result point set, specifically includes: Based on the radar depth values ​​expressed by the radar points in the radar target level result point set, performing an inverse perspective transformation on the pixel points in the visual target level result point set that match the radar points to obtain a first radar point set; The target error is obtained by performing calculation based on the first radar point set and the visual target level result point set.

3. The method according to claim 2, characterized in that The step of performing an inverse perspective transformation on pixel points in the visual target level result point set that match the radar points based on the radar depth values ​​expressed by the radar points in the radar target level result point set to obtain a first radar point set specifically includes: projecting the radar points in the radar target level result point set onto the pixel plane view using the homography matrix to obtain a second radar point set; Matching the second radar point set with the visual target level result point set to obtain at least one first target pair, each of the first target pairs including a radar point in the second radar point set and a pixel point in the visual target result point set; Based on the radar depth values ​​expressed by the radar points in each of the first target pairs and the homography matrix, an inverse perspective transformation is performed on the pixel points in each of the first target pairs to obtain the first radar point set, wherein one radar point in the first radar point set is associated with one pixel point in the first target pair.

4. The method according to claim 2 or 3, characterized in that The performing calculation based on the first radar point set and the visual target level result point set to obtain the target error specifically includes: Projecting pixel points in the visual target-level result point set onto the radar plane view using the homography matrix to obtain a first visual point set, and performing calculation based on the first visual point set and the first radar point set to obtain the first error; And / or, using the homography matrix, projecting radar points in the first radar point set onto the pixel plane view to obtain a third radar point set, and performing calculation based on the third radar point set and the visual target level result point set to obtain the second error.

5. The method according to any one of claims 1 to 4, characterized in that: After updating the homography matrix between the visual sensor and the radar, the method further includes: The updated homography matrix is ​​used to associate the pixel points in the visual target level result point set with the radar points in the radar target level result point set, and based on the association result, the pixel points in the visual target level result point set and the radar points in the radar target level result point set are fused.

6. The method according to claim 5, characterized in that The using the updated homography matrix to associate the pixel points in the visual target level result point set with the radar points in the radar target level result point set specifically includes: Based on the updated homography matrix, projecting the radar points in the radar target level result point set onto the pixel plane view to obtain a fourth radar point set; Matching the fourth radar point set with the visual target level result point set to obtain at least one second target pair, each of the second target pairs including a radar point in the fourth radar point set and a pixel point in the visual target result point set; Projecting pixel points in each second target pair onto the radar plane view based on the radar depth value expressed by the radar point in each second target pair and the updated homography matrix.

7. The method according to claim 6, characterized in that Also includes: Based on the updated homography matrix, project all target pixel points in the visual target-level result point set that do not match the radar points in the fourth radar point set onto the radar plane view; In the radar plane view, the target pixel point is associated with a radar point in the radar target level result point set.

8. A device for improving the accuracy of radar data association, characterized in that: include: an acquisition module, configured to acquire a visual target level result point set and a radar target level result point set, wherein the visual target level result point set is obtained by processing data of the target object collected by a visual sensor, and the radar target level result point set is obtained by processing data of the target object collected by a radar; a processing module, configured to determine a target error between the visual target level result point set and the radar target level result point set based on the radar depth value included in the radar target level result point set, wherein the target error includes: a first error associated with the radar plane view and / or a second error associated with the pixel plane view; The processing module is further configured to update a homography matrix between the visual sensor and the radar based on the target error.

9. The device according to claim 8, characterized in that When determining the target error between the visual target level result point set and the radar target level result point set based on the radar depth value included in the radar target level result point set, the processing module is specifically configured to: Based on the radar depth values ​​expressed by the radar points in the radar target level result point set, performing an inverse perspective transformation on the pixel points in the visual target level result point set that match the radar points to obtain a first radar point set; The target error is obtained by performing calculation based on the first radar point set and the visual target level result point set.

10. The device according to claim 9, characterized in that When the processing module performs an inverse perspective transformation on pixel points in the visual target-level result point set that match the radar points based on the radar depth values ​​expressed by the radar points in the radar target-level result point set to obtain the first radar point set, the processing module is specifically configured to: projecting the radar points in the radar target level result point set onto the pixel plane view using the homography matrix to obtain a second radar point set; Matching the second radar point set with the visual target level result point set to obtain at least one first target pair, each of the first target pairs including a radar point in the second radar point set and a pixel point in the visual target result point set; Based on the radar depth values ​​expressed by the radar points in each of the first target pairs and the homography matrix, an inverse perspective transformation is performed on the pixel points in each of the first target pairs to obtain the first radar point set, wherein one radar point in the first radar point set is associated with one pixel point in the first target pair.

11. The device according to claim 9 or 10, characterized in that When the processing module performs calculation based on the first radar point set and the visual target level result point set to obtain the target error, the processing module is specifically configured to: Projecting pixel points in the visual target-level result point set onto the radar plane view using the homography matrix to obtain a first visual point set, and performing calculation based on the first visual point set and the first radar point set to obtain the first error; And / or, using the homography matrix, projecting radar points in the first radar point set onto the pixel plane view to obtain a third radar point set, and performing calculation based on the third radar point set and the visual target level result point set to obtain the second error.

12. The device according to any one of claims 8 to 11, characterized in that: After updating the homography matrix between the visual sensor and the radar, the processing module is further configured to: The updated homography matrix is ​​used to associate the pixel points in the visual target level result point set with the radar points in the radar target level result point set, and based on the association result, the pixel points in the visual target level result point set and the radar points in the radar target level result point set are fused.

13. The device according to claim 12, characterized in that When the processing module uses the updated homography matrix to associate the pixel points in the visual target level result point set with the radar points in the radar target level result point set, the processing module is specifically configured to: Based on the updated homography matrix, projecting the radar points in the radar target level result point set onto the pixel plane view to obtain a fourth radar point set; Matching the fourth radar point set with the visual target level result point set to obtain at least one second target pair, each of the second target pairs including a radar point in the fourth radar point set and a pixel point in the visual target result point set; Projecting pixel points in each second target pair onto the radar plane view based on the radar depth value expressed by the radar point in each second target pair and the updated homography matrix.

14. The device according to claim 13, characterized in that The processing module is further configured to: Based on the updated homography matrix, project all target pixel points in the visual target-level result point set that do not match the radar points in the fourth radar point set onto the radar plane view; In the radar plane view, the target pixel point is associated with a radar point in the radar target level result point set.

15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that The method comprises computer program instructions, which, when executed by a computing device cluster, enable the computing device cluster to perform the method according to any one of claims 1 to 7, wherein the computing device cluster comprises at least one computing device.

17. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 7, wherein the computing device cluster includes at least one computing device.