Target Recognition Method and System Based on Spatiotemporal Fusion of Radar Data and Camera Data
The spatiotemporal fusion of radar and camera data improves target recognition accuracy and reliability by integrating millimeter wave radar and camera data to overcome weather interference and enhance target detection.
Patent Information
- Application Number
- CN202510488644.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-18
AI Technical Summary
In the prior art, the vehicle surrounding environmental information collected by a single type of sensor is low in credibility, resulting in the inability to accurately and timely identify road targets, and the rate of false detection and missed detection is high.
By fusing millimeter wave radar data and camera data, the spatiotemporal fusion processing method is used, including acquiring image frames and point cloud data, performing algorithm model calculations, spatiotemporal fusion, tracking calibration and projection algorithms, identifying target distance, type and collision time, and performing information fusion processing to improve recognition accuracy.
It reduces the false detection and missed detection rates, improves the accuracy and reliability of target detection, and can accurately identify road targets under different environmental conditions.
Smart Images

Figure CN120044517B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of measurement, and in particular, to a target recognition method and system based on spatio-temporal fusion of radar data and camera data. Background Art
[0002] Intelligent driving often uses different types of sensors to collect the surrounding environment of the vehicle. One is a camera to obtain image information, and the other is a millimeter-wave radar to obtain directional target distance information. The camera can obtain environmental scene information and has great advantages in aspects such as lane departure and vehicle and pedestrian recognition. The millimeter-wave radar sensor is a radar with an operating frequency selected at 77 GHz. It has strong anti-interference ability, high resolution, accurate speed measurement, and the characteristics of working all day and all weather. Therefore, the millimeter-wave radar has excellent advantages in target distance and front anti-collision sensing.
[0003] In the related art, using a camera to collect the surrounding environment of the vehicle will be affected by complex weather such as heavy rain and fog to a certain extent, and the credibility of the target speed measurement is relatively low. Using a millimeter-wave radar to measure and collect the surrounding environment of the vehicle cannot accurately identify the detailed features of the objects in front of the vehicle and environmental factors such as sound and light. In the above ways, the credibility of the surrounding environment information of the vehicle collected by a single type of sensor is relatively low, resulting in the inability to accurately and timely identify road targets. Summary of the Invention
[0004] The present application provides a target recognition method and system based on spatio-temporal fusion of radar data and camera data, which solves the problem that the credibility of the surrounding environment information of the vehicle collected by a single type of sensor in the prior art is relatively low, resulting in the inability to accurately and timely identify road targets, and can reduce the false detection and missed detection rates of a single sensor, and improve the accuracy and reliability of target detection.
[0005] In a first aspect, the present application provides a target recognition method based on spatio-temporal fusion of radar data and camera data, including:
[0006] Obtaining a first image frame collected by a camera at preset time intervals and multiple frames of point cloud data detected and output by a millimeter-wave radar within the preset time intervals;
[0007] Performing algorithm model calculation on the multiple frames of point cloud data to obtain first point cloud data, performing spatio-temporal fusion processing on the first image frame and the first point cloud data, and performing tracking calibration algorithm and projection algorithm calculations on the spatio-temporal fusion processing results respectively to obtain first target recognition information, where the first target recognition information includes the first target distance, the first target type, and the first collision time of a first target object;
[0008] Fuse the first target recognition information with at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion process, output the first target recognition information as the target recognition result. The current time period consists of at least two of the preset time intervals.
[0009] Optionally, the spatio-temporal fusion processing result includes the first three-dimensional coordinates of the first target object in the millimeter-wave radar coordinate system, the relative velocity of the first target object, and the image frame recognition features. Calculating the first target recognition information by respectively performing a tracking calibration algorithm and a projection algorithm on the spatio-temporal fusion processing result includes:
[0010] Perform fusion calculation on the first three-dimensional coordinates according to the compensation coordinates between the millimeter-wave radar and the camera to obtain the relative position of the first target object in the camera coordinate system, and calculate the first target distance of the first target object according to the relative position.
[0011] Calculate the first target type of the first target object according to the image frame recognition features and the relative velocity, calculate the relative acceleration difference of the first target object, and calculate the first collision time of the first target object according to the relative velocity, relative acceleration difference, and the first target distance.
[0012] Optionally, before performing the fusion calculation on the first three-dimensional coordinates according to the compensation coordinates between the millimeter-wave radar and the camera to obtain the relative position of the first target object in the camera coordinate system, it further includes:
[0013] Determine the second three-dimensional coordinates of the camera in the millimeter-wave radar coordinate system according to the actual installation position of the camera, and calculate the compensation coordinates between the millimeter-wave radar and the camera according to the second three-dimensional coordinates.
[0014] Optionally, before calculating the first point cloud data by performing an algorithm model on the multi-frame point cloud data, it further includes:
[0015] Obtain the current environmental information, and perform weight allocation on the first image frame and the multi-frame point cloud data based on the environmental information to determine the corresponding confidence levels in the algorithm model calculation, spatio-temporal fusion processing, tracking calibration algorithm, and projection algorithm calculation.
[0016] Optionally, the environmental information includes the meteorological type and meteorological grade. Performing the weight allocation on the first image frame and the multi-frame point cloud data based on the environmental information includes:
[0017] Query the basic weights corresponding to the first image frame and the multiple frames of point cloud data according to the meteorological type, and determine the adjustment weight coefficient associated with the meteorological type;
[0018] Calculate the adjusted weight according to the meteorological grade and the adjustment weight coefficient, and adjust the basic weights of the first image frame and the multiple frames of point cloud data respectively according to the adjusted weight to obtain the corresponding final weights.
[0019] Optionally, the method further includes:
[0020] Fuse the currently calculated second target recognition information in the next time period with at least two other target recognition information calculated in the next time period;
[0021] Fuse the result of the fusion processing corresponding to the next time period with the result of the fusion processing corresponding to the current time period to obtain an updated fusion processing result, and perform an association judgment on the second target recognition information based on the updated fusion processing result.
[0022] Optionally, after outputting the first target recognition information as the target recognition result, it further includes:
[0023] Determine whether the target recognition result meets the preset display condition and the preset alarm condition. When the preset display condition is met, determine the target display color according to the target recognition result, and send the target recognition result and the target display color to the central control screen for target display;
[0024] When the preset alarm condition is met, send an alarm instruction to the sound alarm to achieve sound alarm.
[0025] In a second aspect, the present application further provides a target recognition device based on the spatio-temporal fusion of radar data and camera data, including:
[0026] An acquisition module, configured to acquire a first image frame collected by a camera and multiple frames of point cloud data detected and output by a millimeter-wave radar within the preset time interval at each preset time interval;
[0027] A model calculation module, configured to perform algorithm model calculation on the multiple frames of point cloud data to obtain first point cloud data;
[0028] A spatio-temporal fusion processing module, configured to perform spatio-temporal fusion processing on the first image frame and the first point cloud data;
[0029] An algorithm calculation module, configured to perform a tracking calibration algorithm and a projection algorithm on the spatio-temporal fusion processing result respectively to obtain first target recognition information, where the first target recognition information includes a first target distance, a first target type, and a first collision time of a first target object;
[0030] An information fusion processing module, configured to perform fusion processing on the first target recognition information and at least two other target recognition information calculated within the current time period;
[0031] An information output module, configured to output the first target recognition information as a target recognition result when the first target recognition information is associated with the result of the fusion processing, where the current time period is composed of at least two of the preset time intervals.
[0032] In a third aspect, the present application further provides a target recognition device based on spatio-temporal fusion of radar data and camera data, and the device includes:
[0033] One or more processors;
[0034] A storage device, configured to store one or more programs,
[0035] When the one or more programs are executed by the one or more processors, the one or more processors implement the target recognition method based on spatio-temporal fusion of radar data and camera data according to the present application.
[0036] In a fourth aspect, the present application further provides a storage medium storing computer-executable instructions, and the computer-executable instructions are used to execute the target recognition method based on spatio-temporal fusion of radar data and camera data according to the present application when executed by a computer processor.
[0037] In this application, the first image frame collected by a camera and multiple frames of point cloud data detected and output by a millimeter-wave radar within a preset time interval are obtained at each preset time interval. The multiple frames of point cloud data are calculated by an algorithm model to obtain the first point cloud data. The first image frame and the first point cloud data are subjected to spatio-temporal fusion processing. The spatio-temporal fusion processing results are respectively calculated by a tracking calibration algorithm and a projection algorithm to obtain the first target recognition information. The first target recognition information includes the first target distance, the first target type, and the first collision time of the first target object. The first target recognition information is fused with at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion processing, the first target recognition information is output as the target recognition result. The current time period consists of at least two preset time intervals. This solution identifies road targets through the spatio-temporal fusion processing of the data collected by the millimeter-wave radar and the camera, solves the problem in the prior art that the vehicle surrounding environment information collected by a single type of sensor has low credibility, resulting in the inability to accurately and timely identify road targets, and can reduce the false detection and missed detection rates of a single sensor, improving the accuracy and reliability of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a flowchart of a target recognition method based on spatio-temporal fusion of radar data and camera data provided by an embodiment of the present application;
[0039] Figure 2 is a schematic diagram of the time axis of data collected by a millimeter-wave radar and a camera provided by an embodiment of the present application;
[0040] Figure 3 is a flowchart of a target recognition method based on spatio-temporal fusion of radar data and camera data including the calculation of the first target recognition information provided by an embodiment of the present application;
[0041] Figure 4 is a schematic diagram of the installation positions of a millimeter-wave radar and a camera provided by an embodiment of the present application;
[0042] Figure 5 is a flowchart of a target recognition method based on spatio-temporal fusion of radar data and camera data including weight assignment provided by an embodiment of the present application;
[0043] Figure 6 is a flowchart of a target recognition method based on spatio-temporal fusion of radar data and camera data including association judgment provided by an embodiment of the present application;
[0044] Figure 7 is a flowchart of a target recognition method based on spatio-temporal fusion of radar data and camera data including screen display and voice alarm provided by an embodiment of the present application;
[0045] Figure 8 It is a block diagram of the module structure of an object recognition device based on the spatio-temporal fusion of radar data and camera data provided by an embodiment of the present application;
[0046] Figure 9 It is a schematic structural diagram of an object recognition device based on the spatio-temporal fusion of radar data and camera data provided by an embodiment of the present application. Specific embodiments
[0047] The following further elaborates on the embodiments of the present application in conjunction with the accompanying drawings and examples. It can be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, rather than limiting the embodiments of the present application. Additionally, it should be noted that for ease of description, only parts related to the embodiments of the present application are shown in the accompanying drawings, rather than all structures.
[0048] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0049] An object recognition method based on the spatio-temporal fusion of radar data and camera data provided by an embodiment of the present application can be applied to the driving scenario of intelligent vehicles. An object recognition method based on the spatio-temporal fusion of radar data and camera data provided by an embodiment of the present application, and the execution entity of each step is a server device.
[0050] Figure 1 It is a flowchart of an object recognition method based on the spatio-temporal fusion of radar data and camera data provided by an embodiment of the present application, as Figure 1 shown, and specifically includes:
[0051] Step S101: Obtain the first image frame collected by the camera at every preset time interval and multiple frames of point cloud data detected and output by the millimeter-wave radar within the preset time interval, and perform algorithm model calculation on the multiple frames of point cloud data to obtain the first point cloud data.
[0052] Among them, the preset time interval can be the time taken for the camera to capture one frame of image. This preset time interval can be flexibly set based on the configuration information of the camera. Every time this preset time interval elapses, the first image frame captured by the camera and multiple frames of point cloud data detected and output by the millimeter-wave radar within this preset time interval can be obtained. Exemplarily, as Figure 2 shown Figure 2 is a schematic diagram of the time axis for the millimeter-wave radar and the camera to collect data provided by an embodiment of the present application. Among them, the preset time interval is 33 ms, that is, the camera captures one frame of image every 33 ms. 01, 02, 03, 03 are the image frames captured by the camera at the corresponding time intervals. Fr1, Fr2... Frn are multiple frames of point cloud data detected and output by the millimeter-wave radar from the 33rd ms to the 66th ms. The point cloud data refers to a series of information such as distance, speed, and angle obtained by the millimeter-wave radar. These information are represented in the three-dimensional space in the form of points, and each point contains three-dimensional coordinates (x, y, z). Using the multiple frames of point cloud data detected and output by the millimeter-wave radar within the preset time interval for algorithm model calculation, the first point cloud data can be obtained. The first point cloud data is used to represent the single-frame point cloud data obtained after fusing multiple frames of point cloud data. In one embodiment, a calculation method for the first point cloud data can be to call the AI module in the self-developed chip to perform convolution and pooling algorithm model calculations on the multiple frames of point cloud data detected and output by the millimeter-wave radar in sequence to obtain the corresponding processed data, and input the processed data corresponding to the multiple frames of point cloud data into the time series model for fusion to obtain the first point cloud data.
[0053] Step S102: Perform spatio-temporal fusion processing on the first image frame and the first point cloud data, and perform tracking calibration algorithm and projection algorithm calculations on the spatio-temporal fusion processing results respectively to obtain the first target recognition information, where the first target recognition information includes the first target distance, the first target type, and the first collision time of the first target object.
[0054] Among them, the tracking and calibration algorithm is a technology that continuously tracks the motion state of an object and calibrates its attributes (such as position, speed, collision time, etc.) through multi-frame data association and dynamic model estimation. The projection algorithm is a coordinate transformation technology that maps radar point cloud data from three-dimensional space to a two-dimensional image plane, and combines visual semantic information to identify the static attributes of an object (such as category, size, pose). Using the tracking and calibration algorithm and the projection algorithm, the first target distance, the first target type, and the first collision time of the first target object can be calculated. The first target object can be a key target preferentially identified through spatio-temporal fusion processing. The first target distance is used to represent the relative distance between the first target object and the vehicle. The first target type is used to represent the classification of the first target object, such as classifications of cars, trucks, pedestrians, bicycles, etc. The first collision time is used to represent the time required for the vehicle to collide with the first target object. In one embodiment, the spatio-temporal fusion processing result is input into the tracking and calibration algorithm and the projection algorithm for calculation to obtain the first target distance, the first target type, and the first collision time of the first target object.
[0055] Step S103: Fuse the first target recognition information with at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion processing, output the first target recognition information as the target recognition result, where the current time period consists of at least two preset time intervals.
[0056] Among them, the time period can be a preset cyclic period, and the time length of the time period is greater than a preset time interval and is an integer multiple of the preset time interval. After calculating the first target recognition information, the first target recognition information is fused with at least two other target recognition information calculated within the current time period. The result of the fusion process can be used to determine whether the first target recognition information is successfully associated. In the case of successful association, the first target recognition information is output as the target recognition result. In one embodiment, a way of association judgment can be that after the fusion process, the similarity between the first target recognition information and the result of the fusion process is calculated. If the similarity is greater than the preset similarity, it is determined that the first target recognition information is associated with the result of the fusion process; otherwise, it is determined that the first target recognition information is not associated with the result of the fusion process. Among them, the result of the fusion process is the overall fused target recognition information obtained by fusing the first target recognition information with at least two other target recognition information calculated within the current time period. Exemplarily, the time period consists of three preset time intervals, arranged in chronological order as preset time interval 1, preset time interval 2, and preset time interval 3, corresponding to the calculated target recognition information 1, target recognition information 2, and target recognition information 3 respectively. The current time is preset time interval 3. Then, the target recognition information 3 is fused with the target recognition information 1 and the target recognition information 2 to obtain the fused target recognition information. The preset similarity is 90%. The similarity between the target recognition information 3 and the fused target recognition information is calculated to be 98%, which is greater than the preset similarity. It is determined that the target recognition information 3 is associated with the fused target recognition information, and the target recognition information 3 is output as the target recognition result.
[0057] In another embodiment, a way of association judgment can be that after the fusion process, the difference data between the first target recognition information and the result of the fusion process is calculated. If the difference data is less than the preset difference standard, it is determined that the first target recognition information is associated with the result of the fusion process; otherwise, it is determined that the first target recognition information is not associated with the result of the fusion process. The difference data includes the target distance difference value, the target type difference, and the collision time difference value.
[0058] In one embodiment, in the case where the first target recognition information is not associated with the result of the fusion process, the first target recognition information is determined as an invalid target recognition result and deleted, and the target recognition for the next time period is continued.
[0059] As can be seen from the above, the first image frame collected by the camera and multiple frames of point cloud data detected and output by the millimeter-wave radar within a preset time interval are obtained at each preset time interval. The multiple frames of point cloud data are calculated by an algorithm model to obtain the first point cloud data. The first image frame and the first point cloud data are subjected to spatio-temporal fusion processing. The tracking calibration algorithm and the projection algorithm are respectively calculated on the spatio-temporal fusion processing result to obtain the first target recognition information. The first target recognition information includes the first target distance, the first target type, and the first collision time of the first target object. The first target recognition information is fused with at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion processing, the first target recognition information is output as the target recognition result. The current time period consists of at least two preset time intervals. This solution identifies road targets through the spatio-temporal fusion processing of the data collected by the millimeter-wave radar and the camera, solves the problem in the prior art that the credibility of the vehicle surrounding environment information collected by a single type of sensor is relatively low, resulting in the inability to accurately and timely identify road targets, and can reduce the false detection and missed detection rates of a single sensor, and improve the accuracy and reliability of target detection.
[0060] Figure 3 The figure is a flowchart of a target recognition method based on the spatio-temporal fusion of radar data and camera data including the calculation of the first target recognition information provided by an embodiment of the present application, as Figure 3 shown, and specifically includes:
[0061] Step S201: Obtain the first image frame collected by the camera and multiple frames of point cloud data detected and output by the millimeter-wave radar within a preset time interval at each preset time interval, and calculate the multiple frames of point cloud data by an algorithm model to obtain the first point cloud data.
[0062] Step S202: Perform spatio-temporal fusion processing on the first image frame and the first point cloud data, and perform fusion calculation on the first three-dimensional coordinates of the first target object in the millimeter-wave radar coordinate system according to the compensation coordinates between the millimeter-wave radar and the camera to obtain the relative position of the first target object in the camera coordinate system, and calculate the first target distance of the first target object according to the relative position.
[0063] Among them, the spatio-temporal fusion processing result includes the first three-dimensional coordinates of the first target object in the millimeter-wave radar coordinate system, the relative speed of the first target object, and the image frame recognition feature. The millimeter-wave radar coordinate system is a coordinate system with the installation position of the millimeter-wave radar as the origin. The camera coordinate system is a coordinate system with the installation position of the camera as the origin. Exemplarily, as Figure 4 shown, Figure 4A schematic diagram of the installation positions of a millimeter-wave radar and a camera provided by an embodiment of the present application. The coordinate system with Oc as the origin is the camera coordinate system, and the coordinate system with Or as the origin is the millimeter-wave radar coordinate system. Other installation positions that meet the system requirements can also be selected, and only the coordinate system configuration needs to be modified. The compensation coordinates between the millimeter-wave radar and the camera can be the coordinate transformation matrix between the two. This compensation coordinate can be used to map the target position in the millimeter-wave radar coordinate system to the camera coordinate system to achieve unified spatial reference. The first target distance of the first target object can be calculated using the relative position of the first target object in the camera coordinate system. Exemplarily, the first three-dimensional coordinate of the first target object in the millimeter-wave radar coordinate system is (x1, y1, z1). The relative position of the first target object in the camera coordinate system is obtained by fusing and calculating the first three-dimensional coordinate according to the compensation coordinates between the millimeter-wave radar and the camera as (x2, y2, z2). The straight-line distance from this relative position to the origin of the camera coordinate system is calculated and determined as the first target distance of the first target object.
[0064] Optionally, before fusing and calculating the first three-dimensional coordinate according to the compensation coordinates between the millimeter-wave radar and the camera to obtain the relative position of the first target object in the camera coordinate system, the second three-dimensional coordinate of the camera in the millimeter-wave radar coordinate system is determined according to the actual installation position of the camera, and the compensation coordinates between the millimeter-wave radar and the camera are calculated according to the second three-dimensional coordinate. Exemplarily, the millimeter-wave radar is installed at the center of the front bumper of the vehicle, and the camera is installed near the rearview mirror. The camera is horizontally offset to the right by Δx = 0.3 m and vertically shifted up by Δz = 1.2 m relative to the millimeter-wave radar, with no longitudinal offset, i.e., Δy = 0 (i.e., the second three-dimensional coordinate of the camera in the millimeter-wave radar coordinate system is (0.3, 0, 1.2)). The pitch angle (around the x-axis) of the camera relative to the millimeter-wave radar is tilted downward by -5°, with no yaw angle and roll angle. The rotation matrix R is calculated according to the camera's pitch rotation of -5° around the x-axis. The translation vector T (-0.3, 0, -1.2) is obtained by transforming the second three-dimensional coordinate (0.3, 0, 1.2) of the camera in the millimeter-wave radar coordinate system. The compensation coordinates between the millimeter-wave radar and the camera are composed of the rotation matrix and the translation vector to obtain , where is the coordinate after being converted to the camera coordinate system, is the coordinate of the target object in the millimeter-wave radar coordinate system.
[0065] Step S203: Calculate the first target type of the first target object according to the recognition features and relative speed of the image frame of the first target object, calculate the relative acceleration difference of the first target object, and calculate the first collision time of the first target object according to the relative speed, relative acceleration difference, and first target distance of the first target object.
[0066] Among them, the image frame recognition feature is the visual feature of the first target object obtained by recognizing the first image frame, such as features like color and shape. The relative speed is the speed of the first target object relative to the current driving vehicle. Using this image frame recognition feature and the relative speed, the first target type of the first target object can be calculated. Exemplarily, a target object is detected in front of the autonomous driving vehicle. The image frame recognition feature of the target object is a two-wheel structure, and the relative speed is 8 m / s. The target types of the two-wheel structure include bicycles and electric bicycles. The relative speed of 8 m / s conforms to the speed feature of an electric bicycle. That is, it is comprehensively determined that the target type of the target object is an electric bicycle. The relative acceleration difference can be the acceleration of the first target object relative to the current driving vehicle. Using this relative acceleration difference, this relative speed, and the first target distance, the first collision time of the first target object can be calculated. In one embodiment, a way to calculate the collision time can be to substitute the relative speed, relative acceleration difference, and first target distance of the first target object into the collision time calculation formula to obtain the first collision time of the first target object, where the collision time calculation formula is , where v is the relative speed of the first target object, a is the relative speed difference of the first target object, and S is the first target distance of the first target object.
[0067] Step S204: Fuse the first target recognition information with at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion processing, output the first target recognition information as the target recognition result, where the current time period consists of at least two preset time intervals.
[0068] As can be seen from the above, after performing spatio-temporal fusion processing on the first image frame and the first point cloud data, the relative position of the first target object in the camera coordinate system is obtained by fusing and calculating the first three-dimensional coordinates of the first target object in the millimeter-wave radar coordinate system according to the compensation coordinates between the millimeter-wave radar and the camera. The first target distance of the first target object is calculated based on the relative position. The first target type of the first target object is calculated based on the image frame recognition feature and relative speed of the first target object. The relative acceleration difference of the first target object is calculated. The first collision time of the first target object is calculated based on the relative speed, relative acceleration difference, and first target distance of the first target object. This solution calculates the distance, type, and collision time of the target object through the spatio-temporal fusion processing result of the data collected by the millimeter-wave radar and the camera, and can improve the accuracy of the recognized target information.
[0069] Figure 5The flowchart of an object recognition method based on spatio-temporal fusion of radar data and camera data with weight assignment provided by an embodiment of this application is as follows Figure 5 shown, and specifically includes:
[0070] Step S301: Obtain the first image frame collected by the camera at every preset time interval and multiple frames of point cloud data detected and output by the millimeter-wave radar within the preset time interval, obtain the current environmental information, and perform weight assignment on the first image frame and the multiple frames of point cloud data based on the environmental information, so as to determine the corresponding confidence levels in algorithm model calculation, spatio-temporal fusion processing, tracking calibration algorithm, and projection algorithm calculation.
[0071] Among them, the environmental information may be the environmental information of the current driving vehicle, such as weather information, light information, etc. The environmental information can be used to perform weight assignment on the first image frame and the multiple frames of point cloud data. The weights of the first image frame and the multiple frames of point cloud data can adjust the confidence levels of the two data sources in the processes of algorithm model calculation, spatio-temporal fusion processing, tracking calibration algorithm, and projection algorithm calculation, so as to optimize the accuracy and robustness of the results.
[0072] Optionally, the environmental information includes meteorological type and meteorological grade. One way of weight assignment can be to query the basic weights corresponding to the first image frame and the multiple frames of point cloud data according to the meteorological type, determine the adjustment weight coefficient associated with the meteorological type, calculate the adjustment weight according to the meteorological grade and the adjustment weight coefficient, and adjust the basic weights of the first image frame and the multiple frames of point cloud data respectively according to the adjustment weight to obtain the corresponding final weights. Among them, the meteorological type is different weather phenomena, such as rainfall, haze and other meteorological types. The meteorological grade is used to represent the severity of the meteorological type. For example, the meteorological grade is divided into the first grade, the second grade, and the third grade, and the corresponding severity increases gradually. Exemplarily, the current meteorological type is rainfall, the rainfall grade is 3, the queried basic weight of the multiple frames of point cloud data during rainfall is 0.7, the basic weight of the first image frame is 0.3, the adjustment weight coefficient associated with rainfall is 0.05, multiplying the rainfall grade by the weight coefficient gives an adjustment weight of 0.15. Since the camera will be affected to a certain extent by complex weather such as heavy rain and fog, adding the basic weight of the multiple frames of point cloud data and the adjustment weight gives a final weight of 0.85 for the multiple frames of point cloud data, and subtracting the adjustment weight from the basic weight of the first image frame gives a final weight of 0.15 for the first image frame.
[0073] In another embodiment, query the preset weight mapping table corresponding to the meteorological type according to the meteorological grade to obtain the weights of the first image frame and the multiple frames of point cloud data, where the preset weight mapping table includes the mapping relationship between different meteorological grades and the weights of the first image frame and the multiple frames of point cloud data.
[0074] Step S302: Perform algorithm model calculation on multiple frames of point cloud data to obtain the first point cloud data. Perform spatio-temporal fusion processing on the first image frame and the first point cloud data, and perform tracking calibration algorithm and projection algorithm calculations on the spatio-temporal fusion processing results respectively to obtain the first target recognition information, where the first target recognition information includes the first target distance, the first target type, and the first collision time of the first target object.
[0075] Step S303: Perform fusion processing on the first target recognition information and at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion processing, output the first target recognition information as the target recognition result, where the current time period consists of at least two preset time intervals.
[0076] As can be seen from the above, before performing algorithm model calculation on multiple frames of point cloud data to obtain the first point cloud data, obtain the current environmental information, and perform weight allocation on the first image frame and multiple frames of point cloud data based on the environmental information, so as to determine the corresponding confidence levels in algorithm model calculation, spatio-temporal fusion processing, tracking calibration algorithm, and projection algorithm calculation. This solution can optimize the accuracy of target recognition by sensor data in different environments through different weighted configurations of the acquisition data of millimeter-wave radar and camera in the actual environment.
[0077] Figure 6 This is a flowchart of a target recognition method based on spatio-temporal fusion of radar data and camera data including association judgment provided by an embodiment of the present application. As Figure 6 shown, it specifically includes:
[0078] Step S401: Obtain the first image frame collected by the camera at every preset time interval and multiple frames of point cloud data detected and output by the millimeter-wave radar within the preset time interval, and perform algorithm model calculation on the multiple frames of point cloud data to obtain the first point cloud data.
[0079] Step S402: Perform spatio-temporal fusion processing on the first image frame and the first point cloud data, and perform tracking calibration algorithm and projection algorithm calculations on the spatio-temporal fusion processing results respectively to obtain the first target recognition information, where the first target recognition information includes the first target distance, the first target type, and the first collision time of the first target object.
[0080] Step S403: Perform fusion processing on the first target recognition information and at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion processing, output the first target recognition information as the target recognition result, where the current time period consists of at least two preset time intervals.
[0081] Step S404: Fuse the second target recognition information currently calculated in the next time period with at least two other target recognition information calculated within the next time period, fuse the result of the fusion processing corresponding to the next time period with the result of the fusion processing corresponding to the current time period to obtain an updated fusion processing result, and perform an association judgment on the second target recognition information based on the updated fusion processing result.
[0082] Among them, the next time period is the next period of the current time period. Every other time period, fuse the result of the fusion processing of this time period with the result of the fusion processing of the previous time period to obtain the target recognition information finally used for association judgment. In one embodiment, when it is determined that the second target recognition information is associated with the updated fusion processing result, the second target recognition information is output as the target recognition result of the next time period. Exemplarily, the time period consists of three preset time intervals. The result of the fusion processing corresponding to the current time period is the fused target recognition information 1. The next time period consists of preset time interval 1, preset time interval 2, and preset time interval 3. Currently, in preset time interval 3 of the next time period, the target recognition information a is calculated, the target recognition information b is calculated in preset time interval 1, and the target recognition information c is calculated in preset time interval 2. Fuse the target recognition information a, target recognition information b, and target recognition information c to obtain the fused target recognition information 2. Fuse the fused target recognition information 1 with the fused target recognition information 2 to obtain the updated fusion processing result. Calculate the similarity between the target recognition information a and the updated fusion processing result. If the similarity is greater than the preset similarity, it is determined that the target recognition information a is associated with the updated fusion processing result, and the target recognition information a is output as the target recognition result of the next preset time period.
[0083] As can be seen from the above, in the next time period of the current period, fuse the second target recognition information currently calculated in the next time period with at least two other target recognition information calculated within the next time period, fuse the result of the fusion processing corresponding to the next time period with the result of the fusion processing corresponding to the current time period to obtain the updated fusion processing result, and perform an association judgment on the second target recognition information based on the updated fusion processing result. By continuously circularly recording and fusing data features, this solution can improve the accuracy of subsequent algorithm results.
[0084] Figure 7 It is a flowchart of a target recognition method based on spatio-temporal fusion of radar data and camera data including screen display and sound alarm provided by an embodiment of the present application. As Figure 7 shown, it specifically includes:
[0085] Step S501: Obtain the first image frame collected by the camera and multiple frames of point cloud data detected and output by the millimeter-wave radar within a preset time interval at every preset time interval, and perform algorithm model calculation on the multiple frames of point cloud data to obtain the first point cloud data.
[0086] Step S502: Perform spatio-temporal fusion processing on the first image frame and the first point cloud data, and perform tracking calibration algorithm and projection algorithm calculations on the spatio-temporal fusion processing results respectively to obtain the first target recognition information, where the first target recognition information includes the first target distance, the first target type, and the first collision time of the first target object.
[0087] Step S503: Perform fusion processing on the first target recognition information and at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion processing, output the first target recognition information as the target recognition result, where the current time period consists of at least two preset time intervals.
[0088] Step S504: Determine whether the target recognition result meets the preset display condition and the preset alarm condition. When the preset display condition is met, determine the target display color according to the target recognition result, and send the target recognition result and the target display color to the central control screen for target display. When the preset alarm condition is met, send an alarm instruction to the sound alarm to achieve sound alarm.
[0089] Among them, the preset display condition can be a condition pre-set for determining whether to display the target recognition result. When the preset display condition is met, the target display color can be determined using the target recognition result. The target display color is the color of the target to be displayed on the central control screen. The preset alarm condition can be a condition pre-set for determining whether a sound alarm is required for the target recognition result. In one embodiment, a way to determine the condition judgment can be to query the corresponding first distance standard value, second distance standard value, first collision time standard value, and second collision time standard value according to the first target type. When the first target distance is less than the first distance standard value and the first collision time is less than the first collision time standard value, it is determined that the target recognition result meets the preset display condition. When the first target distance is less than the second distance standard value and the first collision time is less than the second collision time standard value, it is determined that the target recognition result meets the preset alarm condition, where the first distance standard value is greater than the second distance standard value, and the first collision time standard value is greater than the second collision time standard value. In another embodiment, a way to determine the condition judgment can be to query the corresponding distance standard value and collision time standard value according to the first target type. When the first target distance is less than the distance standard value and the first collision time is less than the collision time standard value, it is determined that the target recognition result meets the preset display condition and the preset alarm condition.
[0090] In one embodiment, a way to determine the target display color can be to input the target recognition result into a trained danger level assessment model to obtain the corresponding danger level, and query the corresponding target display color according to the danger level of the target recognition result. In another embodiment, a way to determine the target display color can be to query the corresponding preset color according to the target type in the target recognition result and determine it as the target display color.
[0091] In one embodiment, a way to generate an alarm instruction can be to perform a danger level assessment on the target recognition result to obtain the corresponding danger level, query the corresponding alarm sound type according to the danger level of the target recognition result, and generate the corresponding alarm instruction according to the alarm sound type. In another embodiment, a way to generate an alarm instruction can be to input the target recognition result into a preset alarm template to obtain the corresponding alarm voice text, and generate the corresponding alarm instruction according to the alarm voice text.
[0092] As can be seen from the above, after the first target recognition information is output as the target recognition result, it is determined whether the target recognition result meets the preset display condition and the preset alarm condition. When the preset display condition is met, the target display color is determined according to the target recognition result, and the target recognition result and the target display color are sent to the central control screen for target display. When the preset alarm condition is met, an alarm instruction is sent to the sound alarm to implement sound alarm. This solution determines whether to perform target display and alarm based on the target recognition result, and can accurately and timely provide auxiliary functions for the driver to ensure driving safety.
[0093] Figure 8 This is a module structure block diagram of a target recognition device based on the spatio-temporal fusion of radar data and camera data provided by an embodiment of the present application. This system is used to execute a target recognition method based on the spatio-temporal fusion of radar data and camera data provided by the above embodiment, and has corresponding functional modules and beneficial effects for executing the method. As Figure 8 shown, this system specifically includes:
[0094] An acquisition module 101, configured to acquire a first image frame collected by a camera at each preset time interval and multiple frames of point cloud data detected and output by a millimeter-wave radar within the preset time interval;
[0095] A model calculation module 102, configured to perform algorithm model calculation on the multiple frames of point cloud data to obtain first point cloud data;
[0096] A spatio-temporal fusion processing module 103, configured to perform spatio-temporal fusion processing on the first image frame and the first point cloud data;
[0097] An algorithm calculation module 104, configured to perform tracking calibration algorithm and projection algorithm calculation on the spatio-temporal fusion processing result respectively to obtain first target recognition information, where the first target recognition information includes a first target distance, a first target type, and a first collision time of a first target object;
[0098] An information fusion processing module 105, configured to perform fusion processing on the first target recognition information and at least two other target recognition information calculated within the current time period;
[0099] An information output module 106, configured to output the first target recognition information as the target recognition result when the first target recognition information is associated with the result of the fusion processing, where the current time period is composed of at least two of the preset time intervals.
[0100] As can be seen from the above solution, the first image frame collected by the camera and multiple frames of point cloud data detected and output by the millimeter-wave radar within a preset time interval are obtained at each preset time interval. The algorithm model is used to calculate the first point cloud data from the multiple frames of point cloud data. The first image frame and the first point cloud data are subjected to spatio-temporal fusion processing. The tracking calibration algorithm and the projection algorithm are respectively used to calculate the first target recognition information from the spatio-temporal fusion processing result. The first target recognition information includes the first target distance, the first target type, and the first collision time of the first target object. The first target recognition information is fused with at least two other target recognition information calculated within the current time period. When the first target recognition information is associated with the result of the fusion processing, the first target recognition information is output as the target recognition result. The current time period consists of at least two preset time intervals. This solution identifies road targets through the spatio-temporal fusion processing of the data collected by the millimeter-wave radar and the camera, solves the problem in the prior art that the vehicle surrounding environment information collected by a single type of sensor has low credibility, resulting in the inability to accurately and timely identify road targets, and can reduce the false detection and missed detection rates of a single sensor, improving the accuracy and reliability of target detection.
[0101] In a possible embodiment, the algorithm calculation module 104 is specifically configured to:
[0102] Perform fusion calculation on the first three-dimensional coordinates according to the compensation coordinates between the millimeter-wave radar and the camera to obtain the relative position of the first target object in the camera coordinate system, and calculate the first target distance of the first target object according to the relative position;
[0103] Calculate the first target type of the first target object according to the image frame recognition feature and the relative speed, calculate the relative acceleration difference of the first target object, and calculate the first collision time of the first target object according to the relative speed, relative acceleration difference, and the first target distance.
[0104] In a possible embodiment, the algorithm calculation module 104 is further configured to:
[0105] Determine the second three-dimensional coordinates of the camera in the millimeter-wave radar coordinate system according to the actual installation position of the camera, and calculate the compensation coordinates between the millimeter-wave radar and the camera according to the second three-dimensional coordinates.
[0106] In a possible embodiment, a weight allocation module is further included, which is specifically configured to:
[0107] Obtain the current environmental information, and perform weight assignment for the first image frame and the multi-frame point cloud data based on the environmental information, so as to determine the corresponding confidence levels in the algorithm model calculation, the spatio-temporal fusion processing, the tracking calibration algorithm, and the projection algorithm calculation.
[0108] In a possible embodiment, the weight assignment module is further configured to:
[0109] Query the basic weights corresponding to the first image frame and the multi-frame point cloud data according to the weather type, and determine the adjustment weight coefficient associated with the weather type;
[0110] Calculate the adjusted weights according to the weather grade and the adjustment weight coefficient, and adjust the basic weights of the first image frame and the multi-frame point cloud data respectively according to the adjusted weights to obtain the corresponding final weights.
[0111] In a possible embodiment, the information fusion processing module 105 is specifically configured to:
[0112] Fuse the currently calculated second target recognition information in the next time period with at least two other target recognition information calculated in the next time period;
[0113] Fuse the result of the fusion processing corresponding to the next time period with the result of the fusion processing corresponding to the current time period to obtain an updated fusion processing result;
[0114] It further includes an association judgment module, which is specifically configured to:
[0115] Perform an association judgment on the second target recognition information based on the updated fusion processing result.
[0116] In a possible embodiment, it further includes a control instruction sending module, which is specifically configured to:
[0117] Determine whether the target recognition result meets the preset display condition and the preset alarm condition. When the preset display condition is met, determine the target display color according to the target recognition result, and send the target recognition result and the target display color to the central control screen for target display;
[0118] When the preset alarm condition is met, send an alarm instruction to the sound alarm to achieve sound alarm.
[0119] Figure 9 This is a schematic structural diagram of a target recognition device based on spatio-temporal fusion of radar data and camera data provided by an embodiment of the present application, as Figure 9As shown, the device includes a processor 201, a memory 202, an input device 203, and an output device 204; the number of processors 201 in the device can be one or more, Figure 9 and one processor 201 is taken as an example herein; the processor 201, the memory 202, the input device 203, and the output device 204 in the device can be connected through a bus or other means, Figure 9 and taking connection through a bus as an example herein. The memory 202, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions or modules corresponding to a target recognition method based on spatio-temporal fusion of radar data and camera data in an embodiment of the present application. The processor 201 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 202, that is, implements the above-mentioned target recognition method based on spatio-temporal fusion of radar data and camera data. The input device 203 can be used to receive input digital or character information, and generate key signal inputs related to user settings and function controls of the device. The output device 204 can include display devices such as a display screen.
[0120] An embodiment of the present application further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a target recognition method based on spatio-temporal fusion of radar data and camera data when executed by a computer processor. The method includes:
[0121] Obtaining a first image frame collected by a camera and multiple frames of point cloud data detected and output by a millimeter-wave radar within the preset time interval at each preset time interval;
[0122] Performing algorithm model calculation on the multiple frames of point cloud data to obtain first point cloud data, performing spatio-temporal fusion processing on the first image frame and the first point cloud data, and performing tracking calibration algorithm and projection algorithm calculation on the spatio-temporal fusion processing result to obtain first target recognition information, where the first target recognition information includes a first target distance, a first target type, and a first collision time of a first target object;
[0123] Fusing the first target recognition information with at least two other target recognition information calculated within the current time period, and outputting the first target recognition information as a target recognition result when the first target recognition information is associated with the result of the fusion processing, where the current time period consists of at least two of the preset time intervals.
[0124] It should be noted that in the above embodiments of the object recognition method and system based on the spatio-temporal fusion of radar data and camera data, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present application.
[0125] Note that the above is only the preferred embodiment of the embodiments of the present application and the applied technical principle. Those skilled in the art will understand that the embodiments of the present application are not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the embodiments of the present application. Therefore, although the embodiments of the present application have been described in more detail through the above embodiments, the embodiments of the present application are not limited to the above embodiments only. Without departing from the concept of the embodiments of the present application, more other equivalent embodiments can be included, and the scope of the embodiments of the present application is determined by the scope of the appended claims.
Claims
1. A target recognition method based on spatio-temporal fusion of radar data and camera data, characterized in that Including: Obtaining a first image frame collected by a camera at each preset time interval and multiple frames of point cloud data detected and output by a millimeter-wave radar within the preset time interval; Obtaining current environmental information, and performing weight allocation for the first image frame and the multiple frames of point cloud data based on the environmental information, so as to determine corresponding confidence levels in algorithm model calculation, spatio-temporal fusion processing, tracking calibration algorithm, and projection algorithm calculation; Performing algorithm model calculation on the multiple frames of point cloud data to obtain first point cloud data, performing spatio-temporal fusion processing on the first image frame and the first point cloud data, and performing tracking calibration algorithm and projection algorithm calculation on the spatio-temporal fusion processing result respectively to obtain first target recognition information, where the first target recognition information includes a first target distance, a first target type, and a first collision time of a first target object, and the spatio-temporal fusion processing result includes a first three-dimensional coordinate of the first target object in the millimeter-wave radar coordinate system, a relative speed of the first target object, and an image frame recognition feature. Among them, performing tracking calibration algorithm and projection algorithm calculation on the spatio-temporal fusion processing result respectively to obtain first target recognition information includes: performing fusion calculation on the first three-dimensional coordinate according to the compensation coordinate between the millimeter-wave radar and the camera to obtain a relative position of the first target object in the camera coordinate system, calculating the first target distance of the first target object according to the relative position, calculating the first target type of the first target object according to the image frame recognition feature and the relative speed, calculating a relative acceleration difference of the first target object, and calculating the first collision time of the first target object according to the relative speed, relative acceleration difference, and the first target distance; Fusing the first target recognition information with at least two other target recognition information calculated within the current time period, and when the first target recognition information is associated with the result of the fusion processing, outputting the first target recognition information as a target recognition result, where the current time period consists of at least two of the preset time intervals.
2. The target recognition method based on spatio-temporal fusion of radar data and camera data according to claim 1, wherein Before performing the fusion calculation on the first three-dimensional coordinate according to the compensation coordinate between the millimeter-wave radar and the camera to obtain the relative position of the first target object in the camera coordinate system, it further includes: Determining a second three-dimensional coordinate of the camera in the millimeter-wave radar coordinate system according to the actual installation position of the camera, and calculating the compensation coordinate between the millimeter-wave radar and the camera according to the second three-dimensional coordinate.
3. The target recognition method based on spatio-temporal fusion of radar data and camera data according to claim 1, wherein, The environmental information includes a meteorological type and a meteorological grade, and the performing weight allocation for the first image frame and the multiple frames of point cloud data based on the environmental information includes: Querying the basic weights corresponding to the first image frame and the multiple frames of point cloud data according to the meteorological type, and determining an adjustment weight coefficient associated with the meteorological type; An adjustment weight is calculated based on the meteorological level and the adjustment weight coefficient, and the basic weights of the first image frame and the multi-frame point cloud data are adjusted according to the adjustment weight to obtain corresponding final weights.
4. The target recognition method based on spatio-temporal fusion of radar data and camera data according to any one of claims 1-2, characterized in that The method further includes: Fusing the currently calculated second target recognition information in the next time period with at least two other target recognition information calculated within the next time period; Fusing the result of the fusion processing corresponding to the next time period with the result of the fusion processing corresponding to the current time period to obtain an updated fusion processing result, and making an association judgment on the second target recognition information based on the updated fusion processing result.
5. The target recognition method based on spatio-temporal fusion of radar data and camera data according to any one of claims 1-2, characterized in that, After outputting the first target recognition information as the target recognition result, it further includes: Determining whether the target recognition result meets a preset display condition and a preset warning condition. When the preset display condition is met, determining a target display color according to the target recognition result, and sending the target recognition result and the target display color to the central control screen for target display; When the preset warning condition is met, sending a warning instruction to the sound warning device to achieve sound warning.
6. An object recognition system based on spatio-temporal fusion of radar data and camera data, characterized in that, It includes: An acquisition module, configured to acquire a first image frame collected by a camera and multiple frames of point cloud data detected and output by a millimeter wave radar within a preset time interval at every preset time interval; A weight distribution module, configured to acquire the current environmental information, and perform weight distribution on the first image frame and the multiple frames of point cloud data based on the environmental information, so as to determine the corresponding confidence levels in algorithm model calculation, spatio-temporal fusion processing, tracking calibration algorithm, and projection algorithm calculation; A model calculation module, configured to perform algorithm model calculation on the multiple frames of point cloud data to obtain first point cloud data; A spatio-temporal fusion processing module, configured to perform spatio-temporal fusion processing on the first image frame and the first point cloud data; An algorithm calculation module, configured to perform tracking calibration algorithm and projection algorithm calculation on the spatio-temporal fusion processing result respectively to obtain first target recognition information, where the first target recognition information includes a first target distance, a first target type, and a first collision time of a first target object, and the spatio-temporal fusion processing result includes a first three-dimensional coordinate of the first target object in the millimeter wave radar coordinate system, a relative speed of the first target object, and an image frame recognition feature. Specifically, the algorithm calculation module is configured to perform fusion calculation on the first three-dimensional coordinate according to the compensation coordinate between the millimeter wave radar and the camera to obtain the relative position of the first target object in the camera coordinate system, perform distance calculation according to the relative position to obtain the first target distance of the first target object, perform type calculation according to the image frame recognition feature and the relative speed to obtain the first target type of the first target object, calculate the relative acceleration difference of the first target object, and perform collision time calculation according to the relative speed, relative acceleration difference, and the first target distance to obtain the first collision time of the first target object; An information fusion processing module, configured to perform fusion processing on the first target recognition information and at least two other target recognition information calculated within the current time period; An information output module, configured to output the first target recognition information as a target recognition result when the first target recognition information is associated with the result of the fusion processing, where the current time period consists of at least two of the preset time intervals.
7. An object recognition device based on spatio-temporal fusion of radar data and camera data, characterized in that, The device includes: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the target recognition method based on spatio-temporal fusion of radar data and camera data according to any one of claims 1-5.
Citation Information
Patent Citations
Track and road obstacle detecting method
CN109298415A
Human body behavior detection method based on millimeter wave radar and monocular vision fusion
CN118247842A