A method, apparatus, device and medium for determining an aerial view data frame
Patent Information
- Application Number
- CN202511274214.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2045-09-08
AI Technical Summary
[0003]然而,在相机数据和激光雷达数据的采集过程中,由于激光雷达和相机的采集帧率不同(例如激光雷达的采集帧率通常为10Hz,而相机的帧率通常为30Hz),导致难以同时获取相机数据和激光雷达数据,从而严重影响了BEV算法的感知精度
[0049]This application provides a method, apparatus, device, and medium for determining bird's-eye view data frames. The method includes: acquiring first acquisition sequences of N lidars, second acquisition sequences of M cameras, and vehicle motion state acquisition sequences; wherein the first acquisition sequences include multiple frames of point cloud data and an average timestamp for each frame of point cloud data, the average timestamp being the average of the acquisition timestamps for each point; the second acquisition sequences include multiple frames of image data and an acquisition timestamp for each frame of image data; the vehicle motion state acquisition sequences include multiple vehicle motion state data and an acquisition timestamp for each vehicle motion state data; fusing the first acquisition sequences of the N lidars to obtain a fused acquisition sequence; wherein the fused acquisition sequence includes multiple frames of fused point cloud data and a fusion timestamp for each frame of fused point cloud data; and for each frame of fused point cloud data, performing the following: in M... In the second acquisition sequence of each camera, M first timestamps closest to the fusion timestamp are determined, along with M frames of target image data. In the vehicle motion state acquisition sequence, the second timestamp closest to the fusion timestamp is determined, along with the corresponding target vehicle motion state data. If the absolute value of the difference between the M first timestamps and the target timestamp is less than a first preset threshold, and the absolute value of the difference between the second timestamp and the target timestamp is less than a second preset threshold, then the acquisition timestamp of each point in the fused point cloud data is determined. The target timestamp is the average of the M first timestamps. Based on the first difference between the acquisition timestamp and the target timestamp of each point, each point is compensated to obtain compensated fused point cloud data. Based on the compensated fused point cloud data, the M frames of target image data, and the target vehicle motion state data, the bird's-eye view data frame is determined. Therefore, the method for determining bird's-eye view data frames provided in this application effectively solves the problem of frame rate differences between LiDAR and camera, thereby improving the time alignment accuracy of multi-source data, providing high-quality and highly consistent bird's-eye view data frames to support subsequent autonomous driving bird's-eye view algorithms, and ensuring the accuracy and reliability of bird's-eye view algorithm training and application.
Smart Images

Figure CN121095906B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to a method, apparatus, device and medium for determining bird's-eye view data frames. Background Technology
[0002] With the rapid development of autonomous driving technology, perception algorithms based on bird's-eye view (BEV) have gradually become the industry mainstream. The operation of BEV algorithms relies on the fusion of camera data and LiDAR data to transform the environmental information around the vehicle into a "sky-view" perspective, helping the autonomous driving system to understand road conditions more comprehensively and accurately, such as identifying obstacles and judging lane boundaries.
[0003] However, during the acquisition of camera and LiDAR data, the different acquisition frame rates of LiDAR and camera (for example, the acquisition frame rate of LiDAR is usually 10Hz, while the frame rate of camera is usually 30Hz) make it difficult to acquire camera and LiDAR data simultaneously, which seriously affects the perception accuracy of the BEV algorithm. Summary of the Invention
[0004] To address the aforementioned issues, this application provides a method, apparatus, device, and medium for determining bird's-eye view data frames, which can improve the perception accuracy of bird's-eye view algorithms.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] In a first aspect, this application discloses a method for determining a bird's-eye view data frame, the method comprising:
[0007] Acquire the first acquisition sequence for each of N lidars, the second acquisition sequence for each of M cameras, and the vehicle motion state acquisition sequence; wherein, the first acquisition sequence includes multiple frames of point cloud data and the average timestamp of each frame of point cloud data, the average timestamp being the average of the acquisition timestamps for each point; the second acquisition sequence includes multiple frames of image data and the acquisition timestamp of each frame of image data; the vehicle motion state acquisition sequence includes multiple vehicle motion state data and the acquisition timestamp of each vehicle motion state data.
[0008] The first acquisition sequences of each of the N lidars are fused to obtain a fused acquisition sequence; wherein, the fused acquisition sequence includes multiple frames of fused point cloud data and a fusion timestamp for each frame of fused point cloud data;
[0009] For each frame of fused point cloud data, the following steps are performed: In the second acquisition sequences of the M cameras, determine the M first timestamps closest to the fusion timestamp and the corresponding M frames of target image data; In the vehicle motion state acquisition sequence, determine the second timestamp closest to the fusion timestamp and the corresponding target vehicle motion state data; If the absolute value of the difference between the M first timestamps and the target timestamp is less than a first preset threshold, and the absolute value of the difference between the second timestamp and the target timestamp is less than a second preset threshold, then determine the acquisition timestamp of each point in the fused point cloud data; Wherein, the target timestamp is the average value of the M first timestamps; Based on the first difference between the acquisition timestamp of each point and the target timestamp, compensate for each point to obtain compensated fused point cloud data;
[0010] Based on the compensated fused point cloud data, the M-frame target image data, and the target vehicle motion state data, a bird's-eye view data frame is determined.
[0011] Optionally, fusing the first acquisition sequences of the N lidars to obtain a fused acquisition sequence includes:
[0012] The main radar is determined from the N lidars;
[0013] For each frame of point cloud data from the main radar, the following steps are performed: In the first acquisition sequence of each of the remaining N-1 lidars, determine the N-1 candidate timestamps closest to the average timestamp of the point cloud data, and the corresponding N-1 candidate point cloud data; if the difference between the N-1 candidate timestamps and the average timestamp is less than a third preset threshold, then determine the point cloud data and the N-1 candidate point cloud data as a candidate point cloud data group; convert the candidate point cloud data group to a unified vehicle coordinate system to obtain a converted point cloud data group; based on the effective data mask area, extract the filtered point cloud data group from the converted point cloud data group;
[0014] The filtered point cloud data groups corresponding to each frame of point cloud data from the main radar are fused to obtain a fused acquisition sequence.
[0015] Optionally, the step of extracting the filtered point cloud data group from the converted point cloud data group based on the effective data mask region includes:
[0016] Determine the effective masking area for each of the N lidars;
[0017] For each point cloud data in the transformed point cloud data group, the filtered point cloud data group is determined by extracting the points of each point cloud data within the effective mask area of the corresponding lidar.
[0018] Optionally, the effective masking area of each of the N lidars is related to the installation position, installation angle, and field of view of each of the N lidars.
[0019] Optionally, the formula for compensating each point based on the first difference between the acquisition timestamp and the target timestamp is as follows:
[0020] ;
[0021] ;
[0022] ;
[0023] ;
[0024] ;
[0025] in, Let k be the coordinates after compensation. Let k be the coordinates before compensation. R is the yaw rotation angle, and R is the turning radius. Let yaw rate be the vehicle's angular velocity. The first difference, For speed.
[0026] Secondly, this application discloses a device for determining bird's-eye view data frames, the device comprising: a sequence acquisition module, a sequence fusion module, a data compensation module, and a data determination module;
[0027] The sequence acquisition module is used to acquire the first acquisition sequence of each of N lidars, the second acquisition sequence of each of M cameras, and the vehicle motion state acquisition sequence; wherein, the first acquisition sequence includes multiple frames of point cloud data and the average timestamp of each frame of point cloud data, the average timestamp being the average of the acquisition timestamps of each point; the second acquisition sequence includes multiple frames of image data and the acquisition timestamp of each frame of image data; the vehicle motion state acquisition sequence includes multiple vehicle motion state data and the acquisition timestamp of each vehicle motion state data.
[0028] The sequence fusion module is used to fuse the first acquisition sequences of each of the N lidars to obtain a fused acquisition sequence; wherein, the fused acquisition sequence includes multiple frames of fused point cloud data and a fusion timestamp for each frame of fused point cloud data;
[0029] The data compensation module is used to perform the following operations on each frame of fused point cloud data: In the second acquisition sequences of the M cameras, determine the M first timestamps closest to the fusion timestamp and the corresponding M frames of target image data; in the vehicle motion state acquisition sequence, determine the second timestamp closest to the fusion timestamp and the corresponding target vehicle motion state data; if the absolute value of the difference between the M first timestamps and the target timestamp is less than a first preset threshold, and the absolute value of the difference between the second timestamp and the target timestamp is less than a second preset threshold, then determine the acquisition timestamp of each point in the fused point cloud data; wherein, the target timestamp is the average of the M first timestamps; and compensate for each point based on the first difference between the acquisition timestamp of each point and the target timestamp to obtain compensated fused point cloud data.
[0030] The data determination module is used to determine the bird's-eye view data frame based on the compensated fused point cloud data, the M-frame target image data, and the target vehicle motion state data.
[0031] Optionally, the sequence fusion module includes: a first fusion module, a second fusion module, and a third fusion module;
[0032] The first fusion module is used to determine the main radar from the N lidars;
[0033] The second fusion module is used to perform the following steps on each frame of point cloud data from the main radar: In the first acquisition sequences of the remaining N-1 lidars, determine the N-1 candidate timestamps closest to the average timestamp of the point cloud data, and the corresponding N-1 candidate point cloud data; if the difference between the N-1 candidate timestamps and the average timestamp is less than a third preset threshold, then determine the point cloud data and the N-1 candidate point cloud data as a candidate point cloud data group; convert the candidate point cloud data group to a unified vehicle coordinate system to obtain a converted point cloud data group; and extract a filtered point cloud data group from the converted point cloud data group based on the effective data mask area.
[0034] The third fusion module is used to fuse the filtered point cloud data groups corresponding to each frame of point cloud data of the main radar to obtain a fused acquisition sequence.
[0035] Optionally, the second fusion module is specifically used to: determine the effective mask area of each of the N lidars; and for each point cloud data in the converted point cloud data group, determine the filtered point cloud data group by extracting the points of each point cloud data in the effective mask area of the corresponding lidar.
[0036] Optionally, the effective masking area of each of the N lidars is related to the installation position, installation angle, and field of view of each of the N lidars.
[0037] Optionally, the formula for compensating each point based on the first difference between the acquisition timestamp and the target timestamp is as follows:
[0038] ;
[0039] ;
[0040] ;
[0041] ;
[0042] ;
[0043] in, Let k be the coordinates after compensation. Let k be the coordinates before compensation. R is the yaw rotation angle, and R is the turning radius. Let yaw rate be the vehicle's angular velocity. The first difference, For speed.
[0044] Thirdly, this application discloses a device for determining bird's-eye view data frames, the device comprising: a memory and a processor;
[0045] The memory is used to store programs;
[0046] The processor is configured to execute the program to implement the steps of the method for determining the bird's-eye view data frame as described in the first aspect.
[0047] Fourthly, this application discloses a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for determining the bird's-eye view data frame as described in the first aspect.
[0048] Compared with the prior art, this application has the following beneficial effects:
[0049] This application provides a method, apparatus, device, and medium for determining bird's-eye view data frames. The method includes: acquiring first acquisition sequences of N lidars, second acquisition sequences of M cameras, and vehicle motion state acquisition sequences; wherein the first acquisition sequences include multiple frames of point cloud data and an average timestamp for each frame of point cloud data, the average timestamp being the average of the acquisition timestamps for each point; the second acquisition sequences include multiple frames of image data and an acquisition timestamp for each frame of image data; the vehicle motion state acquisition sequences include multiple vehicle motion state data and an acquisition timestamp for each vehicle motion state data; fusing the first acquisition sequences of the N lidars to obtain a fused acquisition sequence; wherein the fused acquisition sequence includes multiple frames of fused point cloud data and a fusion timestamp for each frame of fused point cloud data; and for each frame of fused point cloud data, performing the following: in M... In the second acquisition sequence of each camera, M first timestamps closest to the fusion timestamp are determined, along with M frames of target image data. In the vehicle motion state acquisition sequence, the second timestamp closest to the fusion timestamp is determined, along with the corresponding target vehicle motion state data. If the absolute value of the difference between the M first timestamps and the target timestamp is less than a first preset threshold, and the absolute value of the difference between the second timestamp and the target timestamp is less than a second preset threshold, then the acquisition timestamp of each point in the fused point cloud data is determined. The target timestamp is the average of the M first timestamps. Based on the first difference between the acquisition timestamp and the target timestamp of each point, each point is compensated to obtain compensated fused point cloud data. Based on the compensated fused point cloud data, the M frames of target image data, and the target vehicle motion state data, the bird's-eye view data frame is determined. Therefore, the method for determining bird's-eye view data frames provided in this application effectively solves the problem of frame rate differences between LiDAR and camera, thereby improving the time alignment accuracy of multi-source data, providing high-quality and highly consistent bird's-eye view data frames to support subsequent autonomous driving bird's-eye view algorithms, and ensuring the accuracy and reliability of bird's-eye view algorithm training and application. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 A flowchart illustrating a method for determining a bird's-eye view data frame, as provided in an embodiment of this application;
[0052] Figure 2 A flowchart illustrating a method for acquiring a fused acquisition sequence as provided in an embodiment of this application;
[0053] Figure 3 A flowchart of a coordinate transformation provided in an embodiment of this application;
[0054] Figure 4 A flowchart illustrating data filtering provided in this application embodiment;
[0055] Figure 5 A flowchart illustrating data compensation provided in an embodiment of this application;
[0056] Figure 6 A schematic diagram of a device for determining a bird's-eye view data frame provided in an embodiment of this application;
[0057] Figure 7 This is a schematic diagram of a computer-readable medium provided in an embodiment of this application. Detailed Implementation
[0058] As described earlier, during the acquisition of camera and LiDAR data, the different frame rates of LiDAR and camera (e.g., LiDAR typically uses a frame rate of 10Hz, while camera typically uses 30Hz) make it difficult to acquire corresponding LiDAR and camera data simultaneously. Furthermore, due to fluctuations in system load (such as changes in computing resource allocation) and hardware stability, the acquisition interval between two frames of camera and LiDAR data is not entirely stable, and some data frames may even be lost. This further disrupts the temporal correlation between camera and LiDAR data, severely impacting the perception accuracy of the BEV algorithm.
[0059] Through research, the inventors proposed a method, apparatus, device, and medium for determining bird's-eye view data frames. This method effectively solves the problem of frame rate differences between LiDAR and cameras, thereby improving the time alignment accuracy of multi-source data. It provides high-quality and highly consistent bird's-eye view data frames to support subsequent autonomous driving bird's-eye view algorithms, ensuring the accuracy and reliability of bird's-eye view algorithm training and application.
[0060] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0061] See Figure 1 This figure is a flowchart of a method for determining a bird's-eye view data frame according to an embodiment of this application. The method includes:
[0062] S101: Acquire the first acquisition sequence of each of the N lidars, the second acquisition sequence of each of the M cameras, and the vehicle motion state acquisition sequence; wherein, the first acquisition sequence includes multiple frames of point cloud data and the average timestamp of each frame of point cloud data, the average timestamp being the average of the acquisition timestamps of each point; the second acquisition sequence includes multiple frames of image data and the acquisition timestamp of each frame of image data; the vehicle motion state acquisition sequence includes multiple vehicle motion state data and the acquisition timestamp of each vehicle motion state data.
[0063] First, obtain the first acquisition sequence of each of the N (N is a positive integer) lidars. The first acquisition sequence includes multiple frames of point cloud data and the average timestamp of each frame. Since each frame of point cloud data includes points acquired at multiple different times, the average timestamp is the acquisition timestamp of all points in that frame. The average value.
[0064] Secondly, obtain the second acquisition sequence for each of the M cameras (M is a positive integer). The second acquisition sequence includes multiple frames of image data and the acquisition timestamp of each frame. .
[0065] Subsequently, the vehicle motion state acquisition sequence was obtained. The vehicle motion state acquisition sequence includes multiple frames of vehicle motion state data and the acquisition timestamp of each frame of vehicle motion state data. Each frame of vehicle motion data includes vehicle speed. and yaw rate .
[0066] S102: Fuse the first acquisition sequences of each of the N lidars to obtain a fused acquisition sequence; wherein, the fused acquisition sequence includes multiple frames of fused point cloud data and a fused timestamp for each frame of fused point cloud data.
[0067] See Figure 2 This figure is a flowchart of a method for obtaining a fused acquisition sequence according to an embodiment of this application. The fused acquisition sequence is obtained through the following steps A1-A3:
[0068] A1: Determine the main radar from N lidars.
[0069] The N lidar units include forward-facing lidars, side-facing lidars, and rear-facing lidars. In the method for determining the bird's-eye view data frame provided in this application embodiment, the forward-facing lidar is typically selected as the primary lidar. This is because autonomous driving decisions focus more on the road conditions ahead.
[0070] A2: For each frame of point cloud data from the main radar, perform the following steps A21-A24:
[0071] A21: In the first acquisition sequence of the remaining N-1 lidars, determine the N-1 candidate timestamps that are closest to the average timestamp of the point cloud data in that frame, and the corresponding N-1 candidate point cloud data.
[0072] The autonomous driving system is equipped with N LiDARs. Since the scanning frequencies of different LiDARs may be different, and the scanning frequency may fluctuate due to system load, it is necessary to perform preliminary temporal alignment of the point cloud data of the N LiDARs first.
[0073] First, select the main radar. A frame of point cloud data And determine the point cloud data of that frame. Average timestamp This serves as the baseline timestamp. The average timestamp refers to the average of the acquisition timestamps of all points in that frame of point cloud data.
[0074] Subsequently, for the remaining N-1 lidars In their respective first acquisition sequences, the average timestamp of the point cloud data for that frame is determined. The closest N-1 candidate timestamps , , and the corresponding N-1 candidate point cloud data.
[0075] A22: If the difference between the N-1 candidate timestamps and the average timestamp is less than the third preset threshold, then the point cloud data of that frame and the N-1 candidate point cloud data are determined as a candidate point cloud data group.
[0076] Determine N-1 candidate timestamps The average timestamp of this frame of point cloud data from the main radar Whether the difference between them is less than a third preset threshold. The third preset threshold can be 50ms, but this application does not limit it.
[0077] If yes, then the point cloud data of that frame from the main radar and N-1 candidate point cloud data are identified as a candidate point cloud data group (this group of point cloud data has strong temporal correlation). If not, then the point cloud data of that frame from the main radar is discarded (the time deviation is too large).
[0078] After determining the candidate point cloud data group corresponding to the point cloud data of the main radar frame, continue to perform the above A21 step on the other point cloud data of the main radar until the steps of retaining (i.e. determining the candidate point cloud data group) or discarding are completed for all the point cloud data of the main radar frames, and then perform step A23.
[0079] A23: Transform the candidate point cloud data set to a unified vehicle coordinate system to obtain the transformed point cloud data set.
[0080] Since the autonomous driving system is equipped with multiple LiDARs, and each LiDAR has a different installation position and angle (for example, the forward LiDAR is installed at the front of the vehicle, and the side LiDAR is installed on both sides of the vehicle), the coordinate systems of the point cloud data collected by different LiDARs are different. Therefore, it is necessary to convert the candidate point cloud data group to a unified vehicle coordinate system.
[0081] See Figure 3 This figure is a flowchart of a coordinate transformation provided in an embodiment of this application. Taking multiple lidars, including a first lidar, as an example, the following steps are performed for coordinate transformation from B1 to B2:
[0082] B1: Obtain the external parameter matrix of the first lidar. .
[0083] The extrinsic parameter matrix refers to the transformation relationship between the local coordinate system of the first lidar and the unified vehicle coordinate system. For example, the extrinsic parameter matrix can be a 4×4 transformation matrix that includes translation and rotation information.
[0084] B2: In the converted point cloud data set, the T-th point of the first lidar... i Coordinates of each point in the frame point cloud data Using the extrinsic matrix of the first lidar Perform coordinate transformation to obtain the coordinates of each transformed point. And based on the coordinates of each transformed point, synthesize the transformed T-th point. i Frame point cloud data .
[0085] For example, using the extrinsic parameter matrix of the first lidar The formula for coordinate transformation is shown in formula (1) below:
[0086] = (1)
[0087] in, The coordinates of the point before the transformation. These are the transformed point coordinates. This is the extrinsic parameter matrix of the first lidar.
[0088] Subsequently, the above step B2 is performed on the other frames of point cloud data from the first lidar in the transformed point cloud data group until the coordinate transformation of all frames of point cloud data from the first lidar is completed. Then, the above steps B1-B2 are performed on the other lidars corresponding to the transformed point cloud data group until the coordinate transformation of all frames of point cloud data from the other lidars is completed, thus obtaining the transformed point cloud data group.
[0089] A24: Extract the filtered point cloud data group from the transformed point cloud data group based on the effective data mask area.
[0090] When multiple LiDARs are installed on a vehicle, there are often invalid areas due to the vehicle's structure and obstructions. These invalid areas not only increase the computational load but may also interfere with the perception results. Therefore, it is necessary to filter out the invalid areas.
[0091] See Figure 4 This figure is a flowchart of a data filtering method provided in an embodiment of this application. Taking multiple lidars, including a first lidar, as an example, the following data filtering steps (C1-C5) are performed:
[0092] C1: Determine the effective masking area of the first lidar.
[0093] The effective masking area of the first lidar , It typically consists of multiple polygonal regions on the ground plane and is related to the installation location, installation angle, and field of view (FOV) of the first lidar.
[0094] It should be noted that the effective masking areas of adjacent lidars can overlap by 0.5m-1m, thereby avoiding missed detections and reducing redundant overlap.
[0095] C2: Determine the T-th transformed point cloud data set from the first LiDAR. i Frame point cloud data The coordinates of each transformed point in the data Is it within a valid mask area? If yes, proceed to step C3; otherwise, proceed to step C4.
[0096] C3: Retain the coordinates of the transformed point. The point was then identified as valid. Step C5 was then executed.
[0097] C4: Delete the transformed point coordinates And it was determined to be an invalid point.
[0098] C5: By extracting the Tth transformed value from the first lidar iFrame point cloud data For each valid point in the dataset, generate the Tth point. i Point cloud data after frame filtering .
[0099] Subsequently, the above-described step C2 is performed on the other frames of point cloud data from the first LiDAR within the converted point cloud data group, until the filtering step is completed for all frames of point cloud data from the first LiDAR within the converted point cloud data group. Then, the above-described steps C1-C5 are performed on the other LiDARs corresponding to the converted point cloud data group, until the data filtering step is completed for all frames of point cloud data from the other LiDARs, resulting in a filtered point cloud data group.
[0100] It is understandable that the point cloud data in the filtered point cloud data group above is point cloud data that is initially aligned in time, with coordinates in the vehicle coordinate system, and only retains the effective area of the mask mark, as shown in the following formula (2):
[0101] (2)
[0102] A3: The filtered point cloud data groups corresponding to each frame of point cloud data from the main radar are fused to obtain a fused acquisition sequence.
[0103] The fused acquisition sequence can be represented by the following formula (3):
[0104] (3)
[0105] The fused acquisition sequence includes multi-frame fused point cloud data. Fusion timestamps of each frame of fused point cloud data Wherein, the fusion timestamp is the average of the collection timestamps of each point in the fusion point cloud data.
[0106] S103: For each frame of fused point cloud data, perform the following: In the second acquisition sequence of each of the M cameras, determine the M first timestamps closest to the fusion timestamp and the corresponding M frames of target image data; In the vehicle motion state acquisition sequence, determine the second timestamp closest to the fusion timestamp and the corresponding target vehicle motion state data; If the absolute value of the difference between the M first timestamps and the target timestamp is less than a first preset threshold, and the absolute value of the difference between the second timestamp and the target timestamp is less than a second preset threshold, then determine the acquisition timestamp of each point in the fused point cloud data; Wherein, the target timestamp is the average value of the M first timestamps; Based on the first difference between the acquisition timestamp and the target timestamp of each point, compensate for each point to obtain compensated fused point cloud data.
[0107] See Figure 5This figure is a flowchart of a data compensation method provided in an embodiment of this application. The specific data compensation process is as follows:
[0108] D1: In the second acquisition sequence of each of the M cameras, determine the fusion timestamp of the point cloud data to be fused with that frame. The M closest first timestamps and the corresponding M-frame target image data .
[0109] D2: In the vehicle motion state acquisition sequence, determine the fusion timestamp of the point cloud data fused with that frame. The closest second timestamp and the corresponding target vehicle motion state data .
[0110] D3: Determine the target timestamp The target timestamp is the average of M first timestamps.
[0111] D4: Determine if the following conditions are met: the absolute values of the differences between the M first timestamps and the target timestamp are all less than the first preset threshold, and the absolute values of the differences between the second timestamps and the target timestamp are less than the second preset threshold. If yes, proceed to step D5; otherwise, proceed to step D6.
[0112] For example, the first preset threshold can be 15ms (where higher camera time alignment accuracy is required because image semantics are more sensitive to time). The second preset threshold can be 50ms (where motion states have slightly lower time accuracy requirements, but must still be controlled within the threshold). This application does not limit the specific first or second preset threshold.
[0113] D5: Based on the fused point cloud data of this frame, the target image data of frame M, and the target vehicle motion state data, determine the original bird's-eye view data frame. Then, execute step D7.
[0114] If the absolute values of the differences between the M first timestamps and the target timestamp are all less than a first preset threshold, and the absolute values of the differences between the second timestamps and the target timestamps are less than a second preset threshold, then it is determined that the fused point cloud data, M frames of target image data, and target vehicle motion state data satisfy time alignment, and the aligned timestamps are... Subsequently, the original bird's-eye view data frame is recorded as shown in formula (4):
[0115] (4)
[0116] D6: Discard the fused point cloud data of this frame.
[0117] If the absolute value of the difference between the first timestamp and the target timestamp is greater than or equal to the first preset threshold, or if the absolute value of the difference between the second timestamp and the target timestamp is greater than or equal to the second preset threshold, then it is determined that the above-mentioned fused point cloud data, M-frame target image data, and target vehicle motion state data do not meet the time alignment requirement, and the fused point cloud data of that frame needs to be discarded.
[0118] D7: Based on the first difference between the acquisition timestamp and the target timestamp of each point in the frame of fused point cloud data, compensate for each point in the frame of fused point cloud data to obtain compensated fused point cloud data.
[0119] The formula for compensating each point (taking point k as an example) of the fused point cloud data of this frame can be shown in the following formulas (5)-(9):
[0120] (5)
[0121] (6)
[0122] (7)
[0123] (8)
[0124] (9)
[0125] in, The coordinates of point k in the fused point cloud data of this frame are compensated. The coordinates of point k in the fused point cloud data of this frame before compensation. R is the yaw rotation angle, and R is the turning radius. Let yaw rate be the vehicle's angular velocity. The first difference, For speed.
[0126] Subsequently, the above D1 step is performed on other fused point cloud data in the fused acquisition sequence until all compensated fused point cloud data corresponding to all fused point cloud data are obtained.
[0127] S104: Determine the bird's-eye view data frame based on the compensated fused point cloud data, M-frame target image data, and target vehicle motion state data.
[0128] Based on the compensated fused point cloud data, M-frame target image data, and target vehicle motion state data, the bird's-eye view data frame is determined. The bird's-eye view data frame can be represented by the following formula (10). Multiple bird's-eye view data frames can form a bird's-eye view data sequence, which can be represented by the following formula (11):
[0129]
[0130] (10)
[0131] (11)
[0132] In summary, the embodiments of this application provide a method for determining bird's-eye view data frames. This method effectively solves the problem of frame rate differences between LiDAR and cameras, thereby improving the temporal alignment accuracy of multi-source data. It provides high-quality and highly consistent bird's-eye view data frames to support subsequent autonomous driving bird's-eye view algorithms, ensuring the accuracy and reliability of bird's-eye view algorithm training and application.
[0133] See Figure 6 This figure is a schematic diagram of a bird's-eye view data frame determination device provided in an embodiment of this application. The bird's-eye view data frame determination device 600 includes: a sequence acquisition module 601, a sequence fusion module 602, a data compensation module 603, and a data determination module 604;
[0134] The sequence acquisition module 601 is used to acquire the first acquisition sequence of each of N lidars, the second acquisition sequence of each of M cameras, and the vehicle motion state acquisition sequence. The first acquisition sequence includes multiple frames of point cloud data and the average timestamp of each frame, where the average timestamp is the average of the acquisition timestamps for each point. The second acquisition sequence includes multiple frames of image data and the acquisition timestamp of each frame. The vehicle motion state acquisition sequence includes multiple vehicle motion state data and the acquisition timestamp of each vehicle motion state data.
[0135] The sequence fusion module 602 is used to fuse the first acquisition sequences of N lidars to obtain a fused acquisition sequence; wherein, the fused acquisition sequence includes multiple frames of fused point cloud data and a fusion timestamp for each frame of fused point cloud data;
[0136] The data compensation module 603 is used to perform the following operations on each frame of fused point cloud data: In the second acquisition sequence of each of the M cameras, determine the M first timestamps closest to the fusion timestamp and the corresponding M frames of target image data; In the vehicle motion state acquisition sequence, determine the second timestamp closest to the fusion timestamp and the corresponding target vehicle motion state data; If the absolute value of the difference between the M first timestamps and the target timestamp is less than a first preset threshold, and the absolute value of the difference between the second timestamp and the target timestamp is less than a second preset threshold, then determine the acquisition timestamp of each point in the fused point cloud data; Wherein, the target timestamp is the average value of the M first timestamps; Based on the first difference between the acquisition timestamp and the target timestamp of each point, compensate for each point to obtain the compensated fused point cloud data;
[0137] The data determination module 604 is used to determine the bird's-eye view data frame based on the compensated fused point cloud data, the M-frame target image data, and the target vehicle motion state data.
[0138] Optionally, the sequence fusion module 602 includes: a first fusion module, a second fusion module, and a third fusion module;
[0139] The first fusion module is used to determine the main radar from N lidars;
[0140] The second fusion module is used to perform the following steps on each frame of point cloud data from the main radar: In the first acquisition sequence of the remaining N-1 lidars, determine the N-1 candidate timestamps that are closest to the average timestamp of the point cloud data, and the corresponding N-1 candidate point cloud data; If the difference between the N-1 candidate timestamps and the average timestamp is less than a third preset threshold, then determine the point cloud data and the N-1 candidate point cloud data as a candidate point cloud data group; Transform the candidate point cloud data group to a unified vehicle coordinate system to obtain the transformed point cloud data group; Based on the effective data mask area, extract the filtered point cloud data group from the transformed point cloud data group;
[0141] The third fusion module is used to fuse the filtered point cloud data groups corresponding to each frame of point cloud data from the main radar to obtain a fused acquisition sequence.
[0142] Optionally, the second fusion module is specifically used to: determine the effective mask area of each of the N lidars; and for each point cloud data in the converted point cloud data group, determine the filtered point cloud data group by extracting the points of each point cloud data in the effective mask area of the corresponding lidar.
[0143] Optionally, the effective masking area of each of the N lidars is related to the installation position, installation angle, and field of view of each of the N lidars.
[0144] Optionally, the formula for compensating each point based on the first difference between the collection timestamp and the target timestamp is as follows:
[0145] ;
[0146] ;
[0147] ;
[0148] ;
[0149] ;
[0150] in, Let k be the coordinates after compensation. Let k be the coordinates before compensation. R is the yaw rotation angle, and R is the turning radius. Let yaw rate be the vehicle's angular velocity. The first difference, For speed.
[0151] In summary, the embodiments of this application provide a device for determining bird's-eye view data frames. This device effectively solves the problem of frame rate differences between LiDAR and cameras, thereby improving the temporal alignment accuracy of multi-source data. It provides high-quality and highly consistent bird's-eye view data frames to support subsequent autonomous driving bird's-eye view algorithms, ensuring the accuracy and reliability of bird's-eye view algorithm training and application.
[0152] This application also provides a device for determining the corresponding bird's-eye view data frame and a computer-readable medium for implementing the method for determining the bird's-eye view data frame provided in this application.
[0153] The device for determining the bird's-eye view data frame includes a memory and a processor. The memory is used to store instructions or code, and the processor is used to execute the instructions or code to enable the device to perform a method for determining a bird's-eye view data frame according to any embodiment of this application.
[0154] See Figure 7 This figure is a schematic diagram of a computer-readable medium provided in an embodiment of this application. The computer-readable medium 700 stores a computer program 711, which, when executed by a processor, implements the above-described... Figure 1 The steps for determining the bird's-eye view data frame.
[0155] It should be noted that, in the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0156] It should be noted that the machine-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0157] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0158] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
[0159] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0160] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A method for determining a bird's-eye view data frame, characterized in that, The method includes: Acquire the first acquisition sequence for each of N lidars, the second acquisition sequence for each of M cameras, and the vehicle motion state acquisition sequence; wherein, the first acquisition sequence includes multiple frames of point cloud data and the average timestamp of each frame of point cloud data, the average timestamp being the average of the acquisition timestamps for each point; the second acquisition sequence includes multiple frames of image data and the acquisition timestamp of each frame of image data; the vehicle motion state acquisition sequence includes multiple vehicle motion state data and the acquisition timestamp of each vehicle motion state data. The first acquisition sequences of each of the N lidars are fused to obtain a fused acquisition sequence; wherein, the fused acquisition sequence includes multiple frames of fused point cloud data and a fusion timestamp for each frame of fused point cloud data; For each frame of fused point cloud data, the following steps are performed: In the second acquisition sequences of the M cameras, determine the M first timestamps closest to the fusion timestamp and the corresponding M frames of target image data; In the vehicle motion state acquisition sequence, determine the second timestamp closest to the fusion timestamp and the corresponding target vehicle motion state data; If the absolute value of the difference between the M first timestamps and the target timestamp is less than a first preset threshold, and the absolute value of the difference between the second timestamp and the target timestamp is less than a second preset threshold, then determine the acquisition timestamp of each point in the fused point cloud data; Wherein, the target timestamp is the average value of the M first timestamps; Based on the first difference between the acquisition timestamp of each point and the target timestamp, compensate for each point to obtain compensated fused point cloud data; Based on the compensated fused point cloud data, the M-frame target image data, and the target vehicle motion state data, a bird's-eye view data frame is determined.
2. The method according to claim 1, characterized in that, The process of fusing the first acquisition sequences of the N lidars to obtain a fused acquisition sequence includes: The main radar is determined from the N lidars; For each frame of point cloud data from the main radar, the following steps are performed: In the first acquisition sequence of each of the remaining N-1 lidars, determine the N-1 candidate timestamps closest to the average timestamp of the point cloud data, and the corresponding N-1 candidate point cloud data; if the difference between the N-1 candidate timestamps and the average timestamp is less than a third preset threshold, then determine the point cloud data and the N-1 candidate point cloud data as a candidate point cloud data group; convert the candidate point cloud data group to a unified vehicle coordinate system to obtain a converted point cloud data group; based on the effective data mask area, extract the filtered point cloud data group from the converted point cloud data group; The filtered point cloud data groups corresponding to each frame of point cloud data from the main radar are fused to obtain a fused acquisition sequence.
3. The method according to claim 2, characterized in that, The step of extracting the filtered point cloud data group from the transformed point cloud data group based on the effective data mask region includes: Determine the effective masking area for each of the N lidars; For each point cloud data in the transformed point cloud data group, the filtered point cloud data group is determined by extracting the points of each point cloud data within the effective mask area of the corresponding lidar.
4. The method according to claim 3, characterized in that, The effective masking area of each of the N lidars is related to the installation position, installation angle, and field of view of each of the N lidars.
5. The method according to claim 1, characterized in that, The formula for compensating each point based on the first difference between the collection timestamp and the target timestamp is as follows: ; ; ; ; ; in, Let k be the coordinates after compensation. Let k be the coordinates before compensation. R is the yaw rotation angle, and R is the turning radius. Let yaw rate be the vehicle's angular velocity. The first difference, For speed.
6. A device for determining a bird's-eye view data frame, characterized in that, The device includes: a sequence acquisition module, a sequence fusion module, a data compensation module, and a data determination module; The sequence acquisition module is used to acquire the first acquisition sequence of each of N lidars, the second acquisition sequence of each of M cameras, and the vehicle motion state acquisition sequence; wherein, the first acquisition sequence includes multiple frames of point cloud data and the average timestamp of each frame of point cloud data, the average timestamp being the average of the acquisition timestamps of each point; the second acquisition sequence includes multiple frames of image data and the acquisition timestamp of each frame of image data; the vehicle motion state acquisition sequence includes multiple vehicle motion state data and the acquisition timestamp of each vehicle motion state data. The sequence fusion module is used to fuse the first acquisition sequences of each of the N lidars to obtain a fused acquisition sequence; wherein, the fused acquisition sequence includes multiple frames of fused point cloud data and a fusion timestamp for each frame of fused point cloud data; The data compensation module is used to perform the following operations on each frame of fused point cloud data: In the second acquisition sequences of the M cameras, determine the M first timestamps closest to the fusion timestamp and the corresponding M frames of target image data; in the vehicle motion state acquisition sequence, determine the second timestamp closest to the fusion timestamp and the corresponding target vehicle motion state data; if the absolute value of the difference between the M first timestamps and the target timestamp is less than a first preset threshold, and the absolute value of the difference between the second timestamp and the target timestamp is less than a second preset threshold, then determine the acquisition timestamp of each point in the fused point cloud data; wherein, the target timestamp is the average of the M first timestamps; and compensate for each point based on the first difference between the acquisition timestamp of each point and the target timestamp to obtain compensated fused point cloud data. The data determination module is used to determine the bird's-eye view data frame based on the compensated fused point cloud data, the M-frame target image data, and the target vehicle motion state data.
7. The apparatus according to claim 6, characterized in that, The sequence fusion module includes: a first fusion module, a second fusion module, and a third fusion module; The first fusion module is used to determine the main radar from the N lidars; The second fusion module is used to perform the following steps on each frame of point cloud data from the main radar: In the first acquisition sequences of the remaining N-1 lidars, determine the N-1 candidate timestamps closest to the average timestamp of the point cloud data, and the corresponding N-1 candidate point cloud data; if the difference between the N-1 candidate timestamps and the average timestamp is less than a third preset threshold, then determine the point cloud data and the N-1 candidate point cloud data as a candidate point cloud data group; convert the candidate point cloud data group to a unified vehicle coordinate system to obtain a converted point cloud data group; and extract a filtered point cloud data group from the converted point cloud data group based on the effective data mask area. The third fusion module is used to fuse the filtered point cloud data groups corresponding to each frame of point cloud data of the main radar to obtain a fused acquisition sequence.
8. The apparatus according to claim 7, characterized in that, The second fusion module is specifically used to: determine the effective mask area of each of the N lidars; and for each point cloud data in the converted point cloud data group, determine the filtered point cloud data group by extracting the points of each point cloud data in the effective mask area of the corresponding lidar.
9. A device for determining a bird's-eye view data frame, characterized in that, The device includes: a memory and a processor; The memory is used to store programs; The processor is configured to execute the program to implement the steps of the method for determining the bird's-eye view data frame as described in any one of claims 1 to 5.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for determining the bird's-eye view data frame as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Vehicle state sensing method based on thunder-vision fusion
CN117111055A
Vehicle starting behavior prediction method and device, storage medium and program product
CN117115776A