A three-dimensional target detection method and device with multi-sensor time adaptive synchronization
The problem of sensor timestamp misalignment is solved through the method of multi-sensor time adaptive synchronization, and the accuracy and robustness of 3D target detection are improved through feature compensation and fusion processing.
Patent Information
- Application Number
- CN202510948578.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-10
AI Technical Summary
In the existing technology, the timestamp misalignment problem of multimodal sensors causes data distortion, affecting the robust performance of three-dimensional object detection.
By acquiring the RGB image frame and 3D point cloud data frame at the current moment, and using the motion feature map and timestamp to determine the alignment, corresponding feature compensation and fusion processing are performed, including the calculation of position offset maps under camera and radar freezes, to achieve multi-sensor time adaptive synchronization.
The robustness of multimodal data fusion is improved, and the accuracy of three-dimensional target detection is enhanced.
Smart Images

Figure CN120451726B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a three-dimensional target detection method and device with multi-sensor time adaptive synchronization. Background Art
[0002] Multimodal data fusion is an important research direction in 3D object detection. Data from a single sensor has limitations in describing object form and spatial position, making accurate detection difficult. To address these issues, researchers are integrating and analyzing data from different sensors to provide more comprehensive and accurate information, thereby improving 3D object detection.
[0003] In the multimodal fusion process, a common method for synchronizing data from different modalities is timestamp alignment. However, in reality, timestamps from different modalities can deviate. For example, network delays or laser failures can cause radar to freeze; network delays or camera failures can cause camera freezes. Currently, a method called "time lag compensation" can be used to address timestamp misalignment. However, this method, based on historical data, has low physical interpretability. When the time gap is large, it can cause data distortion. This distortion can significantly impact subsequent data processing and analysis, leading to a decrease in the model's robustness. Summary of the Invention
[0004] In view of this, the present application provides a three-dimensional target detection method and device with multi-sensor time adaptive synchronization to solve the above technical problems.
[0005] In a first aspect, an embodiment of the present application provides a three-dimensional target detection method with multi-sensor time adaptive synchronization, comprising:
[0006] Obtain the RGB image frame and 3D point cloud data frame of the target area at the current moment;
[0007] Process the continuous 3D point cloud data frame sequence before the current moment to obtain the motion feature map of the previous moment;
[0008] Using the timestamp of the current RGB image frame and the timestamp of the current 3D point cloud data frame, determine whether the current RGB image frame and the current 3D point cloud data frame are aligned. If so, fuse the current RGB image frame and the current 3D point cloud data frame to obtain the fused BEV feature at the current moment.
[0009] Otherwise, when the misalignment is caused by camera jamming, the position offset map of the pixels under the BEV perspective of the laser radar is calculated using the motion feature map of the previous moment, the image BEV features of the current moment are calculated based on the image BEV features and the position offset map of the previous moment, and the fused BEV features of the current moment are calculated based on the point cloud BEV features of the three-dimensional point cloud data frame at the current moment and the image BEV features of the current moment; when the misalignment is caused by radar jamming, the position offset map of the pixels under the BEV perspective of the laser radar is calculated using the motion feature map of the previous moment, the point cloud BEV features of the current moment are calculated based on the point cloud BEV features and the position offset map of the previous moment, and the fused BEV features of the current moment are calculated based on the point cloud BEV features of the current moment and the image BEV features of the RGB image at the current moment;
[0010] The final fused BEV feature is calculated using the fused BEV features of the two consecutive moments before the current moment and the corresponding motion feature map, and the fused BEV feature of the current moment;
[0011] The final fused BEV features are processed by the detection head to obtain the three-dimensional target detection results.
[0012] In one possible implementation, a continuous sequence of three-dimensional point cloud data frames before a current moment is processed to obtain a motion feature map at the previous moment; including:
[0013] Perform feature extraction on the continuous 3D point cloud data frame sequence before the current moment to obtain the point cloud BEV feature sequence;
[0014] The pre-trained spatiotemporal pyramid network is used to process the point cloud BEV feature sequence to obtain the motion feature map of the previous moment. ,in, Pixels The eigenvalues of Pixels Predicted speeds in both directions; For the current moment, The time interval between the current moment and the previous moment;
[0015] The spatiotemporal pyramid network consists of multiple spatiotemporal convolutional layers and global temporal pooling layers, and each spatiotemporal convolutional layer consists of standard two-dimensional convolution and three-dimensional convolution.
[0016] In one possible implementation, using the timestamp of the current RGB image frame and the timestamp of the current 3D point cloud data frame, determining whether the current RGB image frame and the current 3D point cloud data frame are aligned includes:
[0017] Calculate the absolute value of the difference between the timestamp of the RGB image frame at the current moment and the timestamp of the 3D point cloud data frame at the current moment;
[0018] Determine whether the absolute value of the difference is greater than a preset threshold. If so, determine that the RGB image frame at the current moment and the 3D point cloud data frame at the current moment are not aligned; otherwise, determine that the RGB image frame at the current moment and the 3D point cloud data frame at the current moment are aligned.
[0019] In one possible implementation, the RGB image frame at the current moment and the 3D point cloud data frame at the current moment are fused to obtain the fused BEV feature at the current moment; including:
[0020] Use the point cloud feature extractor to process the three-dimensional point cloud data frame at the current moment to obtain the point cloud BEV feature at the current moment ; For the current moment;
[0021] Use the image feature extractor to process the RGB image frame at the current moment to obtain the image BEV feature at the current moment ;
[0022] Calculate the fused BEV features at the current moment :
[0023] ;
[0024] in, Represents the addition operation.
[0025] In one possible implementation, when the misalignment is caused by camera jamming, a position offset map of pixels under the BEV perspective of the lidar is calculated using a motion feature map at a previous moment, the image BEV features at a current moment are calculated based on the image BEV features and the position offset map at a previous moment, and the fused BEV features at a current moment are calculated based on the point cloud BEV features of the three-dimensional point cloud data frame at a current moment and the image BEV features at a current moment; including:
[0026] Use the point cloud feature extractor to process the three-dimensional point cloud data at the current moment to obtain the point cloud BEV feature at the current moment ;
[0027] Using the motion feature map of the previous moment , calculate from the previous moment To the current moment Position offset map within :
[0028] ;
[0029] Calculate the current time Image BEV features :
[0030] ;
[0031] in, Indicates the last moment Image BEV features;
[0032] Fusion BEV characteristics at the current moment for:
[0033] ;
[0034] in, Represents the addition operation.
[0035] In one possible implementation, when the misalignment is caused by radar jamming, the position offset map of the pixels under the BEV perspective of the lidar is calculated using the motion feature map at the previous moment, the point cloud BEV features at the current moment are calculated based on the point cloud BEV features and the position offset map at the previous moment, and the fused BEV features at the current moment are calculated based on the point cloud BEV features at the current moment and the image BEV features of the RGB image at the current moment; including:
[0036] Use the image feature extractor to process the RGB image frame at the current moment to obtain the image BEV feature at the current moment ;
[0037] Using the motion feature map of the previous moment , calculate from the previous moment To the current moment Position offset map within :
[0038] ;
[0039] Calculate the point cloud BEV features at the current time T :
[0040] ;
[0041] in, Indicates the last moment Point cloud BEV features;
[0042] Fusion BEV characteristics at the current moment for:
[0043] ;
[0044] in, Represents the addition operation.
[0045] In one possible implementation, the final fused BEV feature is calculated using the fused BEV features of the two consecutive moments before the current moment and the corresponding motion feature maps, and the fused BEV feature of the current moment; including:
[0046] Get the last moment Fusion BEV features and the motion feature map of the previous moment ;
[0047] Calculate from the last moment To the current moment Position offset diagram within :
[0048] ;
[0049] Get the moment Fusion features and time Motion feature map ;
[0050] Calculate from time To the current moment Position offset diagram within :
[0051] ;
[0052] Calculate the final fusion BEV features at the current moment :
[0053] ;
[0054] in, Represents the addition operation.
[0055] In a second aspect, an embodiment of the present application provides a three-dimensional target detection device with multi-sensor time adaptive synchronization, comprising:
[0056] An acquisition unit, used to acquire an RGB image frame and a 3D point cloud data frame of the target area at the current moment;
[0057] A processing unit, configured to process a sequence of continuous three-dimensional point cloud data frames before a current moment to obtain a motion feature map of a previous moment;
[0058] A judgment unit; used to use the timestamp of the current RGB image frame and the timestamp of the current 3D point cloud data frame to judge whether the current RGB image frame and the current 3D point cloud data frame are aligned, if so, enter the first modality fusion unit; otherwise, enter the second modality fusion unit;
[0059] The first modality fusion unit is used to fuse the RGB image frame at the current moment and the three-dimensional point cloud data frame at the current moment to obtain the fused BEV feature at the current moment;
[0060] The second modal fusion unit is used to calculate the position offset map of pixels under the BEV perspective of the laser radar using the motion feature map at the previous moment when the misalignment is caused by camera jamming, calculate the image BEV feature at the current moment based on the image BEV feature and the position offset map at the previous moment, and calculate the fused BEV feature at the current moment based on the point cloud BEV feature of the three-dimensional point cloud data frame at the current moment and the image BEV feature at the current moment; when the misalignment is caused by radar jamming, calculate the position offset map of pixels under the BEV perspective of the laser radar using the motion feature map at the previous moment, calculate the point cloud BEV feature at the current moment based on the point cloud BEV feature and the position offset map at the previous moment, and calculate the fused BEV feature at the current moment based on the point cloud BEV feature and the image BEV feature of the RGB image at the current moment;
[0061] The time series fusion unit is used to calculate the final fused BEV feature by using the fused BEV features of the two consecutive moments before the current moment and the corresponding motion feature map, and the fused BEV feature of the current moment;
[0062] The target detection unit is used to process the final fused BEV features using the detection head to obtain three-dimensional target detection results.
[0063] In a third aspect, an embodiment of the present application provides an electronic device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method of the embodiment of the present application when executing the computer program.
[0064] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method of the embodiment of the present application is implemented.
[0065] This application improves the robustness of multimodal data fusion and improves target detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0067] Figure 1 A flowchart of a three-dimensional target detection method with multi-sensor time adaptive synchronization provided in an embodiment of the present application;
[0068] Figure 2 This is a functional structure diagram of a multi-sensor time-adaptive synchronized three-dimensional target detection device provided by an embodiment of the present application;
[0069] Figure 3 This is a functional structure diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.
[0071] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.
[0072] First, a brief introduction to the design concept of the embodiments of the present application is given.
[0073] Motion models can reflect the future speed and direction of an object's movement and provide a wealth of useful information about environmental changes. Using motion models, an object's future position can be predicted. To this end, this application provides a multi-sensor time-adaptive synchronization method for 3D target detection. This method uses a motion model with physical prior information to compensate for missing BEV features and interpolate them. This method fuses multimodal data and time series data in two dimensions, achieving more robust data fusion. Fusion at both the feature and time series levels offers a high degree of physical interpretability and can improve the robustness of the multimodal fusion model.
[0074] After introducing the application scenarios and design concepts of the embodiments of the present application, the technical solutions provided by the embodiments of the present application are described below.
[0075] like Figure 1 As shown, the present application provides a three-dimensional target detection method with multi-sensor time adaptive synchronization, comprising the following steps:
[0076] Step 101: Acquire an RGB image frame and a 3D point cloud data frame of the target area at the current moment;
[0077] Step 102: Process the continuous 3D point cloud data frame sequence before the current moment to obtain the motion feature map of the previous moment;
[0078] The generation of motion map includes the following two steps:
[0079] Time series point cloud (LiDAR) data is usually a collection of point clouds, each point contains information such as three-dimensional position and reflection intensity. First, the time series information of laser scanning is extracted to obtain the point cloud BEV feature map, the size of which is , the number of channels is ; Each grid can be regarded as a pixel, and each pixel is related to the features along the height dimension, retaining the height information and metric space, and containing rich physical prior knowledge.
[0080] The point cloud BEV feature map is input into the spatiotemporal pyramid network (STPN), which consists of multiple spatiotemporal convolution layers and global temporal pooling layers. Each spatiotemporal convolution layer consists of standard 2D convolution and 3D convolution to capture spatial and temporal features respectively. The global temporal pooling layer fuses spatiotemporal features at different levels to obtain the motion feature map of the previous moment. ,in, Pixels The eigenvalues of Pixels Predicted speeds in both directions; For the current moment, The time interval between the current moment and the previous moment;
[0081] Motion Map predicts the motion state of BEV pixels by analyzing the differences between feature frames. For example, moving objects will have spatial features that change between frames, while stationary objects will not change significantly between frames. Therefore, Motion Map can be used to segment moving objects from stationary objects (including background).
[0082] Step 103: Using the timestamp of the current RGB image frame and the timestamp of the current 3D point cloud data frame, determine whether the current RGB image frame and the current 3D point cloud data frame are aligned. If so, proceed to step 104; otherwise, proceed to step 105.
[0083] The judgment process includes:
[0084] Calculate the absolute value of the difference between the timestamp of the RGB image frame at the current moment and the timestamp of the 3D point cloud data frame at the current moment;
[0085] Determine whether the absolute value of the difference is greater than a preset threshold. If so, determine that the RGB image frame at the current moment and the 3D point cloud data frame at the current moment are not aligned; otherwise, determine that the RGB image frame at the current moment and the 3D point cloud data frame at the current moment are aligned.
[0086] Step 104: fusing the RGB image frame at the current moment with the 3D point cloud data frame at the current moment to obtain the fused BEV feature at the current moment;
[0087] In this embodiment, the step includes:
[0088] The Lidar Feat Extractor is used to process the three-dimensional point cloud data frame at the current moment to obtain the point cloud BEV feature at the current moment. ; For the current moment;
[0089] Use the Camera Feat Extractor to process the RGB image frame at the current moment to obtain the BEV features of the image at the current moment ;
[0090] Calculate the fused BEV features at the current moment :
[0091] ;
[0092] in, Represents the addition operation.
[0093] Step 105: When the misalignment is caused by camera jamming, the position offset map of the pixels under the BEV perspective of the laser radar is calculated using the motion feature map at the previous moment, the image BEV feature of the current moment is calculated based on the image BEV feature and the position offset map at the previous moment, and the fused BEV feature of the current moment is calculated based on the point cloud BEV feature of the three-dimensional point cloud data frame at the current moment and the image BEV feature at the current moment; When the misalignment is caused by radar jamming, the position offset map of the pixels under the BEV perspective of the laser radar is calculated using the motion feature map at the previous moment, the point cloud BEV feature of the current moment is calculated based on the point cloud BEV feature and the position offset map at the previous moment, and the fused BEV feature of the current moment is calculated based on the point cloud BEV feature and the image BEV feature of the RGB image at the current moment;
[0094] If the misalignment is caused by camera stuck, the specific steps include:
[0095] Use the point cloud feature extractor to process the three-dimensional point cloud data at the current moment to obtain the point cloud BEV feature at the current moment ;
[0096] Using the motion feature map of the previous moment , calculate from the previous moment To the current moment Position offset diagram within :
[0097] ;
[0098] Calculate the current time Image BEV features :
[0099] ;
[0100] in, Indicates the last moment Image BEV features;
[0101] Fusion BEV characteristics at the current moment for:
[0102] ;
[0103] in, Represents the addition operation.
[0104] If the misalignment is caused by LiDAR stuck, the specific steps include:
[0105] Use the image feature extractor to process the RGB image frame at the current moment to obtain the image BEV feature at the current moment ;
[0106] Using the motion feature map of the previous moment , calculate from the previous moment To the current moment Position offset map within :
[0107] ;
[0108] Calculate the point cloud BEV features at the current time T :
[0109] ;
[0110] in, Indicates the last moment Point cloud BEV features;
[0111] Fusion BEV characteristics at the current moment for:
[0112] ;
[0113] in, Represents the addition operation.
[0114] Step 106: Calculate the final fused BEV feature using the fused BEV features of the two consecutive moments before the current moment and the corresponding motion feature maps, and the fused BEV feature of the current moment;
[0115] In this application, by 、 as well as The fused features of these three moments are fused to obtain the fused features in the time dimension.
[0116] The specific steps are as follows:
[0117] Get the last moment Fusion BEV features and the motion feature map of the previous moment ;
[0118] Calculate from the last moment To the current moment Position offset map within :
[0119] ;
[0120] Get the moment Fusion features and time Motion feature map ;
[0121] Calculate from time To the current moment Position offset map within :
[0122] ;
[0123] Calculate the final fusion BEV features at the current moment :
[0124] ;
[0125] Step 107: Use the detection head to process the final fused BEV features to obtain a three-dimensional target detection result.
[0126] Based on the above embodiments, the present application provides a three-dimensional target detection device with multi-sensor time adaptive synchronization. Figure 2 As shown, the multi-sensor time adaptive synchronization 3D target detection device 200 provided in the embodiment of the present application includes at least:
[0127] An acquisition unit 201 is configured to acquire an RGB image frame and a 3D point cloud data frame of a target area at a current moment;
[0128] The processing unit 202 is used to process the continuous three-dimensional point cloud data frame sequence before the current moment to obtain the motion feature map of the previous moment;
[0129] The judgment unit 203 is used to use the timestamp of the current RGB image frame and the timestamp of the current 3D point cloud data frame to judge whether the current RGB image frame and the current 3D point cloud data frame are aligned. If so, the process enters the first modality fusion unit; otherwise, the process enters the second modality fusion unit.
[0130] The first modality fusion unit 204 is configured to fuse the RGB image frame at the current moment with the 3D point cloud data frame at the current moment to obtain a fused BEV feature at the current moment;
[0131] The second modal fusion unit 205 is used to calculate the position offset map of pixels under the BEV perspective of the laser radar using the motion feature map at the previous moment when the misalignment is caused by camera jamming, calculate the image BEV feature at the current moment based on the image BEV feature and the position offset map at the previous moment, and calculate the fused BEV feature at the current moment based on the point cloud BEV feature of the three-dimensional point cloud data frame at the current moment and the image BEV feature at the current moment; when the misalignment is caused by radar jamming, calculate the position offset map of pixels under the BEV perspective of the laser radar using the motion feature map at the previous moment, calculate the point cloud BEV feature at the current moment based on the point cloud BEV feature and the position offset map at the previous moment, and calculate the fused BEV feature at the current moment based on the point cloud BEV feature and the image BEV feature of the RGB image at the current moment;
[0132] The time series fusion unit 206 is used to calculate the final fused BEV feature using the fused BEV features of the two consecutive moments before the current moment and the corresponding motion feature map, and the fused BEV feature of the current moment;
[0133] The target detection unit 207 is used to process the final fused BEV features using the detection head to obtain a three-dimensional target detection result.
[0134] It should be noted that the principle of solving the technical problem of the multi-sensor time adaptively synchronized three-dimensional target detection device 200 provided in the embodiment of the present application is similar to the method provided in the embodiment of the present application. Therefore, the implementation of the multi-sensor time adaptively synchronized three-dimensional target detection device 200 provided in the embodiment of the present application can refer to the implementation of the method provided in the embodiment of the present application, and the repeated parts will not be repeated.
[0135] Based on the above embodiments, the present application also provides an electronic device, referring to Figure 3 As shown, the electronic device 300 provided in the embodiment of the present application includes at least: a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program, the three-dimensional target detection method with multi-sensor time adaptive synchronization provided in the embodiment of the present application is implemented.
[0136] The electronic device 300 provided in the embodiment of the present application may further include a bus 303 connecting different components (including the processor 301 and the memory 302). The bus 303 represents one or more of several types of bus structures, including a memory bus, a peripheral bus, a local bus, and the like.
[0137] The memory 302 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 3021 and / or a cache memory 3022 , and may further include a read-only memory (ROM) 3023 .
[0138] The memory 302 may also include a program tool 3025 having a set (at least one) of program modules 3024, including but not limited to: an operating subsystem, one or more application programs, other program modules and program data, each of which or some combination may include an implementation of a network environment.
[0139] The electronic device 300 may also communicate with one or more external devices 304 (e.g., keyboards, remote controls, etc.), one or more devices that enable a user to interact with the electronic device 300 (e.g., mobile phones, computers, etc.), and / or any device that enables the electronic device 300 to communicate with one or more other electronic devices 300 (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 305. Furthermore, the electronic device 300 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 306. Figure 3 As shown, the network adapter 306 communicates with other modules of the electronic device 300 via the bus 303. Figure 3 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 300, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, disk array (Redundant Arrays of Independent Disks, RAID) subsystems, tape drives, and data backup storage subsystems.
[0140] It should be noted that Figure 3 The electronic device 300 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0141] The present application also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the multi-sensor time-adaptive synchronization 3D object detection method provided in the present application. Specifically, the executable program can be built into or installed in the electronic device 300. Thus, the electronic device 300 can implement the multi-sensor time-adaptive synchronization 3D object detection method provided in the present application by executing the built-in or installed executable program.
[0142] The method provided in the embodiment of the present application can also be implemented as a program product, which includes a program code. When the program product can be run on the electronic device 300, the program code is used to enable the electronic device 300 to execute the multi-sensor time adaptive synchronization three-dimensional target detection method provided in the embodiment of the present application.
[0143] The program product provided in the embodiments of the present application may adopt any combination of one or more readable media, wherein the readable medium may be a readable signal medium or a readable storage medium, and the readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any combination of the above. Specifically, more specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, RAM, ROM, Erasable Programmable Read Only Memory (EPROM), optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0144] The program product provided in the embodiments of the present application may be a CD-ROM and include program code, and may also be run on a computing device. However, the program product provided in the embodiments of the present application is not limited thereto. In the embodiments of the present application, the readable storage medium may be any tangible medium containing or storing a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0145] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.
[0146] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0147] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of this application and are not intended to limit the scope of the present invention. Although this application has been described in detail with reference to the embodiments, it should be understood by those skilled in the art that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application and should be encompassed by the claims of this application.
Claims
1. A three-dimensional target detection method with multi-sensor time adaptive synchronization, characterized in that: include: Obtain the RGB image frame and 3D point cloud data frame of the target area at the current moment; Process the continuous 3D point cloud data frame sequence before the current moment to obtain the motion feature map of the previous moment; Using the timestamp of the current RGB image frame and the timestamp of the current 3D point cloud data frame, determine whether the current RGB image frame and the current 3D point cloud data frame are aligned. If so, fuse the current RGB image frame and the current 3D point cloud data frame to obtain the fused BEV feature at the current moment. Otherwise, when the misalignment is caused by camera jamming, the position offset map of the pixels under the BEV perspective of the laser radar is calculated using the motion feature map of the previous moment, the image BEV features of the current moment are calculated based on the image BEV features and the position offset map of the previous moment, and the fused BEV features of the current moment are calculated based on the point cloud BEV features of the three-dimensional point cloud data frame at the current moment and the image BEV features of the current moment; when the misalignment is caused by radar jamming, the position offset map of the pixels under the BEV perspective of the laser radar is calculated using the motion feature map of the previous moment, the point cloud BEV features of the current moment are calculated based on the point cloud BEV features and the position offset map of the previous moment, and the fused BEV features of the current moment are calculated based on the point cloud BEV features of the current moment and the image BEV features of the RGB image at the current moment; The final fused BEV feature is calculated using the fused BEV features of the two consecutive moments before the current moment and the corresponding motion feature map, and the fused BEV feature of the current moment; The final fused BEV features are processed by the detection head to obtain the three-dimensional target detection results.
2. The multi-sensor time adaptive synchronization 3D target detection method according to claim 1, characterized in that: Process the continuous 3D point cloud data frame sequence before the current moment to obtain the motion feature map of the previous moment; including: Perform feature extraction on the continuous 3D point cloud data frame sequence before the current moment to obtain the point cloud BEV feature sequence; The pre-trained spatiotemporal pyramid network is used to process the point cloud BEV feature sequence to obtain the motion feature map of the previous moment. ,in, Pixels The eigenvalues of Pixels Predicted speeds in both directions; For the current moment, The time interval between the current moment and the previous moment; The spatiotemporal pyramid network consists of multiple spatiotemporal convolutional layers and global temporal pooling layers, and each spatiotemporal convolutional layer consists of standard two-dimensional convolution and three-dimensional convolution.
3. The multi-sensor time adaptive synchronization 3D target detection method according to claim 1, characterized in that: Using the timestamp of the current RGB image frame and the timestamp of the current 3D point cloud data frame, determine whether the current RGB image frame and the current 3D point cloud data frame are aligned; including: Calculate the absolute value of the difference between the timestamp of the RGB image frame at the current moment and the timestamp of the 3D point cloud data frame at the current moment; Determine whether the absolute value of the difference is greater than a preset threshold. If so, determine that the RGB image frame at the current moment and the 3D point cloud data frame at the current moment are not aligned; otherwise, determine that the RGB image frame at the current moment and the 3D point cloud data frame at the current moment are aligned.
4. The multi-sensor time adaptive synchronization 3D target detection method according to claim 1, characterized in that: The RGB image frame at the current moment and the 3D point cloud data frame at the current moment are fused to obtain the fused BEV features at the current moment; including: Use the point cloud feature extractor to process the three-dimensional point cloud data frame at the current moment to obtain the point cloud BEV feature at the current moment ; For the current moment; Use the image feature extractor to process the RGB image frame at the current moment to obtain the image BEV feature at the current moment ; Calculate the fused BEV features at the current moment : ; in, Represents the addition operation.
5. The method for three-dimensional target detection with multi-sensor time adaptive synchronization according to claim 2, characterized in that: When the misalignment is caused by camera jamming, the position offset map of the pixels under the BEV perspective of the lidar is calculated using the motion feature map at the previous moment. The image BEV features at the current moment are calculated based on the image BEV features and the position offset map at the previous moment. The fused BEV features at the current moment are calculated based on the point cloud BEV features of the current 3D point cloud data frame and the image BEV features at the current moment. This includes: Use the point cloud feature extractor to process the three-dimensional point cloud data at the current moment to obtain the point cloud BEV feature at the current moment ; Using the motion feature map of the previous moment , calculate from the previous moment To the current moment Position offset diagram within : ; Calculate the current time Image BEV features : ; in, Indicates the last moment Image BEV features; Fusion BEV characteristics at the current moment for: ; in, Represents the addition operation.
6. The method for three-dimensional target detection with multi-sensor time adaptive synchronization according to claim 2, characterized in that: When the misalignment is caused by radar jamming, the position offset map of the pixels under the BEV perspective of the lidar is calculated using the motion feature map at the previous moment. The point cloud BEV features at the current moment are calculated based on the point cloud BEV features and the position offset map at the previous moment. The fused BEV features at the current moment are calculated based on the point cloud BEV features at the current moment and the image BEV features of the RGB image at the current moment. This includes: Use the image feature extractor to process the RGB image frame at the current moment to obtain the image BEV feature at the current moment ; Using the motion feature map of the previous moment , calculate from the previous moment To the current moment Position offset diagram within : ; Calculate the point cloud BEV features at the current time T : ; in, Indicates the last moment Point cloud BEV features; Fusion BEV characteristics at the current moment for: ; in, Represents the addition operation.
7. The method for three-dimensional target detection with multi-sensor time adaptive synchronization according to claim 4, 5 or 6, characterized in that: The final fused BEV feature is calculated using the fused BEV features of the two consecutive moments before the current moment and the corresponding motion feature map, and the fused BEV feature of the current moment; including: Get the last moment Fusion BEV features and the motion feature map of the previous moment ; Calculate from the last moment To the current moment Position offset diagram within : ; Get the moment Fusion features and time Motion feature map ; Calculate from time To the current moment Position offset map within : ; Calculate the final fusion BEV features at the current moment : ; in, Represents the addition operation.
8. A three-dimensional target detection device with multi-sensor time adaptive synchronization, characterized in that: include: An acquisition unit, used to acquire an RGB image frame and a 3D point cloud data frame of the target area at the current moment; A processing unit, configured to process a sequence of continuous three-dimensional point cloud data frames before a current moment to obtain a motion feature map of a previous moment; judgment unit; Used to use the timestamp of the current RGB image frame and the timestamp of the current 3D point cloud data frame to determine whether the current RGB image frame and the current 3D point cloud data frame are aligned, and if so, enter the first modality fusion unit; Otherwise, enter the second modality fusion unit; The first modality fusion unit is used to fuse the RGB image frame at the current moment and the three-dimensional point cloud data frame at the current moment to obtain the fused BEV feature at the current moment; The second modal fusion unit is used to calculate the position offset map of pixels under the BEV perspective of the laser radar using the motion feature map at the previous moment when the misalignment is caused by camera jamming, calculate the image BEV feature at the current moment based on the image BEV feature and the position offset map at the previous moment, and calculate the fused BEV feature at the current moment based on the point cloud BEV feature of the three-dimensional point cloud data frame at the current moment and the image BEV feature at the current moment; when the misalignment is caused by radar jamming, calculate the position offset map of pixels under the BEV perspective of the laser radar using the motion feature map at the previous moment, calculate the point cloud BEV feature at the current moment based on the point cloud BEV feature and the position offset map at the previous moment, and calculate the fused BEV feature at the current moment based on the point cloud BEV feature and the image BEV feature of the RGB image at the current moment; The time series fusion unit is used to calculate the final fused BEV feature by using the fused BEV features of the two consecutive moments before the current moment and the corresponding motion feature map, and the fused BEV feature of the current moment; The target detection unit is used to process the final fused BEV features using the detection head to obtain three-dimensional target detection results.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Target detection method and device based on multi-modal sequence data fusion
CN115496977A
Target detection method for low-bandwidth vehicle and road feature fusion known to lighthouse
CN116977957A