Data fusion method and device, computer equipment and storage medium
By extracting and fusion of image and radar data, and aggregating feature based on velocity reference vectors, the problem of scheduling of sensor data timing fusion in autonomous driving and robot autonomous mobile scenes is solved, and the accuracy of target detection is improved.
Patent Information
- Application Number
- CN202311629336.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-25
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-06
AI Technical Summary
In autonomous driving and autonomous mobile scenarios, the timing fusion of sensor data has the problem of space-time misalignment of detection data, resulting in shading and motion blur, affecting the accuracy of target detection.
By acquiring image and radar point cloud data, feature extraction and mapping are performed to a three-dimensional coordinate system, point cloud fusion features are fused, and feature aggregation and fusion are performed based on multiple preset velocity reference vectors to obtain timing fusion features for target detection.
It effectively alleviates the problem of striking motion targets in the scene, improves the accuracy of target detection and the perception effect of the system.
Smart Images

Figure CN119942276A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of environmental perception technology, and in particular to a data fusion method, device, computer equipment and storage medium. Background Art
[0002] In some scenarios, such as the vehicle's autonomous driving scenario, the robot's autonomous movement scenario, etc., it is necessary to perceive the external environment in real time. In order to meet the needs of environmental perception, autonomous vehicles, mobile robots, etc. are usually equipped with various sensors, such as millimeter-wave radar, laser radar, cameras, etc., to perceive the surrounding environment and target information. The information obtained by the above sensors can be used to classify and identify the surrounding environment and objects.
[0003] In the relevant technical solutions, taking the autonomous driving scenario as an example, in the autonomous driving scenario, since both the vehicle itself and the detection target may move, the relative spatial relationship between the on-board sensor and the detection target will also change over time. Therefore, when using multi-frame time-series sensor information for fusion, there is a problem of spatial and temporal misalignment of the detection data, such as data blur and ghosting. Traditional sensor time-series fusion solutions generally do not consider motion compensation, or only consider the motion compensation of the vehicle itself, while ignoring the compensation for moving targets of different speeds in the scene, resulting in a decrease in the vehicle's detection capability. Summary of the invention
[0004] The purpose of this application is to provide a data fusion method to effectively alleviate the problem of ghosting of moving targets in a scene and improve the accuracy of target detection.
[0005] Based on the above purpose, the present application provides a data fusion method, which includes:
[0006] Acquire images captured by at least one image acquisition device and radar point cloud data collected by at least one radar in the test area;
[0007] Extract features from the image data, map the acquired feature map to a preset three-dimensional coordinate system to generate a first visual point cloud feature, map the radar point cloud data to a three-dimensional coordinate system to generate a radar point cloud feature, and fuse the first visual point cloud feature with the radar point cloud feature to generate a point cloud fusion feature;
[0008] Extract the feature of the point cloud fusion to generate the bird's-eye view feature at the current moment;
[0009] Based on multiple preset speed reference vectors, the bird's-eye view features at the current moment are aggregated with the bird's-eye view features at multiple historical moments to obtain the speed-separated bird's-eye view features;
[0010] The speed-separated bird's-eye view features are fused with the current moment's bird's-eye view features to obtain time series fusion features;
[0011] Object detection based on temporal fusion features.
[0012] Furthermore, based on multiple preset speed reference vectors, the bird's-eye view features at the current moment are aggregated with the bird's-eye view features at multiple historical moments to obtain speed-separated bird's-eye view features, including:
[0013] A neural network is used to train the model on the target speed prior distribution data of multiple scenes to obtain multiple speed query vectors;
[0014] A multi-layer perceptron is used to transform the dimensions of multiple speed query vectors to obtain multiple speed reference vectors;
[0015] Based on multiple speed reference vectors, the bird's-eye view features at the current moment are convolved with the bird's-eye view features at multiple historical moments to obtain speed-separated bird's-eye view features;
[0016] The speed-separated bird's-eye view features are fused with the current bird's-eye view features to obtain the time series fusion features, which also include:
[0017] A multi-layer perceptron is used to transform the dimension of the speed-separated bird's-eye view features to obtain speed-separated bird's-eye view features with the same dimension as the bird's-eye view features at the current moment;
[0018] The bird's-eye view features separated by speed in the same dimension are concatenated with the bird's-eye view feature dimension at the current moment to obtain the time series fusion features.
[0019] Furthermore, the velocity-separated bird's-eye view features y k ∈R W×H×1 It is expressed as follows:
[0020]
[0021] Among them, i∈{1,...W},j∈{1,...,H},c∈{1,...,256}, i, j and c represent the index of the bird's-eye view feature in width, height and channel dimensions respectively, and w∈R 3×3×256 is the parameter of the speed-sensitive convolution kernel, Δp∈{(-1,-1),(-1,0),...,(0,1),(1,1)}, α T ∈R is the decay factor related to the timestamp interval, Δp k,T ∈R 2 is the spatial motion offset vector in the bird's-eye view feature. The spatial motion offset vector is calculated by the reference speed, timestamp interval and ego-vehicle pose transformation matrix. TThe bird's-eye view features of the current moment and the bird's-eye view features of multiple historical moments. Further, the acquired feature map is mapped to a preset three-dimensional coordinate system to generate a first visual point cloud feature, the radar point cloud data is mapped to a three-dimensional coordinate system to generate a radar point cloud feature, and the first visual point cloud feature and the radar point cloud feature are fused to generate a point cloud fusion feature, including:
[0022] Mapping each point on the acquired feature map to a three-dimensional coordinate system to generate a first visual point cloud feature, wherein each point feature in the first visual point cloud feature is represented by information combined with the three-dimensional spatial coordinate value, timestamp, and image feature of the corresponding point;
[0023] The radar point cloud data is mapped to a three-dimensional coordinate system to generate radar point cloud features. Each point feature in the radar point cloud feature is represented by the three-dimensional spatial coordinate value of the corresponding point, the timestamp, and the radar's perception attributes.
[0024] The first visual point cloud features are fused with the radar point cloud features to obtain point cloud fusion features.
[0025] Furthermore, the radar point cloud data is mapped to a three-dimensional coordinate system to generate radar point cloud features, including:
[0026] The radar is a laser radar, and laser point cloud data collected by the laser radar is obtained;
[0027] Mapping the laser point cloud data to a three-dimensional coordinate system to generate a first laser point cloud feature, wherein each point feature in the first laser point cloud feature is represented by information combined with a three-dimensional spatial coordinate value, a timestamp, and a signal strength of the corresponding point;
[0028] and / or,
[0029] The radar is a millimeter wave radar, and millimeter wave point cloud data collected by the millimeter wave radar is obtained;
[0030] The millimeter wave point cloud data is mapped to a three-dimensional coordinate system to generate a first millimeter wave point cloud feature, and each point feature in the first millimeter wave point cloud feature is characterized by information combined with the three-dimensional spatial coordinate value, timestamp, speed and signal-to-noise ratio of the corresponding point.
[0031] Furthermore, the first visual point cloud features are fused with the radar point cloud features to obtain point cloud fusion features, including:
[0032] A multi-layer perceptron is used to map each point feature in the first visual point cloud feature to obtain an image feature vector of each point, and the image feature vector, three-dimensional space coordinate value and timestamp information of each point are combined to generate a second visual point cloud feature;
[0033] A multi-layer perceptron is used to map each point feature in the first laser point cloud feature to obtain a laser radar feature vector of each laser radar point, and the laser radar feature vector, three-dimensional space coordinate value and timestamp information of each laser radar point are combined to generate a second laser point cloud feature;
[0034] A multi-layer perceptron is used to map each point feature in the first millimeter-wave point cloud feature to obtain a millimeter-wave radar feature vector for each millimeter-wave radar point, and the millimeter-wave radar feature vector, three-dimensional space coordinate value and timestamp information of each millimeter-wave radar point are combined to generate a second millimeter-wave point cloud feature, wherein the image feature vector, the laser radar feature vector and the millimeter-wave radar feature vector have the same dimension;
[0035] After the second visual point cloud feature, the second laser point cloud feature and the second millimeter wave point cloud feature are dimensional-joined, a point cloud fusion feature is obtained.
[0036] Based on the above purpose, the present application provides a data fusion device, which includes:
[0037] An acquisition module, used to acquire images acquired by at least one image acquisition device and point cloud data acquired by at least one radar in the area to be measured;
[0038] A first fusion module is used to extract features from image data, map the acquired feature map to a preset three-dimensional coordinate system to generate a first visual point cloud feature, map the radar point cloud data to a three-dimensional coordinate system to generate a radar point cloud feature, and fuse the first visual point cloud feature with the radar point cloud feature to generate a point cloud fusion feature;
[0039] The bird's-eye view feature module is used to extract the point cloud fusion features and generate the bird's-eye view features at the current moment;
[0040] A time series feature aggregation module is used to aggregate the bird's-eye view features at the current moment with the bird's-eye view features at multiple historical moments based on multiple preset speed reference vectors to obtain the bird's-eye view features separated by speed;
[0041] The second fusion module is used to fuse the speed-separated bird's-eye view features with the bird's-eye view features at the current moment to obtain a time series fusion feature.
[0042] Furthermore,
[0043] The time series feature aggregation module is specifically used for:
[0044] A neural network is used to train the model on the target speed prior distribution data of multiple scenes to obtain multiple speed query vectors;
[0045] A multi-layer perceptron is used to transform the dimensions of multiple speed query vectors to obtain multiple speed reference vectors;
[0046] Based on multiple speed reference vectors, the bird's-eye view features at the current moment are convolved with the bird's-eye view features at multiple historical moments to obtain speed-separated bird's-eye view features;
[0047] and / or,
[0048] The second fusion module is specifically used for:
[0049] A multi-layer perceptron is used to transform the dimension of the speed-separated bird's-eye view features to obtain speed-separated bird's-eye view features with the same dimension as the bird's-eye view features at the current moment;
[0050] The bird's-eye view features separated by speed in the same dimension are concatenated with the bird's-eye view feature dimension of the current moment to obtain the time series fusion features;
[0051] and / or,
[0052] The first fusion module specifically includes:
[0053] A visual point cloud feature generating unit, used for mapping each point on the acquired feature map to a three-dimensional coordinate system to generate a first visual point cloud feature, wherein each point feature in the first visual point cloud feature is represented by information combined with a three-dimensional spatial coordinate value, a timestamp, and an image feature of the corresponding point;
[0054] A radar point cloud feature generation unit is used to map radar point cloud data to a three-dimensional coordinate system to generate radar point cloud features. Each point feature in the radar point cloud feature is represented by information combined with the three-dimensional spatial coordinate value of the corresponding point, a timestamp, and the perception attributes of the radar.
[0055] The radar point cloud feature generation unit is specifically used for:
[0056] The radar is a laser radar, and laser point cloud data collected by the laser radar is obtained;
[0057] Mapping the laser point cloud data to a three-dimensional coordinate system to generate a first laser point cloud feature, wherein each point feature in the first laser point cloud feature is represented by information combined with a three-dimensional spatial coordinate value, a timestamp, and a signal strength of the corresponding point;
[0058] and / or,
[0059] The radar point cloud feature generation unit is specifically used for:
[0060] The radar is a millimeter wave radar, and millimeter wave point cloud data collected by the millimeter wave radar is obtained;
[0061] The millimeter wave point cloud data is mapped to a three-dimensional coordinate system to generate a first millimeter wave point cloud feature, and each point feature in the first millimeter wave point cloud feature is characterized by information combined with the three-dimensional spatial coordinate value, timestamp, speed and signal-to-noise ratio of the corresponding point.
[0062] Based on the above objectives, the present application provides a computer device, including:
[0063] It includes: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0064] Memory, used to store computer programs;
[0065] The processor is used to implement the steps of the method described above when executing the computer program stored in the memory.
[0066] Based on the above purpose, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the method described above are implemented.
[0067] This application can effectively alleviate the problems of ghosting and motion blur of moving targets in the scene, enhance the perception effect of multiple sensor systems, and improve the accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a functional block diagram of a vehicle provided according to an embodiment of the present application;
[0069] Figure 2 is a schematic diagram of sensor distribution applied to a vehicle according to an embodiment of the present application;
[0070] Figure 3 is a schematic block diagram of a system architecture provided according to an embodiment of the present application;
[0071] Figure 4 is a first flow chart of a data fusion method provided according to an embodiment of the present application;
[0072] Figure 5 is a second flow chart of the data fusion method provided according to an embodiment of the present application;
[0073] Figure 6 It is a schematic diagram of the aggregation of time series features of a bird's-eye view provided according to an embodiment of the present application;
[0074] Figure 7 is a third flow chart of the data fusion method provided according to an embodiment of the present application;
[0075] Figure 8is a fourth flow chart of the data fusion method provided according to an embodiment of the present application;
[0076] Fig. 9 is a fifth flow chart of a data fusion method provided according to an embodiment of the present application;
[0077] Fig.10 is a sixth flow chart of the data fusion method provided according to an embodiment of the present application;
[0078] Fig.11 is a seventh flow chart of a data fusion method provided according to an embodiment of the present application;
[0079] Fig.12 is a flow chart of a target detection method provided according to an embodiment of the present application;
[0080] Fig.13 is a first system block diagram of a data fusion device provided according to an embodiment of the present application;
[0081] Fig.14 is a second system block diagram of a data fusion device provided according to an embodiment of the present application;
[0082] Fig.15 It is a schematic diagram of the structure of a computer device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0083] The present application will be described in detail below in conjunction with the specific implementation modes shown in the accompanying drawings, but these implementation modes do not limit the present application. Structural, methodological, or functional changes made by ordinary technicians in the field based on these implementation modes are included in the protection scope of the present application.
[0084] Figure 1 1 is a functional block diagram of a vehicle provided in an embodiment of the present application. The vehicle includes an image acquisition device 101, a radar 102, and a computing platform 103. The image acquisition device 101 and the radar 102 are respectively connected to the computing platform 103. It should be understood that the above connection method can be a wireless connection, such as a Bluetooth connection, a Wi-Fi connection, etc., or the above connection method can also be a wired connection, such as an optical fiber connection, etc., without limitation.
[0085] The image acquisition device 101 may be any one or more combinations of a camera, a video camera, a scanner, or other devices with a camera function (e.g., a mobile phone, a tablet computer, etc.). The image acquisition device 101 may also include a lens and an image sensor. The scene in the field of view of the image acquisition device 101 is projected onto the image sensor through the optical image generated by the lens, and the image sensor converts it into an electrical signal, and then obtains a visual image after processing such as analog-to-digital conversion.
[0086] The radar 102 may be any one or more combinations of a laser radar sensor, a millimeter wave radar sensor, etc. The radar 102 may detect targets in the current scene using electromagnetic waves to obtain point cloud data. For example, the radar installed on the vehicle may emit detection electromagnetic waves to surrounding obstacles, and receive echo signals generated by obstacles reflecting the detection electromagnetic waves, i.e., the above-mentioned point cloud data.
[0087] The computing platform 103 is used to process the visual image and point cloud data of the area to be tested by the data fusion method of the embodiment of the present application to detect the target object information of the area to be tested. The computing platform 103 may include processors 1031 to 103n (n is a positive integer), and the processor may be a circuit with the ability to read and run instructions, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), etc. In another implementation, the processor may be a hardware circuit that implements a certain function through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by a processor such as an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc. The computing platform 150 can also include a memory, the memory is used to store instructions, and some or all of the processors 131 to 131n can call the instructions in the memory and execute the quality to achieve the corresponding functions.
[0088] In an optional implementation, the image acquisition device 101, radar 102 and computing platform 103 can be set on the same device, such as a vehicle, drone, etc. In another optional implementation, the image acquisition device 101 and radar 102 can be set on the same device, such as a vehicle, drone, etc., and the computing platform 103 can be a terminal device to which the image acquisition device 101 and radar 102 are respectively connected, and the computing platform 103 can communicate with the device where the image acquisition device 101 and radar 102 are set to obtain the visual image and point cloud data of the area to be measured. The terminal device in the embodiment of the present application can be, for example, a mobile phone, a tablet computer, a desktop, a handheld computer, a laptop computer, a super mobile personal computer, a netbook, a cellular phone, a personal digital assistant, an augmented reality\virtual reality device, etc. It can interact with the user through one or more methods such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction or a handwriting device.
[0089] In the present application, a vehicle may include one or more different types of transportation vehicles, and may also include one or more different types of transportation vehicles or movable objects that operate or move on land (e.g., roads, roads, railways, etc.), water (e.g., waterways, rivers, oceans, etc.) or space. For example, a vehicle may include a car, a bicycle, a motorcycle, a train, a subway, an airplane, a ship, an aircraft, a robot or other types of transportation vehicles or movable objects.
[0090] Figure 2 This is a schematic diagram of sensor distribution applied to a vehicle provided in an embodiment of the present application. It should be understood that: Figure 2 This is only a schematic diagram of a sensor distribution method, and there may be other distribution methods, which are not limited in the embodiments of the present application. Figure 2 As shown, the sensors distributed on the vehicle include a millimeter wave radar 201, a camera 202 and a laser radar 203, and may also include Figure 2 Other sensors not shown in the figure are not limited in the embodiments of the present application.
[0091] Figure 3 Schematic diagram of the system architecture provided by the embodiment of the present application. Figure 3 The system includes a sensor and a data fusion device. For example, the sensor may include Figure 1 In one or more of the image acquisition device 101 and the radar 102, the sensor can be Figure 2 The distribution diagram shown is distributed on the vehicle, and can also be deployed in other distribution modes. The data fusion device can be deployed on Figure 1The computing platform 103 shown is detected by the vehicle-mounted computing platform 103, and can also be deployed on a cloud server, detected by the cloud server, and the fusion information is transmitted to the computing platform 103. The sensor is used to sense the vehicle and its surrounding environment and obtain data. For example, the sensor may include one or more cameras and one or more laser radars, and the data output by the sensor may be image data obtained by the camera and point cloud information obtained by the laser radar. The data fusion device can be used to obtain data collected by multiple sensors, and perform feature fusion on the data collected by multiple sensors to obtain fusion information. In one embodiment, the data fusion device can be deployed inside the vehicle in the form of a hardware module and / or a software module, for example, deployed on the computing platform 103, or the data fusion device can be a computing platform located in a cloud server.
[0092] Please refer to Figure 4 , a flow chart of a data fusion method provided in an embodiment of the present application includes:
[0093] S401, obtaining images captured by at least one image acquisition device and radar point cloud data captured by at least one radar in the area to be measured.
[0094] The area to be tested can be understood as the area within which the image acquisition device can capture images. Similarly, the area is also within the area that the radar can detect. For example, in a vehicle automatic driving scenario, the area to be tested can refer to the area in the surrounding environment of the vehicle where target detection is being performed. The image can be acquired by the image acquisition device installed on the vehicle for the area to be tested, and the image acquisition device can send the acquired image to the processor in this embodiment. Point cloud data is data acquired by the radar installed on the vehicle to detect the area to be tested, and the radar can also send the acquired point cloud data to the processor in this embodiment. Point cloud data is a massive set of points that express the spatial distribution of the target and the surface characteristics of the target in the same spatial reference system. After obtaining the spatial position of each reflection point in the area to be tested or on the surface of the object to be tested, a set of points is obtained, which is called a "point cloud". The processor obtains the image and point cloud data, and this embodiment is used to fuse the two.
[0095] S402, extract features from the image data, map the acquired feature map to a preset three-dimensional coordinate system to generate a first visual point cloud feature, map the radar point cloud data to a three-dimensional coordinate system to generate a radar point cloud feature, and fuse the first visual point cloud feature with the radar point cloud feature to generate a point cloud fusion feature.
[0096] Feature extraction refers to the method and process of extracting characteristic information from an image. For example, characteristic information in the image can be extracted by analyzing and transforming the image. As an optional implementation method, the image is input into a pre-trained feature extraction model for feature extraction to obtain one or more image features of the image, that is, to obtain the image features of each pixel in the image. The extracted image features can represent the semantic information of the image. Exemplarily, the feature extraction model can use a convolutional neural network to extract features from the acquired image. For example, the ResNet-101 network is used as the skeleton network for feature extraction. The ResNet-101 network includes a total of 5 network layers, namely, a 7×7 convolutional layer, a first stage, a second stage, a third stage, a fourth stage, a pooling layer, and a fully connected layer. Each stage is composed of a basic residual unit, which is used to extract the features of the input image and obtain output feature maps of different levels.
[0097] The acquired feature map is transformed in perspective, and the coded features of each image feature point on the feature map are mapped to a preset three-dimensional coordinate system to generate the first visual point cloud feature. By generating point cloud features in three-dimensional space based on two-dimensional image features and converting 2D images to 3D space, the 2D semantic information corresponding to the image in the area to be tested can be accurately obtained, providing accurate 3D scene expression for downstream 3D target detection.
[0098] As an example, the LSS (Lift, Splat, Shoot) algorithm is used to explicitly estimate the depth distribution of each image feature, and each feature point on each 2D feature map is "lifted" to a 3D coordinate point to obtain point cloud data containing image features. Using the known camera internal and external parameter matrix, it is converted to a preset three-dimensional coordinate system to obtain the first visual point cloud feature.
[0099] As an optional implementation method, the preset three-dimensional coordinate system is the vehicle's own coordinate system. According to the projection relationship defined by the vehicle's own coordinate system and the intrinsic parameter matrix and extrinsic parameter matrix of the image acquisition device, the image features are mapped to the vehicle's own coordinate system, so that the target positions corresponding to multiple positions in the image and the vehicle's own coordinate system can be determined. Among them, the intrinsic parameter matrix and the extrinsic parameter matrix can be pre-calibrated and stored in the processor. When obtaining the intrinsic parameter matrix and the extrinsic parameter matrix of the image acquisition device, the processor can directly read the intrinsic parameter matrix and the extrinsic parameter matrix locally; or, the intrinsic parameter matrix and the extrinsic parameter matrix of the image acquisition device can also be pre-calibrated and stored in the image acquisition device, and the processor can request the image acquisition device to obtain the intrinsic parameter matrix and the extrinsic parameter matrix.
[0100] Radar point cloud data may be the position information of the point cloud in space, and the position information is represented by the spatial coordinate values of the above-mentioned preset spatial coordinate system. The pixels of the radar point cloud data and the image are mapped to the same three-dimensional coordinate system, so that the radar point cloud data and the image can be represented in the same spatial coordinate. Exemplarily, according to the conversion relationship defined by the vehicle's own coordinate system and the radar's external parameter matrix, the three-dimensional coordinates of the radar point cloud data are mapped to the vehicle's own coordinate system, so that the radar point cloud data and the image data are mapped to the vehicle's own coordinate system, and in the vehicle's own coordinate system, the three-dimensional spatial coordinate values of the radar point cloud data and the image are represented, so that the multimodal data collected by multiple sensors can be expressed in a unified spatial coordinate system, so that the multimodal data representation has uniformity.
[0101] S403: extracting features from the point cloud fusion features to generate bird's-eye view features at the current moment.
[0102] The bird's-eye view feature (Bird's Eye View, BEV) is a point cloud image under the bird's-eye view perspective, which represents the projection of the point cloud on a plane perpendicular to the height direction. Usually before obtaining the bird's-eye view feature, the three-dimensional point cloud data is voxelized, that is, the space is divided into voxels to form a form in which multiple three-dimensional point cloud data correspond to one voxel. Voxel refers to a cuboid when the three-dimensional space is divided into cuboids of fixed size. The voxelized three-dimensional point cloud data is input into the encoder for encoding to obtain the three-dimensional point cloud feature code. The feature code is then projected to the bird's-eye view space (i.e., the three-dimensional space from a bird's-eye view perspective) and converted into bird's-eye view point cloud data. The bird's-eye view point cloud data is encoded according to the combined feature aggregator and the BEV encoder to obtain the bird's-eye view feature, that is, the bird's-eye view feature at the current moment with multimodal information fusion in the embodiment of the present application. The bird's-eye view feature can be obtained in a variety of ways, such as VoxelNet (Voxel Network), etc. For example, the bird's-eye view feature obtained at the current moment through the VoxelNet network is expressed as f t ∈R W*H*C , W represents the width of the bird's-eye view feature, H represents the height of the bird's-eye view feature, and C represents the number of channels of the bird's-eye view feature. The embodiments of the present application do not limit the method for obtaining the bird's-eye view image.
[0103] S404: Based on a plurality of preset speed reference vectors, feature aggregation is performed on the bird's-eye view features at the current moment and the bird's-eye view features at a plurality of historical moments to obtain speed-separated bird's-eye view features.
[0104] Through the preset speed reference vector, the bird's-eye view features of the current moment are interactively aggregated with the bird's-eye view features of multiple historical moments to obtain speed-separated bird's-eye view features, compensate for the moving targets of different speeds in the scene, and compensate for the movement of the target object at historical moments, and obtain the three-dimensional spatial offset of the target object at each historical moment, which can effectively aggregate the time-effect features and improve the efficiency of time series aggregation.
[0105] S405: Fusing the speed-separated bird's-eye view features with the current moment's bird's-eye view features to obtain time series fusion features.
[0106] The speed-separated bird's-eye view features are fused with the bird's-eye view feature dimensions at the current moment to obtain time series fusion features, so as to obtain data of higher dimensions. In this embodiment, based on the above-mentioned fusion of the bird's-eye view features at the current moment, another data fusion is performed to further improve the richness of the data.
[0107] S406: Detect targets based on temporal fusion features.
[0108] The temporal fusion feature is input into the pre-trained target detection model to obtain the target detection result output by the target detection model. For another example, the point cloud fusion feature can also be processed according to the target detection algorithm to obtain the target detection result. The temporal fusion feature carries multimodal data of images, lidar, and millimeter wave radar, effectively alleviating the problem of moving target smearing and further improving the accuracy of the detection results.
[0109] As an optional implementation, Figure 5 As shown, based on multiple preset speed reference vectors, the bird's-eye view features at the current moment are aggregated with the bird's-eye view features at multiple historical moments to obtain the speed-separated bird's-eye view features, and also include:
[0110] S501, using a neural network to perform model training on target speed prior distribution data of multiple scenes to obtain multiple speed query vectors.
[0111] Using the neural network model, the target speed prior distribution data obtained in each scene is used as the input of the model. The target speed prior distribution data of the scene is used to represent the movement of each target in the scene. After the data training of the model, the speed query vector is obtained. K and d q Indicates the vector dimension of the model output. In a typical embodiment, K = 128, d q = 128. The scenarios can be understood as high-speed driving scenarios, park driving scenarios, etc.
[0112] S502: Use a multi-layer perceptron to perform dimension transformation on multiple speed query vectors to obtain multiple speed reference vectors.
[0113] A multi-layer perceptron is used to map and transform the speed query vector to generate a speed reference vector V = {v1, v2, ..., v K}∈R K×2 , v k ∈R 2 represents a set of velocity vectors in the bird's-eye view space. In a typical embodiment, K=128. Based on this implementation, the velocity reference vectors in various scenes can be obtained.
[0114] S503, based on multiple speed reference vectors, perform speed-sensitive convolution on the bird's-eye view features at the current moment and the bird's-eye view features at multiple historical moments to obtain speed-separated bird's-eye view features, wherein the speed-separated bird's-eye view features y k ∈R W×H×1 It is expressed as follows:
[0115]
[0116] Among them, i∈{1,...W},j∈{1,...,H},c∈{1,...,256}, i, j and c represent the index of the bird's-eye view feature in width, height and channel dimensions respectively, and w∈R 3×3×256 is the parameter of the speed-sensitive convolution kernel, Δp∈{(-1,-1),(-1,0),...,(0,1),(1,1)}, α T ∈R is the decay factor related to the timestamp interval, Δp k,T ∈R 2 is the spatial motion offset vector in the bird's-eye view feature. The spatial motion offset vector is calculated by the reference speed, timestamp interval and ego-vehicle pose transformation matrix. T The bird's-eye view features of the current moment and the bird's-eye view features of multiple historical moments. The speed-separated bird's-eye view features can be understood as multi-channel bird's-eye view features formed by aggregating features of multiple frames of time-series bird's-eye view using different target speed priors. Figure 6 The bird's-eye view of the time series feature aggregation diagram shown in f t is the bird's-eye view feature at the current moment, f t-1 、f t-2 …f t-N+1 The MLP represents the bird's-eye view features of multiple historical moments. In this embodiment, the speed reference vector is used as motion compensation, and the bird's-eye view features of the current moment are aggregated with the bird's-eye view features of the historical moments to obtain the three-dimensional spatial motion offset of the target object at each historical moment, which can effectively alleviate the problem of moving target smear and improve the system perception performance.
[0117] S504, using a multi-layer perceptron to perform dimensionality transformation on the speed-separated bird's-eye view features to obtain speed-separated bird's-eye view features with the same dimension as the bird's-eye view features at the current moment.
[0118] S505: splice the bird's-eye view features separated by speed in the same dimension with the bird's-eye view feature dimension at the current moment to obtain a time series fusion feature.
[0119] The dimensions of the bird's-eye view features separated by speed are superimposed with the bird's-eye view features at the current moment to obtain data of higher dimensions, thereby improving the effect of data fusion.
[0120] In the prior art, the complementary characteristics of sensors such as cameras and radars are used to fuse the information obtained by multiple sensors to detect and perceive the vehicle's driving environment, which can effectively make up for the shortcomings of single sensors in measurement accuracy and environmental adaptability, and provide accuracy and robustness of target detection. In this solution, the processing method for multimodal data collected by multiple sensors is relatively customized, not universal, and difficult to expand to more complex multi-sensor scenarios. The present application provides an embodiment, such as Figure 7 As shown, the first visual point cloud feature and the first radar point cloud feature are fused to generate a point cloud fusion feature, including:
[0121] S701. Map each point on the acquired feature map to a three-dimensional coordinate system to generate a first visual point cloud feature. Each point feature in the first visual point cloud feature is represented by information combined with the three-dimensional spatial coordinate value, timestamp and image feature of the corresponding point.
[0122] In this embodiment, each point feature in the first visual point cloud feature is represented by the information of the three-dimensional space coordinate value, timestamp and image feature of the corresponding point. As an optional implementation, the point feature of the first visual point cloud feature can be represented as p = (x, y, z, t, f s )∈R 4+d , where (x, y, z) represents the three-dimensional space coordinates, t represents the timestamp, which can be understood as the current moment, and f s ∈R d To represent the image feature attributes. As an example, the point feature of the first visual point cloud feature can be represented as (x, y, z, t, f CNN )∈R 260 , where f CNN ∈R 256It is represented as the image features after convolutional neural network. (x, y, z) represents the three-dimensional space coordinates, t represents the timestamp, and 260 and 256 represent the dimensions. This is just an example. You can set different dimensions according to the actual situation. The spatial coordinates, timestamp and image feature information are integrated to represent the visual point cloud features, which can improve the richness of image data.
[0123] S702, mapping the radar point cloud data to a three-dimensional coordinate system to generate radar point cloud features, wherein each point feature in the radar point cloud features is represented by information combined with the three-dimensional spatial coordinate value of the corresponding point, a timestamp, and a perception attribute of the radar.
[0124] In this embodiment, each point feature in the radar point cloud feature is represented by the information combined by the three-dimensional spatial coordinate value of the corresponding point, the timestamp, and the radar's perception attributes. Each frame of point cloud data acquired by the radar includes multiple radar reflection point data. A radar reflection point data may include the three-dimensional coordinate information of the radar reflection point, the timestamp information, and may also include data expressing the radar perception attributes, such as speed, the signal-to-noise ratio of the reflected signal, signal strength information, etc. Exemplarily, the point feature of the radar point cloud feature can be expressed as p = (x, y, z, t, f s )∈R 4+d , where (x, y, z) represents the three-dimensional space coordinates, t represents the timestamp, which can be understood as the current moment, and f s ∈R d represents the perception attribute of the radar.
[0125] In this embodiment, a unified form is used to represent visual point cloud features and radar point cloud features, so that the multimodal data collected by multiple sensors can not only be expressed in a unified spatial coordinate system, but also the feature data can be represented in a unified form, so that the representation of multimodal data has universality and uniformity, and can better adapt to more complex multi-sensor fusion scenarios, so that multimodal data can be better integrated and the accuracy of target detection can be improved.
[0126] S703: Fuse the first visual point cloud features with the radar point cloud features to obtain point cloud fusion features.
[0127] The visual point cloud features are fused with the radar point cloud features to improve the fusion effect. For example, the first visual point cloud features and the radar point cloud features can be dimensionally spliced. Dimensional splicing can be understood as superimposing the dimensions of the data, that is, one dimension of the two feature data is superimposed, and the other dimensions remain unchanged. For example, the dimension of the first visual point cloud feature is (K, n), and the dimension of the radar point cloud feature is (K, m). After dimension splicing, a point cloud fusion feature with a dimension of (K, n+m) is obtained.
[0128] In this embodiment, the features of the two-dimensional image are extracted through a neural network model, the extracted image features are represented by a point cloud, the image and radar point cloud data are converted into a unified form representation under a unified coordinate system, and the fused image and radar point cloud data are converted into a unified, accurate, and dense 3D scene expression from a bird's-eye view. This not only makes the representation of multimodal data universal and unified, and adaptable to more complex multi-sensor application scenarios, but also improves the accuracy of the detection results.
[0129] As an optional implementation, the radar is a laser radar, such as Figure 8 The method for mapping radar point cloud data to a three-dimensional coordinate system to generate radar point cloud features includes:
[0130] S801. Obtain laser point cloud data collected by a laser radar.
[0131] Laser radar usually emits electromagnetic waves periodically. When the electromagnetic waves encounter obstacles in the direction of travel, part of the electromagnetic waves will be reflected back in the opposite direction of the emission direction. The electromagnetic waves reflected back by the target objects in the detection area are the echo information received by the laser radar. After signal processing, the required laser point cloud data can be obtained.
[0132] S802, mapping the laser point cloud data to a three-dimensional coordinate system to generate a first laser point cloud feature, wherein each point feature in the first laser point cloud feature is represented by information combined with the three-dimensional spatial coordinate value, timestamp and signal strength of the corresponding point.
[0133] The laser radar reflection point data may include the three-dimensional coordinate information of the laser radar reflection point, and may also include data expressing the sensing attributes of the laser radar, such as the signal strength of the laser radar. The signal strength is expressed as the intensity of the echo signal reflected back after the radar signal emitted by the laser radar reaches the spatial position corresponding to the point cloud data. Exemplarily, according to the projection relationship defined by the vehicle's own coordinate system and the laser radar's external parameter matrix, the three-dimensional coordinates of the laser point cloud data are mapped to the vehicle's own coordinate system. Each point feature in the first laser point cloud feature is represented by information fused by the three-dimensional spatial coordinate value of the corresponding point, the timestamp, and the signal strength. Exemplarily, the point feature of the first laser point cloud feature can be expressed as (x, y, z, t, i)∈R 5 , where (x, y, z) represents the three-dimensional space coordinates, t represents the timestamp, which can be understood as the current moment, and i represents the signal strength of the laser radar.
[0134] As an optional implementation, the radar is a millimeter wave radar, such as Fig. 9 The method for mapping point cloud data to a three-dimensional coordinate system to generate radar point cloud features also includes:
[0135] S901. Acquire millimeter-wave point cloud data collected by a millimeter-wave radar.
[0136] In order to achieve target detection, multiple millimeter wave radars can be deployed on the vehicle, and the scanning range of the multiple millimeter wave radars can cover all around the vehicle body. For example, millimeter wave radars can be installed at the four corners of the vehicle, or millimeter wave radars can be installed at the front, back, left, and right of the vehicle. The millimeter wave point cloud data around the vehicle body can be obtained through the installed millimeter wave radars.
[0137] S902, mapping the millimeter wave point cloud data to a three-dimensional coordinate system to generate a first millimeter wave point cloud feature, wherein each point feature in the first millimeter wave point cloud feature is characterized by information combined with the three-dimensional spatial coordinate value, timestamp, speed, and signal-to-noise ratio of the corresponding point.
[0138] The millimeter-wave radar point cloud data may include the three-dimensional coordinate information of the radar reflection point, and may also include data expressing the perception attributes of the millimeter-wave radar, such as the millimeter-wave radar speed, signal-to-noise ratio, etc. Exemplarily, the three-dimensional coordinates of the millimeter-wave radar point cloud data are mapped to the vehicle's own coordinate system based on the projection relationship defined by the vehicle's own coordinate system and the millimeter-wave radar's external parameter matrix. Each point feature in the first millimeter-wave point cloud feature is represented by information fused by the three-dimensional spatial coordinate value, timestamp, speed, and signal-to-noise ratio of the corresponding point. Exemplarily, the point features in the first millimeter-wave point cloud feature can be expressed as (x, y, z, t, vx, vy, SNR)∈R 7 , where (x, y, z) represents the three-dimensional space coordinates, t represents the timestamp, which can be understood as the current moment, and v x and v y They represent the lateral and longitudinal speeds of the millimeter-wave radar respectively, and SNR (Signal to Noise Ratio) represents the signal-to-noise ratio.
[0139] In this embodiment, a unified form is used to characterize the first visual point cloud feature, the first laser point cloud feature and the first millimeter wave point cloud feature. The feature data can be represented in a unified form, so that the representation of multimodal data has universality and uniformity, and can better adapt to more complex multi-sensor fusion scenarios.
[0140] As an optional implementation, Fig.10 The first visual point cloud feature is fused with the radar point cloud feature to obtain the point cloud fusion feature, and further includes:
[0141] S1001, using a multi-layer perceptron to map each point feature in the first visual point cloud feature to obtain an image feature vector of each point, and fusing the image feature vector, three-dimensional space coordinate value and timestamp information of each point to generate a second visual point cloud feature;
[0142] S1002, using a multi-layer perceptron to map each point feature in the first laser point cloud feature to obtain a laser radar feature vector for each laser radar point, and fusing the laser radar feature vector, three-dimensional space coordinate value and timestamp information of each laser radar point to generate a second laser point cloud feature;
[0143] S1003, using a multi-layer perceptron to map each point feature in the first millimeter-wave point cloud feature to obtain a millimeter-wave radar feature vector for each millimeter-wave radar point, and fusing the millimeter-wave radar feature vector, three-dimensional space coordinate value, and timestamp information of each millimeter-wave radar point to generate a second millimeter-wave point cloud feature, wherein the image feature vector, the laser radar feature vector, and the millimeter-wave radar feature vector have the same dimension;
[0144] S1004, dimensionalally splicing the second visual point cloud feature, the second laser point cloud feature, and the second millimeter wave point cloud feature to obtain a point cloud fusion feature.
[0145] A multilayer perceptron is a feedforward artificial neural network model that maps input point cloud features of multiple different dimensions to point cloud features of a unified dimension. The specific implementation of the multilayer perceptron is not described in detail in the embodiments of the present application. A multilayer perceptron is used to map the first visual point cloud feature, the first laser point cloud feature, and the first millimeter wave point cloud feature, respectively, and image feature vectors, laser radar feature vectors, and millimeter wave radar feature vectors are obtained respectively, and these feature vectors have the same dimension. The image feature vector, three-dimensional spatial coordinate value, and timestamp information of each point are combined to generate a second visual point cloud feature, the laser radar feature vector, three-dimensional spatial coordinate value, and timestamp information of each laser radar point are combined to generate a second laser point cloud feature, and the millimeter wave radar feature vector, three-dimensional spatial coordinate value, and timestamp information of each millimeter wave radar point are combined to generate a second millimeter wave point cloud feature, so the second visual point cloud feature, the second laser point cloud feature, and the second millimeter wave point cloud feature have the same dimension. As an example, the point features of these point cloud features can be uniformly represented as (x, y, z, t, f MLP )∈R (4+q) , f MLP is the feature vector generated by the point cloud feature through the multi-layer perceptron, and q represents the feature dimension after passing through the multi-layer perceptron. In this embodiment, not only a unified form is used to represent the visual point cloud features, laser point cloud features, and millimeter wave point cloud features, but also the dimensions of these point cloud features are the same, so that the representation of multimodal data is more universal and unified, and can better adapt to more complex multi-sensor fusion scenarios.
[0146] The second visual point cloud features, the second laser point cloud features and the second millimeter wave point cloud features are dimensionally spliced to obtain point cloud fusion features to obtain higher dimensional data.
[0147] As an optional implementation, Fig.11 The first visual point cloud feature is fused with the radar point cloud feature to obtain the point cloud fusion feature, and further includes:
[0148] S1101, using a multi-layer perceptron to map the image features of each image feature point to obtain an image feature vector of each point, and combining the image feature vector, three-dimensional space coordinate value and timestamp information of each point to generate a third visual point cloud feature;
[0149] S1102, using a multi-layer perceptron to map the signal strength of each laser radar point to obtain a laser radar feature vector of each laser radar point, and combining the laser radar feature vector, three-dimensional space coordinate value and timestamp information of each laser radar point to generate a third laser point cloud feature;
[0150] S1103, using a multi-layer perceptron to map the speed and signal-to-noise ratio of each millimeter-wave radar point to obtain a millimeter-wave radar feature vector of each millimeter-wave radar point, and combining the millimeter-wave radar feature vector, three-dimensional space coordinate value, and timestamp information of each millimeter-wave radar point to generate a third millimeter-wave point cloud feature, wherein the image feature vector, the laser radar feature vector, and the millimeter-wave radar feature vector have the same dimension;
[0151] S1104, dimensionalally splicing the third visual point cloud feature, the third laser point cloud feature, and the third millimeter wave point cloud feature to obtain a point cloud fusion feature.
[0152] This embodiment uses a multi-layer perceptron to map the image features, the signal strength of the laser radar, and the speed and signal-to-noise ratio of the millimeter-wave radar, respectively, to obtain the corresponding image feature vector, the laser radar feature vector, and the millimeter-wave radar feature vector. The specific implementation method is similar to the above embodiment and will not be repeated here.
[0153] The present application also provides a target detection method, comprising:
[0154] Get the features to be detected.
[0155] The features to be detected can be Figure 4 The time series fusion features obtained by the embodiment.
[0156] Target detection is performed based on the features to be detected.
[0157] In one implementation, the feature to be detected can be input into a pre-trained target prediction model to obtain a target prediction result output by the target prediction model. The specific structure and training process of the target detection model are not limited.
[0158] like Fig.12 The specific embodiment of the target detection method of the present application shown in the figure extracts features of the acquired image through a convolutional neural network, transforms the view angle of the acquired image features, maps them to a preset three-dimensional coordinate system, generates a first visual point cloud feature, and the point features (x, y, z, t, f) of the first visual point cloud feature are CNN ) is characterized by the information combined by the three-dimensional spatial coordinate value, timestamp and image feature of the corresponding point. The acquired laser point cloud data is mapped to a three-dimensional coordinate system to generate a first laser point cloud feature, and the information combined by the three-dimensional spatial coordinate value, timestamp and laser radar signal strength is used to characterize the point feature (x, y, z, t, i) of the first laser point cloud feature. The acquired millimeter-wave point cloud data is mapped to a three-dimensional coordinate system to generate a first millimeter-wave point cloud feature, and the three-dimensional spatial coordinate value, timestamp, speed of the millimeter-wave radar and signal-to-noise ratio of the millimeter-wave point cloud are used to characterize the point feature (x, y, z, t, vx, vy, SNR) of the first millimeter-wave point cloud feature. The first visual point cloud feature, the first laser point cloud feature and the first millimeter-wave point cloud feature are represented in a unified form, so that the representation of multimodal data is universal and unified. The first visual point cloud feature, the first laser point cloud feature and the first millimeter wave point cloud feature are passed through the point cloud feature encoder. The function performed by the point cloud feature encoder is as described above: a multi-layer perceptron is used to perform feature mapping on the first visual point cloud feature, the first laser point cloud feature and the first millimeter wave point cloud feature respectively, and the corresponding feature vectors are obtained and spliced with the corresponding three-dimensional space coordinate values and timestamp dimensions respectively to obtain the second visual point cloud feature, the second laser point cloud feature and the second millimeter wave point cloud feature of the same dimension. After these features are fused, the point cloud fusion feature is obtained, and the point cloud fusion feature is extracted to generate the bird's-eye view feature f at the current moment. t . The bird's-eye view feature f at the current moment t , with bird's-eye view features of multiple historical moments t-1 、f t-2 …f t-N+1 Time series features are aggregated to obtain time series fusion features, and 3D target detection is performed based on the time series fusion features. The time series fusion features carry multimodal data of images, lidar, and millimeter wave radar, which can not only effectively alleviate the problem of moving target smearing, but also further improve the accuracy of detection results.
[0159] like Fig.13 As shown, the present application also provides a data fusion device, the device comprising:
[0160] An acquisition module 1301 is used to acquire images acquired by at least one image acquisition device and point cloud data acquired by at least one radar in the area to be measured;
[0161] The first fusion module 1302 is used to extract features from the image data, map the acquired image feature map to a preset three-dimensional coordinate system to generate a first visual point cloud feature, map the radar point cloud data to a three-dimensional coordinate system to generate a radar point cloud feature, and fuse the first visual point cloud feature with the radar point cloud feature to generate a point cloud fusion feature;
[0162] A bird's-eye view feature module 1303 is used to extract the point cloud fusion features and generate the bird's-eye view features at the current moment;
[0163] A time series feature aggregation module 1304 is used to aggregate the bird's-eye view features at the current moment with the bird's-eye view features at multiple historical moments based on multiple preset speed reference vectors to obtain a speed-separated bird's-eye view feature;
[0164] The second fusion module 1305 is used to fuse the speed-separated bird's-eye view features with the bird's-eye view features at the current moment to obtain a time series fusion feature.
[0165] In an optional implementation, the time series feature aggregation module 1304 is specifically used for:
[0166] A neural network is used to train the model on the target speed prior distribution data of multiple scenes to obtain multiple speed query vectors;
[0167] A multi-layer perceptron is used to transform the dimensions of multiple speed query vectors to obtain multiple speed reference vectors;
[0168] Based on multiple speed reference vectors, the bird's-eye view features at the current moment are convolved with the bird's-eye view features at multiple historical moments to obtain speed-separated bird's-eye view features;
[0169] In an optional implementation, the second fusion module 1305 is specifically configured to:
[0170] A multi-layer perceptron is used to transform the dimension of the speed-separated bird's-eye view features to obtain speed-separated bird's-eye view features with the same dimension as the bird's-eye view features at the current moment;
[0171] The bird's-eye view features separated by speed in the same dimension are concatenated with the bird's-eye view feature dimension at the current moment to obtain the time series fusion features.
[0172] An optional implementation, such as Fig.14 As shown, the first fusion module 1302 specifically includes:
[0173] A visual point cloud feature generating unit 13021 is used to map each point on the acquired feature map to the three-dimensional coordinate system to generate a first visual point cloud feature, wherein each point feature in the first visual point cloud feature is represented by information combined with the three-dimensional spatial coordinate value, timestamp and image feature of the corresponding point;
[0174] The radar point cloud feature generating unit 13022 is used to map the radar point cloud data to the three-dimensional coordinate system to generate radar point cloud features, wherein each point feature in the radar point cloud feature is represented by information combined with the three-dimensional spatial coordinate value of the corresponding point, a timestamp, and the perception attributes of the radar.
[0175] In an optional implementation, the radar point cloud feature generation unit 13022 is specifically used for:
[0176] The radar is a laser radar, and laser point cloud data collected by the laser radar is obtained;
[0177] The laser point cloud data is mapped to a three-dimensional coordinate system to generate a first laser point cloud feature, wherein each point feature in the first laser point cloud feature is represented by information combined with the three-dimensional spatial coordinate value, timestamp and signal strength of the corresponding point.
[0178] In an optional implementation, the radar point cloud feature generating unit 13022 is further configured to:
[0179] The radar is a millimeter wave radar, and millimeter wave point cloud data collected by the millimeter wave radar is obtained;
[0180] The millimeter wave point cloud data is mapped to a three-dimensional coordinate system to generate a first millimeter wave point cloud feature, and each point feature in the first millimeter wave point cloud feature is characterized by information combined with the three-dimensional spatial coordinate value, timestamp, speed and signal-to-noise ratio of the corresponding point.
[0181] In an optional implementation, the first fusion module 1302 is specifically configured to:
[0182] A multi-layer perceptron is used to map each point feature in the first visual point cloud feature to obtain an image feature vector of each point, and the image feature vector, three-dimensional space coordinate value and timestamp information of each point are combined to generate a second visual point cloud feature;
[0183] A multi-layer perceptron is used to map each point feature in the first laser point cloud feature to obtain a laser radar feature vector of each laser radar point, and the laser radar feature vector, three-dimensional space coordinate value and timestamp information of each laser radar point are combined to generate a second laser point cloud feature;
[0184] A multi-layer perceptron is used to map each point feature in the first millimeter-wave point cloud feature to obtain a millimeter-wave radar feature vector for each millimeter-wave radar point, and the millimeter-wave radar feature vector, three-dimensional space coordinate value and timestamp information of each millimeter-wave radar point are combined to generate a second millimeter-wave point cloud feature, wherein the image feature vector, the laser radar feature vector and the millimeter-wave radar feature vector have the same dimension;
[0185] After the second visual point cloud feature, the second laser point cloud feature and the second millimeter wave point cloud feature are dimensional-joined, a point cloud fusion feature is obtained.
[0186] Fig.15 It is a schematic diagram of the hardware structure of a computer device provided in an embodiment of the present application. Fig.15 The computer device shown includes: a processor 1501, a communication interface 1502, a memory 1503 and a communication bus 1504. The processor 1501, the communication interface 1502 and the memory 1503 communicate with each other via the communication bus 1504. Fig.15 The connection method between the processor 1501, the communication interface 1502, and the memory 1503 shown is merely exemplary. During implementation, the processor 1501, the communication interface 1502, and the memory 1503 may also be connected to each other in communication with each other using other connection methods besides the communication bus 1504.
[0187] The memory 1503 can be used to store a computer program 15031, which may include instructions and data to implement the steps of any of the above data fusion methods and target detection methods. In an embodiment of the present application, the memory 1503 may be various types of storage media, such as random access memory (RAM), read only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical storage, and registers. The memory 1503 may include a hard disk and / or memory.
[0188] Processor 1501 may be a general-purpose processor, which may be a processor that performs specific steps and / or operations by reading and executing a computer program (e.g., computer program 15031) stored in a memory (e.g., memory 1503). The general-purpose processor may use data stored in the memory (e.g., memory 1503) in the process of performing the steps and / or operations. A general-purpose processor may be, for example, but not limited to, a central processing unit (CPU). In addition, processor 1501 may also be a special-purpose processor, which may be a processor specially designed to perform specific steps and / or operations. A special-purpose processor may be, for example, but not limited to, an ASIC and an FPGA. In addition, processor 1501 may also be a combination of multiple processors, such as a multi-core processor.
[0189] The communication interface 1502 may include an input / output (I / O) interface, a physical interface, and a logical interface for interconnecting devices within the network device, as well as an interface for interconnecting the network device with other devices (such as network devices). The communication network may be Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 1502 may be a module, a circuit, a transceiver, or any device capable of implementing communication.
[0190] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 1501 or an instruction in the form of software. The method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in the processor for execution. The software module can be located in a mature storage medium in the art such as a random access memory flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1503, and the processor 1501 reads the information in the memory 1503 and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.
[0191] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, any one of the above-mentioned data fusion methods and target detection methods is implemented.
[0192] Although the preferred embodiments of the present application have been disclosed for illustrative purposes, those skilled in the art will appreciate that various modifications, additions and substitutions are possible, without departing from the scope and spirit of the application as disclosed in the accompanying claims.
Claims
1. A data fusion method, characterized in that: The method comprises: Acquire images captured by at least one image acquisition device and radar point cloud data collected by at least one radar in the test area; Extracting features from the image data, mapping the acquired feature map to a preset three-dimensional coordinate system to generate a first visual point cloud feature, mapping the radar point cloud data to the three-dimensional coordinate system to generate a radar point cloud feature, and fusing the first visual point cloud feature with the radar point cloud feature to generate a point cloud fusion feature; Extracting features from the point cloud fusion features to generate bird's-eye view features at the current moment; Based on a plurality of preset speed reference vectors, feature aggregation is performed on the bird's-eye view features at the current moment and the bird's-eye view features at a plurality of historical moments to obtain a speed-separated bird's-eye view feature; Performing feature fusion on the bird's-eye view features separated by speed and the bird's-eye view features at the current moment to obtain a time series fusion feature; Target detection is performed based on the temporal fusion features.
2. The data fusion method according to claim 1, characterized in that: The step of performing feature aggregation on the bird's-eye view features at the current moment and the bird's-eye view features at multiple historical moments based on multiple preset speed reference vectors to obtain the speed-separated bird's-eye view features includes: A neural network is used to train the model on the target speed prior distribution data of multiple scenes to obtain multiple speed query vectors; Using a multi-layer perceptron to perform dimension transformation on the multiple speed query vectors to obtain the multiple speed reference vectors; Based on the multiple speed reference vectors, performing speed-sensitive convolution on the bird's-eye view features at the current moment and the bird's-eye view features at multiple historical moments to obtain the speed-separated bird's-eye view features; The step of fusing the speed-separated bird's-eye view features with the current moment's bird's-eye view features to obtain a time series fusion feature further includes: A multi-layer perceptron is used to perform dimension transformation on the speed-separated bird's-eye view features to obtain speed-separated bird's-eye view features with the same dimension as the bird's-eye view features at the current moment; The speed-separated bird's-eye view features of the same dimension and the bird's-eye view feature dimension of the current moment are concatenated to obtain the time series fusion features.
3. The data fusion method according to claim 1, characterized in that: The bird's-eye view feature of the velocity separation y k ∈R W ×H×1 It is expressed as follows: Among them, i∈{1,...W},j∈{1,...,H},c∈{1,...,256}, i, j and c represent the index of the bird's-eye view feature in width, height and channel dimensions respectively, and w∈R 3×3×256 is the parameter of the speed-sensitive convolution kernel, Δp∈{(-1,-1),(-1,0),...,(0,1),(1,1)}, α T ∈R is the decay factor related to the timestamp interval, Δp k,T ∈R 2 is the spatial motion offset vector in the bird's-eye view feature. The spatial motion offset vector is calculated by the reference speed, timestamp interval and ego-vehicle pose transformation matrix. T It is the bird's-eye view features of the current moment and the bird's-eye view features of multiple historical moments.
4. The data fusion method according to claim 1, characterized in that: The step of mapping the acquired feature map to a preset three-dimensional coordinate system to generate a first visual point cloud feature, mapping the radar point cloud data to the three-dimensional coordinate system to generate a radar point cloud feature, and fusing the first visual point cloud feature with the radar point cloud feature to generate a point cloud fusion feature includes: Mapping each point on the acquired feature map to the three-dimensional coordinate system to generate a first visual point cloud feature, wherein each point feature in the first visual point cloud feature is represented by information combined with the three-dimensional spatial coordinate value, timestamp and image feature of the corresponding point; Mapping the radar point cloud data to the three-dimensional coordinate system to generate radar point cloud features, wherein each point feature in the radar point cloud features is represented by information combined with the three-dimensional spatial coordinate value of the corresponding point, a timestamp, and a perception attribute of the radar; The first visual point cloud feature is fused with the radar point cloud feature to obtain a point cloud fusion feature.
5. The data fusion method according to claim 4, characterized in that: Mapping the radar point cloud data to the three-dimensional coordinate system to generate radar point cloud features includes: The radar is a laser radar, and laser point cloud data collected by the laser radar is obtained; Mapping the laser point cloud data to the three-dimensional coordinate system to generate a first laser point cloud feature, wherein each point feature in the first laser point cloud feature is represented by information combined with a three-dimensional spatial coordinate value, a timestamp, and a signal strength of the corresponding point; and / or, The radar is a millimeter wave radar, and millimeter wave point cloud data collected by the millimeter wave radar is obtained; The millimeter wave point cloud data is mapped to the three-dimensional coordinate system to generate a first millimeter wave point cloud feature, wherein each point feature in the first millimeter wave point cloud feature is characterized by information combined with the three-dimensional spatial coordinate value, timestamp, speed and signal-to-noise ratio of the corresponding point.
6. The data fusion method according to claim 5, characterized in that: The step of fusing the first visual point cloud feature with the radar point cloud feature to obtain a point cloud fusion feature includes: A multi-layer perceptron is used to map each point feature in the first visual point cloud feature to obtain an image feature vector of each point, and the image feature vector, three-dimensional space coordinate value and timestamp information of each point are combined to generate a second visual point cloud feature; A multi-layer perceptron is used to map each point feature in the first laser point cloud feature to obtain a laser radar feature vector of each laser radar point, and the laser radar feature vector, three-dimensional space coordinate value and timestamp information of each laser radar point are combined to generate a second laser point cloud feature; A multi-layer perceptron is used to map each point feature in the first millimeter-wave point cloud feature to obtain a millimeter-wave radar feature vector of each millimeter-wave radar point, and the millimeter-wave radar feature vector, three-dimensional space coordinate value and timestamp information of each millimeter-wave radar point are combined to generate a second millimeter-wave point cloud feature, wherein the image feature vector, the laser radar feature vector and the millimeter-wave radar feature vector have the same dimension; The point cloud fusion feature is obtained by dimensional splicing of the second visual point cloud feature, the second laser point cloud feature and the second millimeter wave point cloud feature.
7. A data fusion device, characterized in that: The device comprises: An acquisition module, used to acquire images acquired by at least one image acquisition device and point cloud data acquired by at least one radar in the test area; A first fusion module is used to extract features from the image data, map the acquired feature map to a preset three-dimensional coordinate system to generate a first visual point cloud feature, map the radar point cloud data to the three-dimensional coordinate system to generate a radar point cloud feature, and fuse the first visual point cloud feature with the radar point cloud feature to generate a point cloud fusion feature; A bird's-eye view feature module is used to extract features from the point cloud fusion features to generate bird's-eye view features at the current moment; A time series feature aggregation module is used to perform feature aggregation on the bird's-eye view features at the current moment and the bird's-eye view features at multiple historical moments based on multiple preset speed reference vectors to obtain a speed-separated bird's-eye view feature; The second fusion module is used to fuse the speed-separated bird's-eye view features with the bird's-eye view features at the current moment to obtain a time series fusion feature.
8. The data fusion device according to claim 7, characterized in that: The time series feature aggregation module is specifically used for: A neural network is used to train the model on the target speed prior distribution data of multiple scenes to obtain multiple speed query vectors; Using a multi-layer perceptron to perform dimension transformation on the multiple speed query vectors to obtain the multiple speed reference vectors; Based on the multiple speed reference vectors, performing speed-sensitive convolution on the bird's-eye view features at the current moment and the bird's-eye view features at multiple historical moments to obtain the speed-separated bird's-eye view features; and / or, The second fusion module is specifically used for: A multi-layer perceptron is used to perform dimension transformation on the speed-separated bird's-eye view features to obtain speed-separated bird's-eye view features with the same dimension as the bird's-eye view features at the current moment; The speed-separated bird's-eye view features of the same dimension are concatenated with the bird's-eye view feature dimension of the current moment to obtain the time series fusion feature; and / or The first fusion module includes: A visual point cloud feature generating unit, used for mapping each point on the acquired feature map to the three-dimensional coordinate system to generate a first visual point cloud feature, wherein each point feature in the first visual point cloud feature is represented by information combined with a three-dimensional spatial coordinate value, a timestamp and an image feature of the corresponding point; A radar point cloud feature generating unit, used for mapping the radar point cloud data to the three-dimensional coordinate system to generate radar point cloud features, wherein each point feature in the radar point cloud features is represented by information combined with the three-dimensional spatial coordinate value of the corresponding point, a timestamp, and a perception attribute of the radar; and / or, The radar point cloud feature generation unit is specifically used for: The radar is a laser radar, and laser point cloud data collected by the laser radar is obtained; Mapping the laser point cloud data to the three-dimensional coordinate system to generate a first laser point cloud feature, wherein each point feature in the first laser point cloud feature is represented by information combined with a three-dimensional spatial coordinate value, a timestamp, and a signal strength of the corresponding point; and / or, The radar point cloud feature generation unit is specifically used for: The radar is a millimeter wave radar, and millimeter wave point cloud data collected by the millimeter wave radar is obtained; The millimeter wave point cloud data is mapped to the three-dimensional coordinate system to generate a first millimeter wave point cloud feature, wherein each point feature in the first millimeter wave point cloud feature is characterized by information combined with the three-dimensional spatial coordinate value, timestamp, speed and signal-to-noise ratio of the corresponding point.
9. A computer device, characterized in that: include: A processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the steps of any one of the methods described in claims 1-6 when executing the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Intelligent automobile operation risk field prediction method and device
CN120299007A
Feature fusion method and device, environment perception method and device, equipment, medium and product
CN120852941A
Aerial view data frame determination method, apparatus and device, and medium
CN121095906A