Cross-modal pose estimation method, device, electronic device and storage medium

By obtaining a bird's-eye view of the target object's point cloud and image data, performing multimodal feature extraction and cross-modal feature association, and using convolutional neural networks and pose regression networks to calculate the relevant volume, the problems of insufficient accuracy and robustness of pose estimation in existing technologies are solved, and high-precision cross-modal pose estimation is achieved.

CN117953058BActive Publication Date: 2025-10-10SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410123069.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-10-10
Estimated Expiration
2044-01-29

AI Technical Summary

Technical Problem

In the existing technology, the pose estimation method based on the bird's-eye view has the problems of missing key features and cumulative errors, resulting in insufficient positioning accuracy and robustness, especially in the inability to meet the high-precision navigation requirements in intelligent connected vehicles.

Method used

By obtaining a bird's-eye view of the target object's point cloud data and image data, multimodal feature extraction and cross-modal feature association are performed, and the relevant volume is calculated using a convolutional neural network and a pose regression network to achieve pose estimation. The hyperbolic tangent layer and the fully connected layer are used for prediction.

Benefits of technology

The accuracy and robustness of pose estimation are improved, the problem of feature misalignment caused by inconsistency in multimodal data is overcome, and high-precision cross-modal pose estimation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117953058B_ABST
    Figure CN117953058B_ABST
Patent Text Reader

Abstract

The application discloses a cross-modal pose estimation method and device, electronic equipment and a storage medium. The method comprises the following steps: obtaining a first bird's-eye view of point cloud data and a second bird's-eye view of image data of a target object in the same frame; performing first feature extraction according to the first bird's-eye view to obtain point cloud bird's-eye features; performing second feature extraction according to the second bird's-eye view to obtain image bird's-eye features; performing cross-modal feature association based on the point cloud bird's-eye features and the image bird's-eye features to obtain a relevant volume of the target object in an optical flow; inputting the relevant volume into a pose regression network to perform pose estimation and obtain pose information of the target object; wherein the pose regression network comprises a hyperbolic tangent layer and a plurality of fully connected layers, and the hyperbolic tangent layer serves as an output layer. The application can accurately perform cross-modal pose estimation through multi-modal feature extraction, cross-modal feature association and pose estimation, and can be widely applied to the technical field of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a cross-modal pose estimation method, device, electronic device and storage medium. Background Art

[0002] Pose estimation provides information about an agent's own position and is a key step in environmental perception for connected vehicles. Given that global navigation satellite systems cannot always meet the high-precision requirements for navigation, ego-vehicle localization methods leverage sensors embedded within the vehicle to improve positioning accuracy and robustness.

[0003] E-vehicle perception devices, including cameras and LiDAR, can accurately acquire three-dimensional position information and measure the real-world environment, enabling multimodal perception of traffic scenarios. Methods based on a bird's-eye view (BEV) can eliminate the problem of inconsistent multimodal data and provide a more comprehensive understanding of the interaction between the ego vehicle and its surroundings. BEV features have demonstrated promising results in various tasks, including 3D object detection, semantic segmentation, and trajectory planning, and also provide new solutions for positioning tasks.

[0004] Conventional mapping and localization methods project lidar point clouds onto image coordinates to create a unified representation, but this approach can result in the loss of key features. Furthermore, the accumulated errors from not tracking vehicle motion can cause a mismatch between sensor data and map information, leading to localization errors that negatively impact the robustness and accuracy of the algorithm.

[0005] Application Contents

[0006] The main purpose of the embodiments of the present application is to propose a cross-modal pose estimation method, device, electronic device and storage medium that can accurately perform cross-modal pose estimation.

[0007] To achieve the above objectives, an embodiment of the present application provides a cross-modal pose estimation method, which includes:

[0008] Acquire a first bird's-eye view of the point cloud data and a second bird's-eye view of the image data of the same frame of the target object;

[0009] Performing a first feature extraction based on the first bird's-eye view to obtain a point cloud bird's-eye view feature;

[0010] Performing a second feature extraction based on the second bird's-eye view image to obtain a bird's-eye view feature of the image;

[0011] Based on the bird's-eye view features of point clouds and images, cross-modal feature association is performed to obtain the relevant volume of the target object in the optical flow;

[0012] The relevant volume is input into the pose regression network for pose estimation to obtain the pose information of the target object; wherein, the pose regression network includes a hyperbolic tangent layer and multiple fully connected layers, and the hyperbolic tangent layer serves as the output layer.

[0013] In some embodiments, performing first feature extraction based on the first bird's-eye view to obtain a point cloud bird's-eye view feature includes:

[0014] Input the first bird's-eye view image into a voxel extractor, and output a voxel representation of the point cloud data;

[0015] The voxel representation is input into the point cloud backbone network based on point pillars, and the point cloud bird's-eye view features are output.

[0016] In some embodiments, performing a second feature extraction based on the second bird's-eye view to obtain an image bird's-eye view feature includes:

[0017] Input the second bird's-eye view image into the image feature extractor, and output an image feature map;

[0018] Depth distribution prediction is performed based on the image feature map to obtain depth probability;

[0019] Generate a three-dimensional frustum point cloud of the target object based on the image data;

[0020] According to the depth probability, the image feature map is scaled by outer product in combination with the 3D frustum point cloud to generate a 3D feature point cloud;

[0021] The three-dimensional feature point cloud is flattened in the vertical direction to obtain the bird's-eye view features of the image.

[0022] In some embodiments, cross-modal feature association is performed based on the point cloud bird's-eye view features and the image bird's-eye view features to obtain the relevant volume of the target object in the optical flow, including:

[0023] The point cloud bird's-eye view features and image bird's-eye view features are input into the convolutional neural network respectively, and the corresponding outputs are point cloud features and image features;

[0024] According to the point cloud features and image features, the relevant volume is calculated.

[0025] In some embodiments, the point cloud bird's-eye view features and the image bird's-eye view features are respectively input into a convolutional neural network, and the corresponding outputs are point cloud features and image features, including:

[0026] The bird's-eye view features of the point cloud are input into the convolutional neural network for downsampling and information extraction, and the point cloud features are output;

[0027] The bird's-eye view features of the image are input into the convolutional neural network for downsampling and information extraction, and the image features are output;

[0028] Among them, the convolutional neural network includes three convolutional layers, and forms a pyramid structure through three convolutional layers.

[0029] In some embodiments, the relevant volume is calculated based on the point cloud features and the image features, including:

[0030] Based on the point cloud features and image features, the matching cost of the associated corresponding features is calculated and stored in the correlation volume; among them, the correlation volume is used to measure the correlation of cross-modal features.

[0031] In some embodiments, the relevant volume is input into a pose regression network to perform pose estimation to obtain pose information of the target object, including:

[0032] Based on the relevant volume, translation and rotation are predicted through multiple fully connected layers, and then the predicted pose information of the target object is output through the hyperbolic tangent layer.

[0033] To achieve the above objectives, another aspect of the present application provides a cross-modal pose estimation device, comprising:

[0034] The first module is used to obtain a first bird's-eye view of the point cloud data and a second bird's-eye view of the image data of the same frame of the target object;

[0035] The second module is used to extract the first feature according to the first bird's-eye view to obtain the point cloud bird's-eye view feature;

[0036] The third module is used to extract the second feature according to the second bird's-eye view image to obtain the bird's-eye view feature of the image;

[0037] The fourth module is used to perform cross-modal feature association based on the bird's-eye view features of the point cloud and the bird's-eye view features of the image to obtain the relevant volume of the target object in the optical flow;

[0038] The fifth module is used to input the relevant volume into the pose regression network for pose estimation to obtain the pose information of the target object; wherein, the pose regression network includes a hyperbolic tangent layer and multiple fully connected layers, and the hyperbolic tangent layer serves as the output layer.

[0039] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned method is implemented.

[0040] The embodiments of the present application include at least the following beneficial effects: The present application provides a cross-modal pose estimation method, device, electronic device and storage medium. The scheme obtains a first bird's-eye view of point cloud data and a second bird's-eye view of image data of the same frame of the target object; extracts a first feature based on the first bird's-eye view to obtain a point cloud bird's-eye feature; extracts a second feature based on the second bird's-eye view to obtain an image bird's-eye feature; performs cross-modal feature association based on the point cloud bird's-eye feature and the image bird's-eye feature to obtain a related volume of the target object in the optical flow; inputs the related volume into a pose regression network for pose estimation to obtain the pose information of the target object; wherein the pose regression network includes a hyperbolic tangent layer and multiple fully connected layers, and the hyperbolic tangent layer serves as the output layer. The embodiments of the present application first describe the construction of two modal features, representing point cloud data and real-time images respectively, through multimodal feature extraction, cross-modal feature association and pose estimation. Then, the related volumes of the two modalities are calculated based on the optical flow idea, and finally pose estimation is performed based on the related volumes. The embodiments of the present application can accurately perform cross-modal pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of a cross-modal pose estimation method provided in an embodiment of the present application;

[0042] Figure 2 Schematic diagram of the overall process of cross-modal pose estimation provided by the embodiment of the present application;

[0043] Figure 3 is a structural diagram of a cross-modal pose estimation device provided in an embodiment of the present application;

[0044] Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0046] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0047] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0049] The cross-modal pose estimation method provided in the embodiment of the present application relates to the field of data processing technology. The cross-modal pose estimation method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the cross-modal pose estimation method, etc., but is not limited to the above forms.

[0050] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0051] Figure 1 is an optional flowchart of the cross-modal pose estimation method provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S100 to S500.

[0052] S100, obtaining a first bird's-eye view of point cloud data and a second bird's-eye view of image data of a target object in the same frame;

[0053] For example, in some specific embodiments, step S100 is implemented as follows: obtaining a bird's-eye view of the radar point cloud; obtaining a bird's-eye view of the camera image.

[0054] S200, extracting a first feature based on the first bird's-eye view to obtain a point cloud bird's-eye view feature;

[0055] It should be noted that in some embodiments, step S200 may include: inputting the first bird's-eye view image into a voxel extractor to output a voxel representation of the point cloud data; inputting the voxel representation into a point cloud backbone network based on point pillars to output point cloud bird's-eye view features. It should be noted that a point cloud backbone network based on PointPillars is employed, using FC-64-64 and BatchNorm as its PointNet. Pillars are created for LiDAR points within the range x∈[-10m,70m] and y∈[-40m,40m], with each pillar representing a 0.25m×0.25m spatial region. A default 2D multi-scale feature CNN is used to obtain spatial features at 0.5 times the resolution of the original pillars. Specifically, for PointPillars, FC-64-64 and BatchNorm are used as its PointNet. Unlike the original PointPillars, which directly constructs dense pillars specified by hyperparameters, pillars are sparsely represented. A sparse PointNet is also used to process the corresponding sparse pillar features. This enables efficient processing of all columns in space and time.

[0056] For example, in some specific embodiments, for the laser radar point cloud map, {p i |i=1,2,…,n} represents; the current frame laser point cloud data can be input into the voxel extractor , get the voxel representation of the radar point cloud

[0057]

[0058] According to the obtained voxel representation, a point cloud backbone network based on point pillars is used Perform feature extraction and flatten along the height direction to create a bird's-eye view feature of the point cloud

[0059]

[0060] S300, performing a second feature extraction based on the second bird's-eye view image to obtain a bird's-eye view feature of the image;

[0061] It should be noted that, in some embodiments, step S300 may include: inputting the second bird's-eye view into the image feature extractor, and outputting an image feature map; performing depth distribution prediction based on the image feature map to obtain a depth probability; generating a three-dimensional frustum point cloud of the target object based on the image data; performing outer product scaling on the image feature map in combination with the three-dimensional frustum point cloud based on the depth probability to generate a three-dimensional feature point cloud; and flattening the three-dimensional feature point cloud along the vertical direction to obtain an image bird's-eye view feature.

[0062] For example, in some specific embodiments, the current frame camera image data may be Input image feature extractor , get the image feature map

[0063]

[0064] Where: C represents the number of feature map channels, H and W are the image width and height pixel values;

[0065] Using the Depth Estimation Module Obtain the depth distribution of the predicted features and obtain the depth probability (distribution); Based on the image data, use the camera internal parameters and external reference Generate a 3D frustum point cloud in the vehicle's own coordinate system; perform outer product scaling on the obtained image features based on the depth probability to generate a 3D feature point cloud

[0066]

[0067] According to the obtained three-dimensional feature point cloud, use the feature pooling module Flatten along the vertical direction to obtain the bird's-eye view features of the image

[0068]

[0069] S400, performing cross-modal feature association based on the point cloud bird's-eye view features and the image bird's-eye view features to obtain the relevant volume of the target object in the optical flow;

[0070] It should be noted that, in some embodiments, step S400 may include: inputting the point cloud bird's-eye view features and the image bird's-eye view features into the convolutional neural network respectively, and outputting the corresponding point cloud features and image features; and calculating the relevant volume based on the point cloud features and the image features.

[0071] In some embodiments, the point cloud bird's-eye view features and the image bird's-eye view features are respectively input into the convolutional neural network, and the corresponding outputs are point cloud features and image features, which can include: inputting the point cloud bird's-eye view features into the convolutional neural network for downsampling and information extraction, and outputting point cloud features; inputting the image bird's-eye view features into the convolutional neural network for downsampling and information extraction, and outputting image features; wherein the convolutional neural network includes three convolutional layers, and the three convolutional layers form a pyramid structure.

[0072] For example, in some specific embodiments, a convolutional pyramid module consisting of three convolutional layers can be used based on the obtained point cloud bird's-eye view features. Perform downsampling and information extraction to obtain point cloud features

[0073]

[0074] According to the obtained image bird's eye view features, the same convolution pyramid module is used to perform down-sampling and information extraction to obtain image features

[0075]

[0076] In some embodiments, according to the point cloud features and the image features, the related volume is calculated, which can include: calculating the matching cost of the related features and storing in the related volume according to the point cloud features and the image features; wherein the related volume is used to measure the correlation of the cross-modal features.

[0077] Exemplarily, in some specific embodiments, the matching cost of the related features can be calculated and stored in the related volume according to the obtained point cloud and image features, which is used to measure the correlation of the cross-modal features:

[0078]

[0079] wherein: is the related volume, represents a column vector whose length is c(·) represents the column operation, represents a feature map in which one feature pixel is represented by

[0080] For each feature pixel, only the correlation of the feature pixels within a certain range d around the same position in needs to be calculated; because after down-sampling, one pixel offset of the top layer corresponds to 2 L--1 pixels in the full-resolution bird's eye view image, the value range of d can be set to be small, and the dimension of the obtained related volume is:

[0081]

[0082] wherein: H 3th and W 3th are the height and width of the 3rd layer features, respectively;

[0083] S500, input the related volume into the pose regression network to perform pose estimation, and obtain the pose information of the target object.

[0084] It should be noted that in some embodiments, step S500 can include: based on the related volume, predicting the translation and rotation through multiple fully connected layers, and then outputting the predicted pose information of the target object through the hyperbolic tangent layer.

[0085] For example, in some specific embodiments, the six-degree-of-freedom camera pose can be predicted using several fully connected layers based on the correlation volume constructed in step S3.1; to ensure that the predicted translation and rotation fall within the appropriate range, an additional tanh layer (hyperbolic tangent layer) is added as the output layer:

[0086]

[0087] in: represents the translation along the xyz direction, Represents a quaternion array for rotation, Represents multiple fully connected layers used to predict poses.

[0088] In some preferred implementations, a series of convolutional layers and concat layers with a stride of 1 can be integrated into the pose regression network to implement the processing flow and specific implementation of steps S200 to S400. Specifically, based on the obtained bird's-eye view of the lidar point cloud and the bird's-eye view of the camera image, a series of convolutional layers and concat layers with a stride of 1 are used in the channel dimension to obtain a correlation volume that integrates multi-level information; the function of the concat layer is to splice two or more feature maps in the channel or num dimension.

[0089] It should be noted that, unless otherwise specified, if the data processing logic is clear, the embodiments of the present application do not limit the specific composition structure of the application modules and networks (such as feature pooling modules, etc.).

[0090] In order to explain the principles of the technical solution of this application in detail, the overall process of this application is described below in combination with some specific embodiments. It is easy to understand that the following is an explanation of the technical principles of this application and cannot be regarded as a limitation of this application.

[0091] In some specific embodiments, such as Figure 2 As shown, the embodiment of the present application realizes cross-modal pose estimation based on bird's-eye view features, and the specific process is as follows:

[0092] S1 bird's-eye view feature extraction: generating corresponding bird's-eye view features based on the point cloud data and image data of the current frame;

[0093] S2 cross-modal feature association uses the correlation volume in the optical flow to calculate the similarity between the lidar point cloud features and the camera image features;

[0094] S3 pose estimation inputs the cross-modal feature correlation volume into the pose regression network to obtain the six-degree-of-freedom pose of the camera in the map.

[0095] As a preferred embodiment, step S1 may include:

[0096] S1.1 Obtain a bird's-eye view of the radar point cloud;

[0097] S1.2 obtain a bird's-eye view of the camera image;

[0098] As a preferred embodiment, step S1.1 may include:

[0099] S1.1.1 For lidar point cloud maps, you can use {p i |i=1,2,…,n} represents;

[0100] S1.1.2 Input the current frame laser point cloud data into the voxel extractor , get the voxel representation of the radar point cloud

[0101]

[0102] S1.1.3 Based on the voxel representation obtained in step S1.1.2, use the point cloud backbone network based on the point column Perform feature extraction and flatten along the height direction to create a bird's-eye view feature of the point cloud

[0103]

[0104] As a preferred embodiment, step S1.2 may include:

[0105] S1.2.1 The current frame camera image data Input image feature extractor , get the image feature map

[0106]

[0107] Where: C represents the number of feature map channels, H and W are the image width and height pixel values;

[0108] S1.2.2 Using the Depth Estimation Module Obtain the depth distribution of the predicted features and obtain the depth probability (distribution); Based on the image data, use the camera internal parameters and external reference Generate a 3D frustum point cloud in the vehicle's own coordinate system; perform outer product scaling on the image features obtained in step S1.2.1 based on the depth probability to generate a 3D feature point cloud

[0109]

[0110] S1.2.3 Based on the three-dimensional feature point cloud obtained in step S1.2.2, use the feature pooling module Flatten along the vertical direction to obtain the bird's-eye view features of the image

[0111]

[0112] As a preferred embodiment, step S2 may include:

[0113] S2.1 Based on the bird's-eye view features obtained in steps S1.1.3 and S1.2.3, a convolutional neural network is used to perform downsampling and information fusion;

[0114] S2.2 calculates the correlation volume of the cross-modal features based on the point cloud and image features obtained in S2.1;

[0115] As a preferred embodiment, step S2.1 may include:

[0116] S2.1.1 Based on the bird’s-eye view features of the point cloud obtained in step S1.1.3, a convolutional pyramid module consisting of three convolutional layers is used. Perform downsampling and information extraction to obtain point cloud features

[0117]

[0118] S2.1.2 Based on the bird's-eye view features of the image obtained in step S1.2.3, use the same convolutional pyramid module as in S2.1.1 Perform downsampling and information extraction to obtain image features

[0119]

[0120] As a preferred embodiment, step S2.2 may include:

[0121] S2.2.1 Based on the point cloud and image features obtained in step S2.1, calculate the matching cost of the associated corresponding features and store them in the relevant volume to measure the cross-modal correlation:

[0122]

[0123] in: is the relevant volume, Represents a column vector The length of , c(·) represents the columnization operation, Representation feature map A feature pixel in ;

[0124] S2.2.2 For each feature pixel, it is only necessary to calculate its The correlation of feature pixels within a certain range d around the same position in the image; because after downsampling, a one-pixel shift in the top layer corresponds to 2 in the full-resolution bird's-eye view image. L-1 pixels, so the value range of d can be set to a smaller value, and the relevant volume dimension obtained is:

[0125]

[0126] Among them: H 3th and W 3th are the height and width of the third layer features respectively;

[0127] As a preferred embodiment, step S3 may include:

[0128] S3.1 Based on the bird's-eye view of the lidar point cloud and the camera image obtained in step S1, a series of convolutional layers and concat layers with a stride of 1 are applied to the channel dimension to obtain a correlation volume that fuses multi-level information. Taking the integration of convolutional layers and concat layers as an example, the feature extraction and cross-modal feature association mentioned above can be achieved.

[0129] S3.2 uses several fully connected layers to predict the 6DOF camera pose based on the correlation volume constructed in step S3.1. To ensure that the predicted translation and rotation fall within the appropriate range, an additional tanh layer (hyperbolic tangent layer) is added as the output layer:

[0130]

[0131] in: represents the translation along the xyz direction, Represents a quaternion array for rotation, Represents multiple fully connected layers used to predict poses.

[0132] In summary, compared with the prior art, the beneficial effects of the embodiments of the present application include at least:

[0133] 1. This application first describes the construction of two-modal BEV features, representing a live image and a pre-built LiDAR map, through multimodal BEV feature extraction, cross-modal feature association, and pose estimation. Then, inspired by optical flow, we compute the correlation volume between these two modalities. Based on this correlation volume, the network predicts the global ego-vehicle pose in the map.

[0134] 2. Use optical flow technology to align lidar point cloud features with camera image features. By matching in BEV space, the problem of feature misalignment caused by multimodal data differences is overcome.

[0135] See also Figure 3The present application also provides a cross-modal pose estimation device 500, which can implement the above-mentioned cross-modal pose estimation method. The device includes:

[0136] The first module 510 is configured to obtain a first bird's-eye view of the point cloud data and a second bird's-eye view of the image data of the same frame of the target object;

[0137] The second module 520 is configured to extract first features based on the first bird's-eye view image to obtain a point cloud bird's-eye view feature;

[0138] The third module 530 is configured to extract a second feature based on the second bird's-eye view image to obtain a bird's-eye view feature of the image;

[0139] A fourth module 540 is configured to perform cross-modal feature association based on the point cloud bird's-eye view features and the image bird's-eye view features to obtain a relevant volume of the target object in the optical flow;

[0140] The fifth module 550 is used to input the relevant volume into the pose regression network for pose estimation to obtain the pose information of the target object; wherein the pose regression network includes a hyperbolic tangent layer and multiple fully connected layers, and the hyperbolic tangent layer serves as the output layer.

[0141] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0142] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the cross-modal pose estimation method. The electronic device can be any smart terminal, such as a tablet computer or an in-vehicle computer.

[0143] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0144] See also Figure 4 , Figure 4 The hardware structure of an electronic device 600 according to another embodiment is shown. The electronic device includes:

[0145] The processor 601 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0146] The memory 602 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 602 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 602 and is called by the processor 601 to execute the cross-modal pose estimation method of the embodiments of this application.

[0147] Input / output interface 603, used to implement information input and output;

[0148] Communication interface 604, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0149] Bus 605 , which transmits information between various components of the device (e.g., processor 601 , memory 602 , input / output interface 603 , and communication interface 604 );

[0150] The processor 601 , the memory 602 , the input / output interface 603 and the communication interface 604 are connected to each other in communication within the device via a bus 605 .

[0151] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned cross-modal pose estimation method.

[0152] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0153] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0154] The cross-modal pose estimation method, cross-modal pose estimation device, electronic device and storage medium provided in the embodiment of the present application are as follows: obtaining a first bird's-eye view of the point cloud data and a second bird's-eye view of the image data of the same frame of the target object; performing a first feature extraction based on the first bird's-eye view to obtain a point cloud bird's-eye feature; performing a second feature extraction based on the second bird's-eye view to obtain an image bird's-eye feature; performing cross-modal feature association based on the point cloud bird's-eye feature and the image bird's-eye feature to obtain a relevant volume of the target object in the optical flow; inputting the relevant volume into a pose regression network to perform pose estimation to obtain the pose information of the target object; wherein the pose regression network includes a hyperbolic tangent layer and multiple fully connected layers, and the hyperbolic tangent layer serves as the output layer. The embodiment of the present application first describes the construction of two modal features, representing point cloud data and real-time images respectively, through multimodal feature extraction, cross-modal feature association and pose estimation. Then, the relevant volumes of the two modalities are calculated based on the optical flow idea, and finally pose estimation is performed based on the relevant volumes. The embodiment of the present application can accurately perform cross-modal pose estimation.

[0155] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0156] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0157] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0158] Those skilled in the art will appreciate that all or some of the steps, devices, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0159] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the numbers used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0160] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0161] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0162] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0163] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0164] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.

[0165] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A cross-modal pose estimation method, characterized in that: The method comprises: Acquire a first bird's-eye view of the point cloud data and a second bird's-eye view of the image data of the same frame of the target object; Performing a first feature extraction based on the first bird's-eye view to obtain a point cloud bird's-eye view feature; Performing a second feature extraction based on the second bird's-eye view to obtain an image bird's-eye view feature; The step of extracting a second feature based on the second bird's-eye view to obtain an image bird's-eye view feature includes: Inputting the second bird's-eye view image into an image feature extractor, and outputting an image feature map; Performing depth distribution prediction based on the image feature map to obtain a depth probability; generating a three-dimensional frustum point cloud of the target object based on the image data; Performing outer product scaling on the image feature map based on the depth probability and in combination with the three-dimensional frustum point cloud to generate a three-dimensional feature point cloud; Flattening the three-dimensional feature point cloud along the vertical direction to obtain the image bird's-eye view feature; Performing cross-modal feature correlation based on the point cloud bird's-eye view features and the image bird's-eye view features to obtain a related volume of the target object in the optical flow; The performing cross-modal feature association based on the point cloud bird's-eye view features and the image bird's-eye view features to obtain the relevant volume of the target object in the optical flow includes: Inputting the point cloud bird's-eye view features and the image bird's-eye view features into a convolutional neural network respectively, and outputting point cloud features and image features accordingly; Calculating the relevant volume according to the point cloud features and the image features; The step of calculating the relevant volume based on the point cloud features and the image features includes: Calculating matching costs of associated corresponding features based on the point cloud features and the image features and storing the costs in the correlation volume; wherein the correlation volume is used to measure the correlation of cross-modal features; The relevant volume is input into a pose regression network for pose estimation to obtain the pose information of the target object; wherein the pose regression network includes a hyperbolic tangent layer and multiple fully connected layers, and the hyperbolic tangent layer serves as an output layer.

2. The method according to claim 1, characterized in that The extracting a first feature according to the first bird's-eye view to obtain a point cloud bird's-eye view feature includes: Inputting the first bird's-eye view image into a voxel extractor, and outputting a voxel representation of the point cloud data; The voxel representation is input into a point cloud backbone network based on point pillars, and the point cloud bird's-eye view feature is output.

3. The method according to claim 1, characterized in that The step of inputting the point cloud bird's-eye view features and the image bird's-eye view features into a convolutional neural network and outputting corresponding point cloud features and image features includes: Inputting the point cloud bird's-eye view features into the convolutional neural network for downsampling and information extraction, and outputting the point cloud features; Inputting the bird's-eye view features of the image into the convolutional neural network for downsampling and information extraction, and outputting the image features; The convolutional neural network includes three convolutional layers, and a pyramid structure is formed by the three convolutional layers.

4. The method according to any one of claims 1 to 3, characterized in that Inputting the relevant volume into a pose regression network to perform pose estimation to obtain pose information of the target object includes: Based on the relevant volume, translation and rotation are predicted through multiple fully connected layers, and then the predicted position information of the target object is output through the hyperbolic tangent layer.

5. A cross-modal pose estimation device, characterized in that: The device comprises: The first module is used to obtain a first bird's-eye view of the point cloud data and a second bird's-eye view of the image data of the same frame of the target object; A second module is configured to extract a first feature based on the first bird's-eye view to obtain a point cloud bird's-eye view feature; A third module is configured to extract a second feature based on the second bird's-eye view image to obtain a bird's-eye view feature of the image; The step of extracting a second feature based on the second bird's-eye view to obtain an image bird's-eye view feature includes: Inputting the second bird's-eye view image into an image feature extractor, and outputting an image feature map; Performing depth distribution prediction based on the image feature map to obtain a depth probability; generating a three-dimensional frustum point cloud of the target object based on the image data; Performing outer product scaling on the image feature map based on the depth probability and in combination with the three-dimensional frustum point cloud to generate a three-dimensional feature point cloud; Flattening the three-dimensional feature point cloud along the vertical direction to obtain the image bird's-eye view feature; A fourth module is configured to perform cross-modal feature association based on the point cloud bird's-eye view features and the image bird's-eye view features to obtain a relevant volume of the target object in the optical flow; The performing cross-modal feature association based on the point cloud bird's-eye view features and the image bird's-eye view features to obtain the relevant volume of the target object in the optical flow includes: Inputting the point cloud bird's-eye view features and the image bird's-eye view features into a convolutional neural network respectively, and outputting point cloud features and image features accordingly; Calculating the relevant volume according to the point cloud features and the image features; The step of calculating the relevant volume based on the point cloud features and the image features includes: Calculating matching costs of associated corresponding features based on the point cloud features and the image features and storing the costs in the correlation volume; wherein the correlation volume is used to measure the correlation of cross-modal features; The fifth module is used to input the relevant volume into a pose regression network for pose estimation to obtain the pose information of the target object; wherein the pose regression network includes a hyperbolic tangent layer and multiple fully connected layers, and the hyperbolic tangent layer serves as the output layer.

6. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 4 when executing the computer program.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Modeling method based on oblique photography of unmanned aerial vehicle

    CN114359503A

  • Complex road target detection method based on multi-modal fusion aerial view

    CN117058646A