Method and device for simultaneously estimating dynamic and static poses in 4D space

By proposing a simultaneous estimation method of dynamic and static position in 4D space in autonomous driving technology, the problem of unsupervised detection box generation method in autonomous driving is solved, and more accurate perception and understanding capabilities are achieved to meet the application needs of autonomous driving.

CN120070836APending Publication Date: 2025-05-30TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411941470.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The unsupervised detection box generation method has problems in the autonomous driving technology, such as the inability to introduce semantic information, resulting in poor environmental perception and understanding capabilities, the lack of matching processing of static scenes leads to scene flow calculation deviations, and the low accuracy of the 3D detection box makes it difficult to meet the application needs of autonomous driving.

Method used

A simultaneous estimation method for dynamic and static poses in 4D space is proposed. By identifying and separating static points and dynamic points in point cloud data, inter-frame poses are calculated, and dynamic point cloud labels are generated to realize simultaneous estimation of dynamic and static poses in 4D space under unsupervised conditions.

Benefits of technology

It realizes the classification of dynamic and static points in point clouds, and uses the unsupervised static point position estimation network model and the unsupervised scene flow network model to calculate the matching relationship between the two frame point clouds, generate a 3D detection frame, improves the tracking accuracy of dynamic object motion information, and at the same time estimates the position information of static objects and dynamic objects, which helps the data closed-loop system of the perception function module in autonomous driving to more accurately understand and track the positions of different objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070836A_ABST
    Figure CN120070836A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving, in particular to a 4D space dynamic and static pose simultaneous estimation method and device, and the method comprises the steps: recognizing static points and dynamic points in two frames of point cloud data, separating the static points and the dynamic points from the two frames of point cloud data, and determining the two frames of static points and the two frames of dynamic points; calculating an inter-frame pose between the two frames of point cloud data according to the two frames of static points, and generating a dynamic point cloud label according to the two frames of dynamic points; and estimating the dynamic and static poses of the target object in the 4D space based on the inter-frame poses and the dynamic point cloud labels. According to the method, the dynamic and static points in the point cloud data can be separated, the unsupervised static point pose estimation network model and the unsupervised scene flow network model are constructed by using the dynamic and static points respectively, the matching relation of two frames of point clouds is calculated in an unsupervised mode, and then pose information of a static object and pose information of a dynamic object are estimated at the same time. And the automatic driving system is helped to more accurately understand and track the positions of different objects in the scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and particularly to a method and device for simultaneously estimating dynamic and static poses in a 4D space. Background Art

[0002] The Computing Brain Development System (CBDES) is the core of implementing autonomous driving technology. Since the module weights of the CBDES functional software are learned in a data-driven manner, it is necessary to build a data closed-loop system to achieve data collection, annotation, simulation, and functional module update. The point cloud data of lidar provides three-dimensional information of objects in the real world. How to generate accurate target detection boxes and label information from the point cloud data and more precisely understand and track the position of target objects in the scene is one of the key issues in realizing the data closed-loop system of the perception functional module in the autonomous driving perception model.

[0003] In related technologies, with the development of deep learning, the number of point cloud object detection methods based on deep learning has gradually increased. The vast majority of detection box generation algorithms still use supervised networks, mainly focusing on designing more powerful convolutional or feature extraction operators / modules to enhance the point cloud segmentation performance in a fully supervised setting, and these methods have also achieved good results. However, since the training of supervised networks relies on a large number of manual labels, and collecting and annotating large-scale data sets is very expensive and time-consuming, unsupervised detection box generation methods have emerged and achieved certain success.

[0004] However, there are many problems with the unsupervised detection box generation methods in related technologies. The lack of semantic information leads to poor environmental perception and understanding ability, or the lack of matching processing for static scenes results in large deviations in scene flow calculation, and the accuracy of 3D detection boxes is also relatively low, making it difficult to meet the application requirements of the data closed-loop system of the perception functional module and autonomous driving in actual scenarios, which urgently need to be solved. Summary of the Invention

[0005] The present application provides a method and device for simultaneously estimating dynamic and static poses in a 4D space to solve the problems in related technologies, such as the large number of problems with unsupervised detection box generation methods, the poor environmental perception and understanding ability due to the inability to introduce semantic information, or the large deviation in scene flow calculation caused by the lack of matching processing for static scenes, and the relatively low accuracy of 3D detection boxes, which are difficult to meet the application requirements of the data closed-loop system of the perception functional module and autonomous driving in actual scenarios.

[0006] The first aspect embodiment of this application provides a method for simultaneously estimating the dynamic and static poses in a 4D space, including the following steps: identifying static points and dynamic points in two frames of point cloud data to separate the static points and the dynamic points from the two frames of point cloud data, and determining two frames of static points and two frames of dynamic points; calculating the inter-frame pose between the two frames of point cloud data based on the two frames of static points, and generating a dynamic point cloud label based on the two frames of dynamic points; estimating the dynamic and static poses of the target object in the 4D space based on the inter-frame pose and the dynamic point cloud label.

[0007] Optionally, in an embodiment of this application, before calculating the inter-frame pose between the two frames of point cloud data, it further includes: constructing an unsupervised static point pose estimation network model for calculating the inter-frame pose between the two frames of point cloud data based on the two frames of static points.

[0008] Optionally, in an embodiment of this application, the generating the dynamic point cloud label based on the two frames of dynamic points includes: generating a first dynamic object detection box and a second dynamic object detection box based on the two frames of dynamic points; generating the dynamic point cloud label based on the first dynamic object detection box and the second dynamic object detection box.

[0009] Optionally, in an embodiment of this application, the generating the dynamic point cloud label based on the first dynamic object detection box and the second dynamic object detection box includes: optimizing the first dynamic object detection box and the second dynamic object detection box by using the two frames of dynamic points to obtain an optimized first dynamic object detection box and a second dynamic object detection box; constructing an unsupervised scene flow network model based on the optimized first dynamic object detection box and the second dynamic object detection box in combination with the unsupervised static point pose estimation network model; obtaining the scene flow information of the target object through the unsupervised scene flow network model, and determining the dynamic point cloud label according to the scene flow information and the inter-frame pose.

[0010] Optionally, in an embodiment of this application, the identifying static points and dynamic points in two frames of point cloud data to separate the static points and the dynamic points from the two frames of point cloud data, and determining two frames of static points and two frames of dynamic points includes: identifying the static points and the dynamic points according to a target semantic segmentation model; separating the static points and the dynamic points from the two frames of point cloud data according to a preset point cloud semantic segmentation network.

[0011] The second aspect of the present application provides an apparatus for simultaneously estimating the dynamic and static poses in a 4D space, including: a separation module, configured to identify static points and dynamic points in two frames of point cloud data, so as to separate the static points and the dynamic points from the two frames of point cloud data, and determine two frames of static points and two frames of dynamic points; a generation module, configured to calculate the inter-frame pose between the two frames of point cloud data according to the two frames of static points, and generate a dynamic point cloud label according to the two frames of dynamic points; an estimation module, configured to estimate the dynamic and static poses of the target object in the 4D space based on the inter-frame pose and the dynamic point cloud label.

[0012] Optionally, in an embodiment of the present application, it further includes: a construction module, configured to construct an unsupervised static point pose estimation network model for calculating the inter-frame pose between the two frames of point cloud data based on the two frames of static points before calculating the inter-frame pose between the two frames of point cloud data.

[0013] Optionally, in an embodiment of the present application, the generation module includes: a first generation unit, configured to generate a first dynamic object detection box and a second dynamic object detection box according to the two frames of dynamic points; a second generation unit, configured to generate the dynamic point cloud label based on the first dynamic object detection box and the second dynamic object detection box.

[0014] Optionally, in an embodiment of the present application, the second generation unit includes: an optimization subunit, configured to optimize the first dynamic object detection box and the second dynamic object detection box by using the two frames of dynamic points to obtain an optimized first dynamic object detection box and an optimized second dynamic object detection box; a construction subunit, configured to construct an unsupervised scene flow network model based on the optimized first dynamic object detection box and the optimized second dynamic object detection box in combination with the unsupervised static point pose estimation network model; a determination subunit, configured to obtain the scene flow information of the target object through the unsupervised scene flow network model, so as to determine the dynamic point cloud label according to the scene flow information and the inter-frame pose.

[0015] Optionally, in an embodiment of the present application, the separation module includes: an identification unit, configured to identify the static points and the dynamic points according to a target semantic segmentation model; a separation unit, configured to separate the static points and the dynamic points from the two frames of point cloud data according to a preset point cloud semantic segmentation network.

[0016] The third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the method for simultaneously estimating the dynamic and static poses in a 4D space as described in the above embodiments.

[0017] The fourth aspect of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above method for simultaneously estimating the dynamic and static poses in a 4D space.

[0018] The fifth aspect of the present application provides a computer program product including a computer program, which when executed is used to implement the above method for simultaneously estimating the dynamic and static poses in a 4D space.

[0019] The embodiments of the present application can distinguish and separate static points and dynamic points in point cloud data, calculate the inter-frame pose between two frames of static points through the separated static points and dynamic points, and construct a dynamic point cloud label, so as to simultaneously estimate the dynamic and static poses in a 4D space without supervision. Thus, the dynamic and static points in the point cloud are classified, and the unsupervised static point pose estimation network model and the unsupervised scene flow network model constructed by using the information of both can calculate the matching relationship between two frames of point clouds in an unsupervised manner and generate a 3D detection box. The dynamic point cloud label generated by the 3D detection box can better track the motion information of dynamic objects, and simultaneously estimate the pose information of static objects and dynamic objects, which helps the data closed-loop system of the perception function module in autonomous driving to more accurately understand and track the positions of different objects in the scene. Thus, it solves the problems in the related art that there are many problems in the unsupervised detection box generation method, the lack of semantic information cannot be introduced, resulting in poor perception and understanding ability of the environment, or the lack of matching processing for the static scene leads to large deviations in scene flow calculation, and the accuracy of the 3D detection box is also low, making it difficult to meet the application requirements of the data closed-loop system of the perception function module and autonomous driving in the actual scene.

[0020] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0022] Figure 1 is a schematic framework diagram of a system for simultaneously estimating the dynamic and static poses in a 4D space according to an embodiment of the present application;

[0023] Figure 2 is a flowchart of a method for simultaneously estimating the dynamic and static poses in a 4D space according to an embodiment of the present application;

[0024] Figure 3 is a schematic framework diagram of the construction process of an unsupervised static point pose estimation network model according to an embodiment of the present application;

[0025] Figure 4 It is a schematic framework diagram of the feature interaction module according to an embodiment of the present application;

[0026] Figure 5 It is a schematic structural diagram of the device for simultaneously estimating the dynamic and static poses in a 4D space provided according to an embodiment of the present application;

[0027] Figure 6 It is a schematic structural diagram of an electronic device provided according to an embodiment of the present application.

[0028] Reference numerals:

[0029] 10 - Device for simultaneously estimating the dynamic and static poses in a 4D space: 100 - Separation module, 200 - Generation module, and 300 - Estimation module; 601 - Memory, 602 - Processor, and 603 - Communication interface. Detailed implementation manners

[0030] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0031] The following describes a method, apparatus, electronic device, and storage medium for simultaneously estimating dynamic and static poses in a 4D space according to embodiments of the present application. In view of the problems in the related art's unsupervised detection box generation method mentioned in the above background technology, where there are many problems, semantic information cannot be introduced resulting in poor environmental perception and understanding ability, or the lack of matching processing for static scenes leads to large deviations in scene flow calculation, and the accuracy of 3D detection boxes is also low, making it difficult to meet the application requirements of the perception function module data closed-loop system and autonomous driving in actual scenarios, the present application provides a method for simultaneously estimating dynamic and static poses in a 4D space. In this method, static points and dynamic points in point cloud data can be distinguished and separated, and the inter-frame pose between two frames of static points can be calculated through the separated static points and dynamic points, and a dynamic point cloud label can be constructed, thereby realizing the simultaneous estimation of dynamic and static poses in a 4D space without supervision. Thus, the dynamic and static points in the point cloud are classified, and the unsupervised static point pose estimation network model and the unsupervised scene flow network model constructed using the information of both can calculate the matching relationship between two frames of point clouds and generate 3D detection boxes in an unsupervised manner. The dynamic point cloud label generated by the 3D detection box can better track the motion information of dynamic objects, and realize the simultaneous estimation of the pose information of static and dynamic objects, which helps the perception function module data closed-loop system in autonomous driving to more accurately understand and track the positions of different objects in the scene. Thus, the problems in the related art's unsupervised detection box generation method, such as many problems, poor environmental perception and understanding ability due to the inability to introduce semantic information, large deviations in scene flow calculation due to the lack of matching processing for static scenes, and low accuracy of 3D detection boxes, making it difficult to meet the application requirements of the perception function module data closed-loop system and autonomous driving in actual scenarios, are solved.

[0032] Before explaining the method for simultaneously estimating dynamic and static poses in a 4D space according to embodiments of the present application, the computing platform and development system (Computing Brain DEvelopment System, CBDES) and the system for simultaneously estimating dynamic and static poses in a 4D space involved in embodiments of the present application will be explained.

[0033] The Computing Brain Development System (CBDES) is the core of implementing autonomous driving technology. The computing platform provides the ability to process and store large-scale sensor data, and the development system can run data processing algorithms on the computing platform to extract information about the environment. CBDES consists of two parts: the Computing Base Brain (CBB) and the Graphical ADAS-AD Software Developer (GAASD). CBB is composed of computing platform hardware, a real-time kernel, middleware, and functional software. The functional software is the core of this product, aiming to provide basic algorithm components and frameworks for various intelligent driving systems. Based on the functional software library, the GAASD tool realizes the graphical and efficient development of algorithms through a graphical way.

[0034] Since the module weights of the CBDES functional software are learned in a data-driven manner, it is necessary to build a data closed-loop system to achieve data collection, annotation, simulation, and functional module update. The data closed-loop link in mainstream technologies mainly consists of five parts, namely Corner Case (extreme cases), data backhaul, ground truth annotation, model update, and OTA (Over-the-Air) distribution. Based on this, the embodiments of this application can realize the data closed-loop system of the perception functional module by simultaneously estimating the 4D spatial dynamic and static poses of the ground truth annotation.

[0035] Figure 1 It is a schematic diagram of the framework of the system for simultaneously estimating the dynamic and static poses in 4D space according to an embodiment of this application. As Figure 1 shown, the system for simultaneously estimating the dynamic and static poses in 4D space in the embodiments of this application can be but is not limited to being divided into three parts: dynamic and static segmentation, pose matrix calculation, and 3D detection box generation. Decoupling the main functional modules in the process of simultaneously estimating the 4D spatial dynamic and static poses facilitates adding new functional modules and has good scalability.

[0036] Specifically, Figure 2 It is a flowchart of a method for simultaneously estimating the dynamic and static poses in 4D space provided by an embodiment of this application.

[0037] As Figure 2 shown, the method for simultaneously estimating the dynamic and static poses in 4D space includes the following steps:

[0038] In step S201, identify the static points and dynamic points in two frames of point cloud data to separate the static points and dynamic points from the two frames of point cloud data, and determine two frames of static points and two frames of dynamic points.

[0039] It can be understood that the perception function module and the data closed-loop system are importantly related to the estimation of the dynamic and static poses of objects in the 4D space. The 4D pose estimation not only involves the position and orientation of an object in the three-dimensional space (i.e., the static pose), but also includes the changes of the object over time (i.e., the dynamic pose), that is, the estimation of the position, orientation, and motion state of an object in the four-dimensional space (three-dimensional space + time). The data output by the perception function module can be used as the input for the 4D pose estimation, and at the same time, the result of the 4D pose estimation can also provide more accurate target position and attribute information for the perception function module. Further, the data output by the 4D pose estimation and the perception function module are important inputs for the data closed-loop system, and the data closed-loop system can feedback the processed data to the 4D pose estimation and the perception function module to optimize and improve the performance of the 4D pose estimation and the perception function module.

[0040] In some embodiments, the estimation of the dynamic and static poses of objects in the 4D space needs to be realized by using certain data information. To ensure the accuracy of the estimation results of the dynamic and static poses of objects in the 4D space, the present application can, but is not limited to, using the point cloud data based on lidar as the data support.

[0041] The point cloud data of lidar provides the three-dimensional information of objects in the real world. How to generate accurate target detection boxes and label information from the point cloud data is one of the key problems of the autonomous driving perception model. However, since the point cloud data is irregular and disordered in the spatial dimension, it is relatively difficult to extract useful information from the disordered point cloud data. Secondly, since the perception target may be blocked by other objects in the environment, and there is certain noise interference in the sensor, these situations will make the target boundary in the point cloud unclear, increasing the difficulty of perception target detection. In addition, since the point cloud data set contains tens of thousands of points, processing such a large-scale data requires efficient algorithms and rich computing resources. Generally speaking, compared with image detection, the target detection task of point cloud is more difficult.

[0042] Based on this, the embodiments of the present application can divide the target detection task of point cloud into the target detection part of dynamic points and the target detection part of static points, and realize the target detection of point cloud without changing the total target detection task.

[0043] In view of the fact that the four-dimensional space involves the time dimension, a single frame of point cloud data cannot distinguish whether the dynamic body is moving in the driving scene. Therefore, the embodiments of the present application can, but are not limited to, obtaining two frames of point cloud data, identifying the static points and dynamic points in these two frames of point cloud data, and then separating the static points and dynamic points from the two frames of point cloud data to obtain the static points in the two frames of point cloud data and the dynamic points in the two frames of point cloud data.

[0044] Optionally, in an embodiment of the present application, static points and dynamic points in two frames of point cloud data are identified to separate static points and dynamic points from the two frames of point cloud data, and two frames of static points and two frames of dynamic points are determined, including: identifying static points and dynamic points according to a target semantic segmentation model; separating static points and dynamic points from the two frames of point cloud data according to a preset point cloud semantic segmentation network.

[0045] It can be understood that the preset point cloud semantic segmentation network can be understood here as a pre-established point cloud semantic segmentation network that can simultaneously input two frames of point cloud data for separating dynamic points and static points.

[0046] In the actual execution process, when the present application identifies static points and dynamic points in two frames of point cloud data, it mainly but not limited to uses a target semantic segmentation model to identify static points and dynamic points. Among them, the target semantic segmentation model can be understood here as a model established according to certain semantic labels that can finely divide each pixel in an image to achieve extremely high image segmentation accuracy, and it can even accurately identify target objects in complex backgrounds. It should be noted that specific semantic labels can be set or selected by those skilled in the art according to actual situations, and only exemplary descriptions are given in the embodiments of the present application without specific limitations.

[0047] After identifying static points and dynamic points in two frames of point cloud data, the embodiments of the present application can use a certain point cloud semantic segmentation network to separate static points and dynamic points in the two frames of point cloud data. Specifically, considering that the point cloud semantic segmentation result of a single frame cannot distinguish whether a dynamic object is moving in a driving scene. For example, both a parked car and a moving car will be classified as dynamic objects. Therefore, the embodiments of the present application can incorporate the temporal information in the point cloud data into the point cloud semantic segmentation network, thereby realizing inputting two frames of point cloud data into a certain point cloud semantic segmentation network at the same time, separating real dynamic points and static points, and obtaining two frames of dynamic points and two frames of static points.

[0048] Step S202, calculate the inter-frame pose between the two frames of point cloud data according to the two frames of static points, and generate a dynamic point cloud label according to the two frames of dynamic points.

[0049] As a possible implementation manner, after obtaining dynamic points and static points in two frames of point cloud data, the embodiments of the present application can respectively use dynamic point information and static point information to construct a certain unsupervised model, and then estimate the inter-frame pose between the two frames of point cloud data using the two frames of static points and generate a dynamic point cloud label using the two frames of dynamic points in an unsupervised manner.

[0050] Optionally, in an embodiment of the present application, before calculating the inter-frame pose between two frames of point cloud data, it includes: constructing an unsupervised static point pose estimation network model for calculating the inter-frame pose between two frames of point cloud data based on two frames of static points.

[0051] In some embodiments, in the process of generating dynamic point cloud labels using dynamic points, if the pose transformation between two frames is not considered, it may cause the static points to calculate velocity information, which will ultimately lead to a very high error rate in the generated labels. Therefore, the present application can consider temporal information and only perform pose transformation on static points in two frames, thereby constructing an unsupervised static point pose estimation network model for calculating the inter-frame pose between two frames of point cloud data.

[0052] When calculating the inter-frame pose based on static points, point cloud registration needs to be performed first. Considering that the registration effect of the traditional point-to-point point cloud registration scheme has a high dependence on the quality of point cloud features, the embodiments of the present application can, but are not limited to, based on the probability distribution scheme, use the Gaussian mixture model (GMM) as the benchmark method, and use the maximum likelihood method to find the best match between point clouds under this benchmark method. Among them, according to the multi-modal probability distribution, the point cloud registration can be decomposed into two processes of fitting and registration, which can be, but are not limited to, expressed as follows:

[0053] First, the multi-modal probability distribution can be expressed as follows:

[0054]

[0055] Based on this multi-modal probability distribution, the fitting process can be, but are not limited to, expressed as follows:

[0056]

[0057] And, based on this multi-modal probability distribution, the registration process can be, but are not limited to, expressed as follows:

[0058]

[0059] Among them, x represents the observed data, J is the number of sub-Gaussian models in the mixture model, j = 1, 2,..., N, π j is the probability that the observed data belongs to the j-th sub-model, π j ≥0, N(x|μ j ,∑ j ) is the Gaussian distribution density function of the j-th sub-model, which can be, but are not limited to, expressed as:

[0060]

[0061] Among them, argmax(f(x)) is the variable point x corresponding to the maximum value of f(x).

[0062] Furthermore, the embodiments of the present application can also combine the GMM model with deep learning, design a loss function, and update and train the network through backpropagation in a self-supervised manner. Finally, a pose transformation matrix of static points with high accuracy is obtained, providing correct pose parameters between two frames for the overall framework of label generation, thereby completing the construction of an unsupervised static point pose estimation network model.

[0063] For example, the present application can be based on the characteristic that two frames of point cloud data are sampled from the same GMM model. First, a GMM model is constructed using the two frames of point cloud, and then the rotation matrix between the two frames of point cloud is calculated through the rotation relationship between the GMM model and the two frames of point cloud. The process can be but is not limited to being expressed as follows:

[0064] Figure 3 It is a framework schematic diagram of the construction process of an unsupervised static point pose estimation network model according to an embodiment of the present application. As Figure 3 shown, the embodiments of the present application can first input the frame into the Feature Interaction (FI) module to learn the posterior by embedding self-point cloud and cross-point cloud information; then use the P module to calculate parameters such as the mean, covariance, and mixing coefficients of the GMM to construct a GMM model; finally, use the T module to estimate the transformation matrix and output the pose parameters of two frames of static points to complete the construction of the unsupervised static point pose estimation network model. Among them, the model construction process is trained in an unsupervised manner.

[0065] Figure 4 It is a framework schematic diagram of the feature interaction module according to an embodiment of the present application. As Figure 4 shown, the FI module in the embodiments of the present application can be but is not limited to being composed of two self-attention (self) blocks, a cross-attention (CA) block, and a convolutional layer, and its main function is to calculate the posterior probability. In the embodiments of the present application, the posterior probability describes the probability that each point in the point cloud belongs to a certain Gaussian distribution. For example, when calculating the posterior probability through the FI module, two frames of point cloud X∈R 1×3 , Y∈R 3×3 can be used as inputs, and the calculation formula can be but is not limited to being expressed as follows:

[0066] α ijk = [α ik , α jk = [FI(x i |X, Y), FI(y j |Y, X)]

[0067] Among them, X and Y represent two frames of point cloud, and αik Denote the point \(x\) in the point cloud \(X\) i , \(\alpha\) jk Denote the point \(y\) in the point cloud \(Y\) j , \(\alpha\) ik and \(\alpha\) jk are all probabilities belonging to the \(k\)-th Gaussian component, and this probability can be calculated by the FI module; FI(\(x\) i |X,Y) represents the posterior probability that the point \(x\) i belongs to a certain Gaussian component calculated by the FI module, and FI(\(y\) j |Y,X) represents the posterior probability that the point \(y\) j belongs to a certain Gaussian component calculated by the FI module.

[0068] After inputting two frames of point cloud data, the embodiments of the present application can first use two self and CA blocks to extract the feature information of the two point clouds, then use a convolutional layer to obtain the final feature of each point, and finally obtain a weight vector between 0 and 1 through softmax. Among them, the weight vector is \(L\)-dimensional, representing the probability that the point belongs to which component of the Gaussian mixture model, that is, the posterior probability.

[0069] The P module can calculate the model parameters of the GMM according to the posterior probability. The principle is to use the posterior probability to calculate the mean, covariance matrix and mixing coefficient of each Gaussian component \(l\) as the parameters of this Gaussian function. Its calculation process can be but is not limited to being expressed as follows:

[0070] (1) Calculate the weighted average of the two frames of point clouds \((X,Y)\) in the Gaussian component \(l\), and its formula can be but is not limited to being expressed as follows:

[0071]

[0072] Among them, \(\alpha\) il is the posterior probability that the point \(x\) i belongs to the Gaussian component \(l\), and \(\alpha\) jl is the posterior probability that the point \(y\) j belongs to the Gaussian component \(l\), and \(N\) and \(M\) are the number of points in each point cloud.

[0073] (2) Calculate the mean and covariance matrix of the Gaussian component \(l\), and its formula can be but is not limited to being expressed as follows:

[0074]

[0075] Among them,

[0076] (3) Calculate the mixing coefficients of different Gaussian distributions, and its formula can be but is not limited to being expressed as follows:

[0077]

[0078] Among them, π l represents the proportion of the l-th Gaussian component in the entire model, and all the mixing coefficients π l must satisfy the condition:

[0079] The T module can calculate the transformation matrix based on the calculated GMM parameters. Among them, the transformation relationship between the point cloud and the GMM model can be but is not limited to being expressed as:

[0080] V = R X X + t X , V = R Y Y + t Y ,

[0081] Among them, V is the virtual scene represented by the GMM, and R X is the rotation formula of the point cloud X, and R Y is the rotation formula of the point cloud Y.

[0082] Perform singular value decomposition on the model parameter matrix of the GMM, then the formula of the rotation matrix can be but is not limited to being expressed as:

[0083]

[0084] Among them, U X and V X respectively represent the left singular matrix and the right singular matrix in the singular value decomposition (SVD) of the matrix W X Λ X P X Λ X μ, and ∑ X is a diagonal matrix, and the diagonal elements are its singular values.

[0085] After determining the rotation matrix, the embodiments of the present application can calculate the translation matrix using the following formula:

[0086]

[0087] Finally, the transformation relationship between two frames of point clouds can be but is not limited to being expressed as follows:

[0088]

[0089] Among them, the rotation matrix is and the translation matrix is

[0090] Thus, the embodiments of the present application can obtain the pose parameters of two frames of static points, thereby completing the construction of the unsupervised static point pose estimation network model.

[0091] Furthermore, in the embodiments of the present application, the loss value in the construction process of the unsupervised static point pose estimation network model can be but is not limited to being composed of the chamfer distance loss and the cycle transformation loss.

[0092] Among them, the chamfer distance loss in the embodiments of the present application can be but is not limited to calculating the pose transformation matrix of two point clouds through the T module, performing pose transformation on one of the point clouds and then calculating the chamfer distance with the other point cloud to obtain the chamfer distance loss function, and its function can be but is not limited to being expressed as follows:

[0093]

[0094] And, due to the transformation matrix T estimated by the T module for transforming from point cloud X to point cloud Y X,Y and the transformation matrix T for transforming from point cloud Y to point cloud X Y,X These two transformation matrices are invertible. Therefore, the embodiments of the present application can use these two cyclic transformations as another unsupervised loss, and its formula can be but is not limited to being expressed as follows:

[0095] loss ct =‖T XY T YX -I‖ 2 ,

[0096] Then the final target loss function can be but is not limited to being expressed as:

[0097] loss=loss cf +loss ct .

[0098] Optionally, in an embodiment of the present application, generating a dynamic point cloud label according to two frames of dynamic points includes: generating a first dynamic object detection frame and a second dynamic object detection frame according to two frames of dynamic points; generating a dynamic point cloud label based on the first dynamic object detection frame and the second dynamic object detection frame.

[0099] Those skilled in the art can understand that dynamic labels can perform refined identification and tracking of objects based on the detection frame information of the objects. Since the detection frame provides the position, shape, and time-varying information of the object in three-dimensional space, dynamic labels can accordingly classify and mark the objects more accurately. This helps to effectively distinguish and track multiple objects in a complex scene in an unsupervised environment. And the "dynamic" characteristic of the dynamic label enables it to be automatically updated as the object changes. When the detection frame information of the object changes, such as position movement, shape change, or speed change, etc., the dynamic label will be adjusted accordingly to reflect the latest state of the object. This real-time update mechanism helps to improve the accuracy of object change information recognition, enabling the unsupervised object recognition system to respond to object changes more timely.

[0100] In some embodiments, the embodiments of the present application can generate dynamic point cloud labels based on dynamic points in two frames of point cloud data, so as to use the dynamic point cloud labels to refine the recognition and tracking of dynamic objects.

[0101] Specifically, the embodiments of the present application can generate different dynamic object detection boxes, namely a first dynamic object detection box and a second dynamic object detection box, based on two frames of dynamic points by using a neural network, and then generate dynamic point cloud labels for an unsupervised network available for the dynamic object detection boxes.

[0102] Optionally, in an embodiment of the present application, generating dynamic point cloud labels based on the first dynamic object detection box and the second dynamic object detection box includes: optimizing the first dynamic object detection box and the second dynamic object detection box by using two frames of dynamic points to obtain optimized first and second dynamic object detection boxes; based on the optimized first and second dynamic object detection boxes, combining an unsupervised static pose estimation network model to construct an unsupervised scene flow network model; obtaining scene flow information of a target object through the unsupervised scene flow network model, so as to determine dynamic point cloud labels according to the scene flow information and the inter-frame pose.

[0103] Based on the related descriptions of other embodiments, it can be understood that the embodiments of the present application can generate different dynamic object detection boxes based on two frames of dynamic points, and generate dynamic point cloud labels for an unsupervised network by using different dynamic object detection boxes.

[0104] In the actual execution process, after the first dynamic object detection box and the second dynamic object detection box are generated in the present application, in order to generate dynamic point cloud labels in an unsupervised manner, it is possible but not limited to use the convex hull of the dynamic points as self-supervision, continuously optimize the neural network for generating the first dynamic object detection box and the second dynamic object detection box, and finally generate two frames of optimized dynamic object detection boxes, namely the optimized first dynamic object detection box and the second dynamic object detection box, so as to construct a self-supervised neural network model based on this process.

[0105] Furthermore, based on this self-supervised neural network model, the embodiments of the present application can also transform the dynamic object detection box in the first frame, i.e., the first dynamic object detection box, into the second frame through the inter-frame pose transformation matrix estimated by the unsupervised static pose estimation network model, and input the point cloud transformation result of the first frame and the point cloud data in the dynamic object detection box in the second frame, i.e., the second dynamic object detection box, into the dynamic scene flow network to obtain the scene flow information of the dynamic object between the two frames, and construct an unsupervised scene flow network model.

[0106] Based on this unsupervised scene flow network model, the embodiments of the present application can obtain the scene flow information between two frames of point cloud data. Moreover, the embodiments of the present application can also obtain accurate data associations between 3D targets by combining the scene flow information between two frames of point cloud data and the inter-frame pose information obtained from static points, thereby effectively improving the accuracy of dynamic point cloud labels.

[0107] In addition, the dynamic point cloud label generation process in the embodiments of the present application can be constructed as a complete network framework. The scene flow information generated by the unsupervised scene flow network model can continuously optimize the network framework for generating dynamic point cloud labels, so that the dynamic point cloud labels in the embodiments of the present application can maintain good accuracy in an unsupervised situation, improve the generation efficiency of dynamic point cloud labels, and at the same time can eliminate the cumbersome steps of manual label annotation, saving a large amount of manpower and material resources.

[0108] Step S203: Estimate the static and dynamic poses of the target object in the 4D space based on the inter-frame pose and the dynamic point cloud label.

[0109] It can be understood that the target object here refers to a static object or a dynamic object for which static pose estimation or dynamic pose estimation needs to be performed.

[0110] In some other embodiments, after obtaining the inter-frame pose and the dynamic point cloud label generated in an unsupervised situation, the embodiments of the present application can use the unsupervised static point pose estimation network model and the unsupervised scene flow network model constructed during the generation of the inter-frame pose and the dynamic point cloud label to simultaneously estimate the pose of the static object or the motion information of the dynamic object, and estimate the pose of the dynamic object based on the motion information.

[0111] By simultaneously estimating the poses of static and dynamic objects, the perception function module data closed-loop system can more accurately understand and track the positions of these elements in the scene, so as to better plan and make decisions, and thus achieve safer and more efficient autonomous driving.

[0112] The method for simultaneously estimating the dynamic and static poses in a 4D space proposed according to the embodiments of the present application can distinguish and separate static points and dynamic points in point cloud data, calculate the inter-frame pose between two frames of static points through the separated static points and dynamic points, and construct a dynamic point cloud label, thereby realizing the simultaneous estimation of the dynamic and static poses in the 4D space without supervision. Thus, the classification of dynamic and static points in the point cloud is realized. The unsupervised static point pose estimation network model and the unsupervised scene flow network model constructed by using the information of both can calculate the matching relationship between two frames of point clouds in an unsupervised manner and generate a 3D detection box. The dynamic point cloud label generated by the 3D detection box can better track the motion information of dynamic objects, realize the simultaneous estimation of the pose information of static objects and dynamic objects, and help the data closed-loop system of the perception function module in autonomous driving to more accurately understand and track the positions of different objects in the scene. Thus, the problems existing in the unsupervised detection box generation method in the related art are solved. There are many problems, the lack of semantic information cannot cause poor perception and understanding ability of the environment, or the lack of matching processing of the static scene leads to a large deviation in scene flow calculation, and the accuracy of the 3D detection box is also low, which is difficult to meet the application requirements of the data closed-loop system of the perception function module and autonomous driving in the actual scene, etc.

[0113] Next, a device for simultaneously estimating the dynamic and static poses in a 4D space proposed according to the embodiments of the present application will be described with reference to the accompanying drawings.

[0114] Figure 5 It is a schematic structural diagram of a device for simultaneously estimating the dynamic and static poses in a 4D space according to the embodiments of the present application.

[0115] As Figure 5 shown, the device 10 for simultaneously estimating the dynamic and static poses in a 4D space includes: a separation module 100, a generation module 200, and an estimation module 300.

[0116] Among them, the separation module 100 is used to identify static points and dynamic points in two frames of point cloud data, so as to separate static points and dynamic points from the two frames of point cloud data, and determine two frames of static points and two frames of dynamic points.

[0117] The generation module 200 is used to calculate the inter-frame pose between two frames of point cloud data according to two frames of static points, and generate a dynamic point cloud label according to two frames of dynamic points.

[0118] The estimation module 300 is used to estimate the dynamic and static poses of the target object in the 4D space based on the inter-frame pose and the dynamic point cloud label.

[0119] Optionally, in an embodiment of the present application, it further includes: a construction module, which is used to construct an unsupervised static point pose estimation network model for calculating the inter-frame pose between two frames of point cloud data based on two frames of static points before calculating the inter-frame pose between two frames of point cloud data.

[0120] Optionally, in an embodiment of the present application, the generation module 200 includes: a first generation unit and a second generation unit.

[0121] Among them, the first generation unit is used to generate a first dynamic object detection frame and a second dynamic object detection frame according to two frames of dynamic points.

[0122] The second generation unit is used to generate a dynamic point cloud label based on the first dynamic object detection frame and the second dynamic object detection frame.

[0123] Optionally, in an embodiment of the present application, the second generation unit includes: an optimization subunit, a construction subunit, and a determination subunit.

[0124] The optimization subunit is used to optimize the first dynamic object detection frame and the second dynamic object detection frame by using two frames of dynamic points to obtain the optimized first dynamic object detection frame and the second dynamic object detection frame.

[0125] The construction subunit is used to construct an unsupervised scene flow network model based on the optimized first dynamic object detection frame and the second dynamic object detection frame in combination with an unsupervised static point pose estimation network model.

[0126] The determination subunit is used to obtain the scene flow information of the target object through the unsupervised scene flow network model to determine the dynamic point cloud label according to the scene flow information and the inter-frame pose.

[0127] Optionally, in an embodiment of the present application, the separation module 100 includes: an identification unit and a separation unit.

[0128] Among them, the identification unit is used to identify static points and dynamic points according to the target semantic segmentation model.

[0129] The separation unit is used to separate static points and dynamic points from two frames of point cloud data according to a preset point cloud semantic segmentation network.

[0130] It should be noted that the foregoing explanation of the embodiment of the method for simultaneously estimating the dynamic and static poses in the 4D space is also applicable to the device for simultaneously estimating the dynamic and static poses in the 4D space of this embodiment, and will not be repeated here.

[0131] The dynamic and static pose simultaneous estimation device in 4D space proposed according to the embodiments of the present application can distinguish and separate static points and dynamic points in point cloud data, calculate the inter-frame pose between two frames of static points through the separated static points and dynamic points, and construct a dynamic point cloud label, so as to realize the simultaneous estimation of dynamic and static poses in 4D space without supervision. Thus, the dynamic and static points in the point cloud are classified. The unsupervised static point pose estimation network model and the unsupervised scene flow network model constructed by using the information of both can calculate the matching relationship between two frames of point clouds in an unsupervised manner and generate a 3D detection box. The dynamic point cloud label generated by this 3D detection box can better track the motion information of dynamic objects, and realize the simultaneous estimation of the pose information of static objects and dynamic objects, which helps the data closed-loop system of the perception function module in autonomous driving to more accurately understand and track the positions of different objects in the scene. Thus, it solves the problems in the related art that there are many problems in the unsupervised detection box generation method, the lack of semantic information leads to poor environmental perception and understanding ability, or the lack of matching processing for the static scene results in large deviations in scene flow calculation, and the accuracy of the 3D detection box is also low, making it difficult to meet the application requirements of the data closed-loop system of the perception function module and autonomous driving in the actual scene, etc.

[0132] Figure 6 The structural schematic diagram of the electronic device provided by the embodiment of the present application. The electronic device may include:

[0133] A memory 601, a processor 602, and a computer program stored on the memory 601 and executable on the processor 602.

[0134] When the processor 602 executes the program, it implements the dynamic and static pose simultaneous estimation method in 6D space provided in the above embodiment.

[0135] Furthermore, the electronic device further includes:

[0136] A communication interface 603 for communication between the memory 601 and the processor 602.

[0137] The memory 601 is used to store a computer program executable on the processor 602.

[0138] The memory 601 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0139] If the memory 601, the processor 602, and the communication interface 603 are implemented independently, the communication interface 603, the memory 601, and the processor 602 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used in Figure 6 to represent it, but it does not mean that there is only one bus or one type of bus.

[0140] Optionally, in a specific implementation, if the memory 601, the processor 602, and the communication interface 603 are integrated on a single chip, the memory 601, the processor 602, and the communication interface 603 can communicate with each other through an internal interface.

[0141] The processor 602 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0142] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for simultaneously estimating the dynamic and static poses in a 4D space as described above is implemented.

[0143] The embodiments of the present application further provide a computer program product, including a computer program, and the computer program can run computer instructions, and when the computer instructions are executed by a processor, the method for simultaneously estimating the dynamic and static poses in a 4D space provided by the embodiments of the present application is implemented.

[0144] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0145] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0146] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions may be executed in a manner that is not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in the reverse order, which should be understood by those skilled in the art to which the embodiments of this application pertain.

[0147] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0148] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0149] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0150] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0151] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for simultaneously estimating static and dynamic posture in 4D space, characterized in that: The following steps are involved: Identify static points and dynamic points in two frames of point cloud data to separate the static points and the dynamic points from the two frames of point cloud data, and determine two frames of static points and two frames of dynamic points; Calculate the inter-frame pose between the two frames of point cloud data according to the two frames of static points, and generate dynamic point cloud labels according to the two frames of dynamic points; Based on the inter-frame pose and the dynamic point cloud label, the dynamic and static pose of the target object in the 4D space is estimated.

2. The method according to claim 1, characterized in that Before calculating the inter-frame pose between the two frames of point cloud data, the method further includes: An unsupervised static point pose estimation network model is constructed based on the two frames of static points for calculating the inter-frame pose between the two frames of point cloud data.

3. The method according to claim 2, characterized in that Generating a dynamic point cloud label according to the two frames of dynamic points includes: Generate a first dynamic object detection frame and a second dynamic object detection frame according to the two frames of dynamic points; The dynamic point cloud label is generated based on the first dynamic object detection frame and the second dynamic object detection frame.

4. The method according to claim 3, characterized in that The generating the dynamic point cloud label based on the first dynamic object detection frame and the second dynamic object detection frame includes: Optimizing the first dynamic object detection frame and the second dynamic object detection frame using the two frames of dynamic points to obtain optimized first dynamic object detection frame and second dynamic object detection frame; Based on the optimized first dynamic object detection frame and the second dynamic object detection frame, in combination with the unsupervised static point pose estimation network model, an unsupervised scene flow network model is constructed; The scene flow information of the target object is obtained through the unsupervised scene flow network model, so as to determine the dynamic point cloud label according to the scene flow information and the inter-frame pose.

5. The method according to claim 1, characterized in that The identifying of static points and dynamic points in two frames of point cloud data to separate the static points and the dynamic points from the two frames of point cloud data and determining two frames of static points and two frames of dynamic points includes: Identify the static points and the dynamic points according to a target semantic segmentation model; The static points and the dynamic points are separated from the two frames of point cloud data according to a preset point cloud semantic segmentation network.

6. A device for simultaneously estimating static and dynamic posture in 4D space, characterized in that: include: A separation module, used for identifying static points and dynamic points in two frames of point cloud data, so as to separate the static points and the dynamic points from the two frames of point cloud data, and determine two frames of static points and two frames of dynamic points; A generating module, used for calculating the inter-frame pose between the two frames of point cloud data according to the two frames of static points, and generating dynamic point cloud labels according to the two frames of dynamic points; The estimation module is used to estimate the dynamic and static pose of the target object in the 4D space based on the inter-frame pose and the dynamic point cloud label.

7. The device according to claim 6, characterized in that Also includes: A construction module is used to construct an unsupervised static point pose estimation network model for calculating the inter-frame pose between the two frames of point cloud data based on the two frames of static points before calculating the inter-frame pose between the two frames of point cloud data.

8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for simultaneously estimating the dynamic and static posture of a 4D space as described in any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method for simultaneously estimating dynamic and static pose in a 4D space as described in any one of claims 1 to 5.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed, it is used to implement the method for simultaneously estimating the dynamic and static posture in a 4D space as described in any one of claims 1 to 5.