Point cloud scene flow estimation method under feature self-supervision

Through the point cloud scene flow estimation method under feature self-supervisation, the Setconv++ layer processing and self-supervisation loss terms are used, and combined with the random walk algorithm to refine it, the problem of corresponding errors in the source point cloud structure and matching errors in the similar structure is solved, and more accurate point cloud scene flow estimation is achieved.

CN119941784APending Publication Date: 2025-05-06JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411778639.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, when some structures in the source point cloud correspond to the target point cloud that have not been collected, they will construct incorrect soft corresponding points, resulting in incorrect point cloud scene flow estimation; at the same time, areas with adjacent similar structures in the target point cloud are easily selected to construct incorrect soft corresponding points.

Method used

The point cloud scene flow estimation method under feature self-supervisation is used to obtain two sets of point cloud data at adjacent moments through a laser scanner, and the feature matrix is ​​extracted using 3-layer Setconv++ layers to calculate the feature cosine similarity matrix to build soft corresponding points, and the self-supervised loss terms and random walk algorithm are used to refine it to obtain the corrected point cloud scene flow.

Benefits of technology

It effectively reduces the mismatch of similar structures, avoids misest estimates caused by the missing part of the source point cloud area in the target point cloud, and ensures the accuracy of the point cloud scene flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941784A_ABST
    Figure CN119941784A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud scene flow estimation method under feature self-supervision, and belongs to the technical field of three-dimensional vision. According to the point cloud scene flow estimation method under feature self-supervision, two groups of point cloud data at adjacent moments are acquired through a laser scanner; the two groups of point cloud data are used as the input of a point feature extraction network and need to be processed by three layers of Setconv + + layers to obtain a feature matrix; calculating a feature cosine similarity matrix between the two groups of point clouds, and constructing soft corresponding points to obtain an initial point cloud scene flow; using a self-supervised loss item training model to obtain a point feature model and related parameters in an optimal transmission problem; using a refining process based on a random walk algorithm to obtain a corrected point cloud scene flow; according to the point cloud scene flow estimation method under feature self-supervision, the distinction degree of the point features between the points in the adjacent similar structures is enhanced by using the features obtained after each stage of processing of the feature extraction network, and wrong matching of the similar structures is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of three-dimensional vision technology, and in particular to a point cloud scene flow estimation method under feature self-supervision. Background Art

[0002] Scene flow is a three-dimensional vector field, which is an extension of optical flow in three-dimensional space. It is used to describe the motion pattern of each point in the scene from the next moment. It plays a key role in accurately perceiving the motion in a three-dimensional environment. The information contained in scene flow can be used in many tasks that require three-dimensional dynamic scene understanding, such as robotics, autonomous driving, and human-computer interaction. The earliest work related to scene flow estimation focused on continuous RGB image pairs and RGB-D image pairs. With the development of 3D sensors such as LiDAR and laser 3D scanners in recent years, high-density and high-precision point cloud data can be obtained, and people have begun to shift their attention to estimating scene flow directly from point clouds.

[0003] Compared with traditional methods, methods based on deep learning often have better accuracy, better generalization ability and stronger robustness. Methods using fully supervised models require a sufficient amount of point cloud data with true value annotations for training. However, the true value annotation of point cloud scene flow is very expensive, and the use of synthetic point cloud data cannot completely simulate and replace real-world data, which limits the application of models trained on synthetic point cloud data when generalized to real-world point cloud data. This has promoted the continuous development of self-supervised learning methods.

[0004] The most advanced self-supervised method SCOOP constructs soft corresponding points by learning point feature representation, and calculates the initial estimated point cloud scene flow based on it, and then directly optimizes the residual flow refinement vector instead of training another network for flow refinement. Experimental observations have found that there are still two typical problems in the method: First, some structures in the source point cloud will construct wrong soft corresponding points based on other points in the target point cloud when they correspond to the target point cloud that has not been collected; second, in the target point cloud, there are some adjacent areas with similar structures, and it is also easy to select points in them to construct wrong soft corresponding points. These wrong soft corresponding points will generate wrong estimated flows. Summary of the invention

[0005] In view of the above problems existing in the prior art, the present invention is proposed.

[0006] Therefore, the technical problem to be solved by the present invention is that some structures in the source point cloud, when corresponding to the target point cloud that has not been collected, will construct erroneous soft corresponding points by relying on other points in the target point cloud, and that in areas where there are some adjacent similar structures in the target point cloud, it is also easy to select points therein to construct erroneous soft corresponding points.

[0007] To achieve the above object, the present invention provides the following technical solution: a point cloud scene flow estimation method under feature self-supervision, comprising:

[0008] Acquire two sets of point cloud data at adjacent moments by using a laser scanner;

[0009] The two sets of point cloud data are used as the input of the point feature extraction network and need to be processed by three layers of Setconv++ to obtain the feature matrix;

[0010] Calculate the feature cosine similarity matrix between the two sets of point clouds, construct soft corresponding points, and obtain the initial point cloud scene flow;

[0011] The model is trained using the self-supervised loss term to obtain the point feature model and the relevant parameters in the optimal transmission problem;

[0012] Using a refinement process based on a random walk algorithm, the corrected point cloud scene flow is obtained.

[0013] As a further solution of the present invention: the processing method of the three Setconv++ layers is the same, except that each layer concatenates the result of the previous layer and the coordinate difference from the center point to the points in the neighborhood in the feature dimension;

[0014] Among them, each Setconv++ layer processing steps include:

[0015] After three rounds of 2D convolution, instance normalization, and Leaky ReLU activation, feature matrices at three different levels are obtained;

[0016] The feature matrices of three different levels are concatenated in the feature dimension to obtain concatenated multi-layer features;

[0017] Through one-dimensional convolution and instance normalization, the concatenated multi-layer features are reduced in dimension to obtain fused multi-layer features.

[0018] Add the fused multi-layer features and the third-layer features in the feature dimension and perform maximum pooling in the neighborhood to obtain the most significant features;

[0019] Concatenate the most significant features with the concatenated multi-layer features in the feature dimension;

[0020] The output features of the Setconv++ layer are obtained by using two-dimensional convolution, instance normalization, and maximum pooling within the neighborhood.

[0021] As a further solution of the present invention: using the transmission cost constructed by the cosine similarity between the point features to construct the weight, and then using the weight to construct the soft corresponding point;

[0022] The calculation formula of the weight is:

[0023]

[0024] Among them, w ij represents the weight of each j-th point in the target point cloud in the process of constructing the soft corresponding point of the i-th point in the source point cloud at the next moment, T * is the optimal transmission matrix, index t is the k point in the target point cloud t The set of index values ​​corresponding to the points.

[0025] The calculation formula of the soft corresponding point is:

[0026]

[0027] Among them, is the soft corresponding point of the i-th point in the source point cloud at the next moment, represents the jth point in the target point cloud.

[0028] As a further solution of the present invention: the steps of using the self-supervised loss term to train the model and obtain the point feature model and the relevant parameters in the optimal transmission problem include:

[0029] The confidence of the constructed soft corresponding points is calculated using the cosine similarity matrix and the weights used in the construction process of the soft corresponding points.

[0030] According to the confidence of the soft corresponding point, the confidence loss term, the chamfer distance loss term, the influence weight of the neighbor point on the center point, and the influence weight of the center point on the overall flow smoothness loss term are calculated respectively;

[0031] By influencing the weights, the flow smoothing loss term is calculated;

[0032] Construct the overall training loss based on the chamfer distance loss, confidence loss and flow smoothness loss;

[0033] Back-propagation is performed through the overall training loss to obtain the training point feature model and the relevant learnable parameters in the optimal transmission problem.

[0034] As a further solution of the present invention: the soft corresponding point confidence formula is:

[0035] ui =max(∑ j∈index w ij S ij ,0)

[0036] Among them, u i represents the confidence of the soft corresponding point of the i-th point in the source point cloud, S ij Represents the cosine similarity between the i-th point in the source point cloud and the j-th point in the target point cloud;

[0037] The chamfer distance loss term L dist The formula is:

[0038]

[0039] Among them, m represents the number of points in the point cloud, PC t+1 represents a point cloud composed of the soft corresponding points of each point in the source point cloud at the next moment, PC t+1 Represents the point cloud at the next moment.

[0040] As a further solution of the present invention: the confidence loss term L dist The formula is:

[0041]

[0042] Where m represents the number of points in the point cloud, u i Represents the confidence of the soft corresponding point of the i-th point in the source point cloud.

[0043] As a further solution of the present invention: the influence weight includes the weight x of the flow smoothing loss at each point in the overall flow smoothing loss term and the weight x of the flow smoothing loss at each point in the overall flow smoothing loss term

[0044] The weight x of the flow smoothing loss at each point in the overall flow smoothing loss term is:

[0045] x=softmax(u i ),i∈1,2,...,m

[0046] The weight of the flow smoothing loss at each point in the overall flow smoothing loss term for:

[0047]

[0048] Among them, N(i) represents the neighboring point corresponding to the i-th point in the source point cloud.

[0049] As a further solution of the present invention: the flow smoothing loss term L smooth_new for:

[0050]

[0051] Among them, k s Indicates the number of selected neighboring points, index s Represents the index of the neighboring points, and flow represents the preliminary estimated point cloud scene flow.

[0052] As a further solution of the present invention: the overall training loss is:

[0053] L total =λL dist +μL conf +νL smooth_new

[0054] Among them, λ, μ and γ are hyperparameters.

[0055] As a further solution of the present invention: the modified point cloud scene flow Flow rw for:

[0056] Flow rw =P 3 Flow (∞) +P 1 Flow alter +P 2 (valid Flow (∞) +(1-valid)Flow alter )

[0057] Among them, P 1 The mask matrix P represents the flow in the initial estimated point cloud scene flow with confidence in the interval [0,0.2]. 2 The mask matrix P represents the flow in the initial estimated point cloud scene flow with confidence in the interval [0.2, 0.8] 3 Represents the mask matrix of the flow in the initial estimated point cloud scene flow with confidence in the interval [0.8,1], Flow (∞) Represents the point cloud scene flow after fine-tuning, Flow alter Represents the reconstructed point cloud scene flow, valid represents the proportion of the combination, and is obtained by setting the values ​​in the confidence u in the interval [0,0.2] to 0.

[0058] Compared with the prior art, the beneficial effects of the present invention are as follows: the feature self-supervised point cloud scene flow estimation method utilizes the features obtained after processing at various levels of the feature extraction network to enhance the distinction between point features of adjacent similar structures, and effectively reduce the mismatch of similar structures; in addition, the present invention is based on a refinement module of a random walk algorithm, which can further fine-tune the initial estimated flow in the refinement module while reconstructing the low-confidence estimated flow therein, thereby avoiding the misestimation caused by the missing of some areas in the source point cloud in the target point cloud; the newly designed flow smoothing loss term is used as part of the self-supervised loss to promote the estimated point cloud scene flow to remain smooth while avoiding the expansion of misestimation to the surrounding area. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0060] Figure 1 This is a schematic diagram of the overall process described in the embodiment provided by the present invention.

[0061] Figure 2 Each Setconv++ layer processing step described in the embodiment provided by the present invention.

[0062] Figure 3 This is a schematic flow chart of step S4 in the embodiment provided by the present invention.

[0063] Figure 4 This is a schematic diagram showing that the corresponding area of ​​the local area in the source point cloud described in the embodiment of the present invention is not captured by the target point cloud.

[0064] Figure 5 A schematic diagram of a local area of ​​adjacent similar structures described in an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.

[0066] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0067] Secondly, the present invention is described in detail with reference to the schematic diagram. When describing the embodiments of the present invention in detail, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.

[0068] Furthermore, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.

[0069] Example 1

[0070] like Figures 1-2 As shown, the present invention provides a technical solution: a point cloud scene flow estimation method under feature self-supervision, comprising: S1: acquiring two sets of point cloud data at adjacent moments by a laser scanner;

[0071] S2: The two sets of point cloud data are used as the input of the point feature extraction network and need to be processed by three layers of Setconv++ to obtain the feature matrix;

[0072] S3: Calculate the feature cosine similarity matrix between the two sets of point clouds, construct soft corresponding points, and obtain the initial point cloud scene flow;

[0073] S4: Use the self-supervised loss term to train the model to obtain the point feature model and related parameters in the optimal transmission problem;

[0074] S5: Use the refinement process based on the random walk algorithm to obtain the corrected point cloud scene flow.

[0075] Among them, the design of the Setconv++ layer enhances the distinguishability of point features between points in adjacent similar structures, effectively reducing the mismatch of similar structures; the design of the random walk refinement module is to reconstruct the low-confidence estimation flow while further fine-tuning the point cloud scene flow, reducing the erroneous estimation flow caused by the missing actual corresponding areas of some structures in the source point cloud in the target point cloud; the new flow smoothing loss term is designed as part of the self-supervised loss, which promotes the estimated point cloud scene flow to remain smooth while avoiding the expansion of erroneous estimates to the surrounding areas.

[0076] It should be noted that the point cloud data is acquired by a laser scanning device to obtain point cloud data at different times, from which point clouds scanned at two adjacent times are selected. The two sets of point cloud data are respectively called source point cloud and target point cloud.

[0077] It should be noted that, based on the feature matrix corresponding to the two sets of point cloud data at adjacent moments output by the point feature extraction network, the transmission cost of the points between the two sets of point clouds is constructed by calculating the point feature cosine similarity matrix S between the two sets of point clouds; the relaxed version of the optimal transmission problem is constructed and solved using the transmission cost to obtain the optimal transmission matrix; the k points with the highest quality transmitted from each point in the source point cloud to the target point cloud are obtained from the optimal transmission matrix t points, calculate k respectively t The weights corresponding to the points are constructed by using the transmission cost constructed by the cosine similarity between the point features, and then the weights are used to construct the soft corresponding points. The specific weight calculation formula is:

[0078]

[0079] Among them, w ij represents the weight of each j-th point in the target point cloud in the process of constructing the soft corresponding point of the i-th point in the source point cloud at the next moment, T * is the optimal transmission matrix, index t is the k point in the target point cloud t The set of index values ​​corresponding to the points.

[0080] It should be noted that, with the calculated weights, the soft corresponding points of each point in the source point cloud at the next moment can be calculated. The calculation formula of the soft corresponding points is:

[0081]

[0082] Among them, is the soft corresponding point of the i-th point in the source point cloud at the next moment, represents the jth point in the target point cloud.

[0083] Furthermore, the processing methods of the three Setconv++ layers are the same, except that each layer concatenates the result of the previous layer and the coordinate difference from the center point to the points in the neighborhood in the feature dimension; the two sets of point cloud data are used as the input of the point feature extraction network and need to be processed by three Setconv++ layers;

[0084] Among them, each Setconv++ layer processing steps include:

[0085] S21: Input the two sets of point cloud data into the first Setconv++ layer. Each Setconv++ layer is processed by three rounds of two-dimensional convolution, instance normalization and Leaky ReLU activation to obtain feature matrices at three different levels.

[0086] S22: concatenate the feature matrices of three different levels in the feature dimension to obtain concatenated multi-layer features;

[0087] S23: The concatenated multi-layer features are subjected to dimensionality reduction processing on the feature dimension through one-dimensional convolution and instance normalization to obtain fused multi-layer features; specifically, the dimension of the point feature is reduced to the required size through one-dimensional convolution processing; instance normalization is to normalize the features in the channel dimension to enhance the network's ability to capture features; one-dimensional convolution and instance normalization are used to reduce the feature dimension of the concatenated multi-layer features, which belongs to the prior art and will not be elaborated here.

[0088] S24: Add the fused multi-layer features and the third-layer features in the feature dimension and perform maximum pooling in the neighborhood to obtain the most significant features. Here, maximum pooling specifically refers to selecting the maximum value in the neighborhood on each feature channel;

[0089] S25: concatenate the most significant feature with the concatenated multi-layer features in the feature dimension;

[0090] S26: Use two-dimensional convolution, instance normalization and maximum pooling in the neighborhood to obtain the output features of the Setconv++ layer; it should be noted that the output features of the Setconv++ layer are used as the output of the point feature extraction network, and maximum pooling specifically refers to selecting the maximum value in the neighborhood on each feature channel.

[0091] The above S21~S26 describe the operations performed in a Setconv++ layer, and a total of three Setconv++ layers are required for processing. The result of the previous layer and the coordinate difference from the center point to the point in the neighborhood are concatenated in the feature dimension as the input of the next layer. The first layer uses the point coordinates and the coordinate difference from the center point to the point in the neighborhood to be concatenated in the feature dimension as the input of the next layer.

[0092] It should be noted that the point feature extraction network is composed of three Setconv++ layers. After the point cloud input is processed by the point feature extraction network, the three-dimensional coordinate information is converted into high-dimensional feature information.

[0093] It should be noted that the random walk algorithm refers to the scene flow at each point in the point cloud, which will be adjusted to a certain extent according to the estimated scene flow at the surrounding points and the corresponding confidence (reliability); specifically, for a part of the point cloud scene flow with low confidence, it is reconstructed by the reliable estimated flow around it, and the high-confidence estimated flow will be adjusted based on the high-confidence estimated flow around it; the points in the middle are mixed.

[0094] The refinement process of the random walk algorithm is as follows: In step S5, the trained point feature extraction network and the relevant parameters in the optimal transmission problem are used to execute steps S1 to S3 to obtain a preliminary estimated point cloud scene flow, and the refined process based on the random walk algorithm is used to obtain the corrected point cloud scene flow Flowrw for:

[0095] Flow rw =P 3 Flow (∞) +P 1 Flow alter +P 2 (valid Flow (∞) +(1-valid)Flow alter )

[0096] Among them, P 1 The mask matrix P represents the flow in the initial estimated point cloud scene flow with confidence in the interval [0,0.2]. 2 The mask matrix P represents the flow in the initial estimated point cloud scene flow with confidence in the interval [0.2, 0.8] 3 Represents the mask matrix of the flow in the initial estimated point cloud scene flow with confidence in the interval [0.8,1], Flow (∞) Represents the point cloud scene flow after fine-tuning, Flow alter Represents the reconstructed point cloud scene flow, valid represents the proportion of the combination, and is obtained by setting the values ​​in the confidence u in the interval [0,0.2] to 0.

[0097] In Flow rw Based on this, the residual flow is refined to obtain the final point cloud scene flow Flow final , point cloud scene flow rw for:

[0098] Flow final =Flow rw +Flow r

[0099] Among them, Flow r represents the residual flow.

[0100] Optional, fine-tuned point cloud scene flow used in the refinement process (∞) The construction process is as follows:

[0101] Flow (∞) =(1-τ)(I-τM -1 ) -1 Flow (0)

[0102] Among them, Flow (0) represents the preliminary estimated point cloud scene flow, M is the transition probability matrix, and τ∈[0,1] is used to control the intensity of random walk.

[0103] Optional, reconstructed point cloud scene flow used in the refinement process alter , in Flow (0) It is constructed on the basis of, as shown in the following formula:

[0104] Flow alter =M*Flow (∞)

[0105] Optional, Flow (∞) with Flow alter The transfer matrix M during the construction process is:

[0106]

[0107]

[0108] where v i , v j The distribution represents the i-th and j-th nodes in the graph constructed using the point cloud constructed at the previous moment.

[0109] This feature-self-supervised point cloud scene flow estimation method uses the features obtained after processing at each level of the feature extraction network to enhance the distinction between point features in adjacent similar structures, effectively reducing the mismatch of similar structures; in addition, the present invention also designs a refinement module based on the random walk algorithm. While the refinement module can further fine-tune the initial estimation flow, it can also reconstruct the low-confidence estimation flow therein, avoiding the erroneous estimation caused by the missing of some areas in the source point cloud in the target point cloud. Finally, the present invention designs a refinement module based on the random walk algorithm. While the refinement module can further fine-tune the initial estimation flow, it can also reconstruct the low-confidence estimation flow therein, avoiding the erroneous estimation caused by the missing of some areas in the source point cloud in the target point cloud.

[0110] Example 2

[0111] like Figure 3 As shown, the difference between this embodiment and the previous embodiment is that it further records the steps of using the self-supervised loss term to train the model, obtain the point feature model and related parameters in the optimal transmission problem, and improves the point cloud scene flow estimation processing method under feature self-supervision.

[0112] Specifically, the steps of using the self-supervised loss term to train the model and obtain the point feature model and the relevant parameters in the optimal transmission problem include:

[0113] S41: Calculate the confidence of the constructed soft corresponding points by using the cosine similarity matrix and the weights occupied in the process of constructing the soft corresponding points; Specifically, calculate the confidence of the constructed soft corresponding points by using the cosine similarity matrix S and the weights w occupied in the process of constructing the soft corresponding points, which is equivalent to the confidence of the estimated point cloud scene flow; Based on the calculated confidence, further calculate to obtain the chamfer distance loss item;

[0114] It should be noted that the confidence formula of the soft corresponding point is:

[0115] u i =max(∑ j∈index w ij S ij ,0)

[0116] Among them, u i represents the confidence of the soft corresponding point of the i-th point in the source point cloud, S ij Represents the cosine similarity between the i-th point in the source point cloud and the j-th point in the target point cloud;

[0117] It should be noted that the chamfer distance loss term L dist The formula is:

[0118]

[0119] Among them, m represents the number of points in the point cloud, PC t+1 represents a point cloud composed of the soft corresponding points of each point in the source point cloud at the next moment, PC t+1 Represents the point cloud at the next moment.

[0120] S42: according to the confidence of the soft corresponding point, respectively calculate the influence weights of the confidence loss term, the chamfer distance loss term and the overall flow smoothing loss term corresponding to the center point of the neighboring point;

[0121] It should be noted that the confidence loss term L dist The formula is:

[0122]

[0123] Where m represents the number of points in the point cloud, u i Represents the confidence of the soft corresponding point of the i-th point in the source point cloud.

[0124] In this embodiment, the influence weight includes the weight x of the flow smoothing loss at each point in the overall flow smoothing loss term and the weight x of the flow smoothing loss at each point in the overall flow smoothing loss term.

[0125] It should be noted that the weight x of the flow smoothing loss at each point in the overall flow smoothing loss term is:

[0126] x=softmax(u i ),i∈1,2,...,m

[0127] It should be noted that the weight of the flow smoothing loss at each point in the overall flow smoothing loss term is for:

[0128]

[0129] Among them, N(i) represents the neighboring point corresponding to the i-th point in the source point cloud.

[0130] S43: Calculate the flow smoothing loss term by influencing the weight;

[0131] It should be noted that, with the obtained weights x and Calculate the flow smoothing loss term, the flow smoothing loss term L smooth_new for:

[0132]

[0133] Among them, k s Indicates the number of selected neighboring points, index s Represents the index of the neighboring points, and flow represents the preliminary estimated point cloud scene flow.

[0134] S44: Construct the overall training loss based on the chamfer distance loss term, the confidence loss term and the flow smoothness loss term;

[0135] It should be noted that the overall training loss is:

[0136] L total =λL dist +μL conf +νL smooth_new

[0137] Among them, λ, μ and γ are hyperparameters.

[0138] S45: Back propagate through the overall training loss to obtain the training point feature model and the relevant learnable parameters in the optimal transmission problem.

[0139] The feature-self-supervised point cloud scene flow estimation method utilizes the features obtained after processing at each level of the feature extraction network to enhance the distinction between point features of adjacent similar structures, and effectively reduces the mismatch of similar structures. In addition, the present invention is based on a refinement module of the random walk algorithm. While further fine-tuning the initial estimation flow in the refinement module, it can also reconstruct the low-confidence estimation flow therein, avoiding the misestimation caused by the missing of some areas in the source point cloud in the target point cloud.

[0140] Example 3

[0141] The present embodiment is different from the previous embodiment in that the present embodiment provides a point cloud target tracking method, which obtains point cloud scene flow based on a point cloud scene flow estimation method under feature self-supervision, and implements point cloud target tracking based on it.

[0142] Example 4

[0143] The difference between this embodiment and the previous embodiment is that this embodiment records a point cloud target tracking system, including:

[0144] A point cloud data acquisition module is used to acquire two sets of point clouds at adjacent moments;

[0145] A data processing module, which uses a point cloud scene flow estimation method under feature self-supervision to obtain a point cloud scene flow and / or a point cloud target tracking method to process the point cloud obtained by the point cloud data acquisition module;

[0146] The output display module is used to visualize the point cloud scene flow estimation results.

[0147] Example 5

[0148] As attached Figures 4-5 The difference between this embodiment and the previous embodiment is that this embodiment records the training and testing of this method and the prior art SCOOP method on the Kaggle platform, proving that this method can better learn the local structure of the point cloud, reduce the possibility of selecting points in adjacent similar structures during the construction of soft corresponding points, make the obtained soft corresponding points more reliable, and improve the accuracy of the estimated point cloud scene flow in the area.

[0149] It is carried out on the kaggle platform. The selected GPU model is P100, and the environment consisting of python3.10.12, cuda11.8, cudnn8.9 and pytorch3.10.12 is used on Ubuntu22.04.

[0150] The experimental results are shown in the attached figure. Figure 4 This is the case where the corresponding area of ​​the local area in the source point cloud is not collected by the target point cloud, where the green points are the point clouds obtained by processing the source point cloud scene flow true value, and the red points are the point clouds obtained by processing the source point cloud scene flow true value by estimating the point cloud scene flow true value. The SCOOP method produces obvious errors, which will be affected by this part of the area and calculate the wrong soft correspondence, causing the estimated point cloud scene flow in this area to deviate from the true value. The method of the present invention can avoid the influence of this area and obtain relatively correct prediction results.

[0151] Figure 5This is the case where there are local areas with adjacent similar structures, where the green points are point clouds obtained by processing the source point cloud with the true value of the point cloud scene flow, and the red points are point clouds obtained by processing the source point cloud with the estimated true value of the point cloud scene flow. The SCOOP method produces obvious errors, as it uses points in adjacent similar structures in the process of constructing soft corresponding points, resulting in incorrect soft correspondences, which ultimately reduces the accuracy of the estimated point cloud scene flow in this area.

[0152] Through the above comparison, it is not difficult to find that this method can better learn the local structure of the point cloud, reduce the possibility of selecting points in adjacent similar structures during the construction of soft corresponding points, make the obtained soft corresponding points more reliable, and improve the accuracy of the estimated point cloud scene flow in this area.

[0153] Importantly, it should be noted that the construction and arrangement of the present application shown in a number of different exemplary embodiments are only exemplary. Although only a few embodiments are described in detail in this disclosure, it should be readily understood by those who refer to this disclosure that many modifications are possible, for example, the size, scale, structure, shape and proportion of various elements, and parameter values ​​such as temperature, pressure, etc., mounting arrangements, use of materials, color, directional changes, etc., without substantially departing from the novel teachings and advantages of the subject matter described in the application. For example, the element shown as integrally formed can be composed of multiple parts or elements, the position of the element can be inverted or otherwise changed, and the nature or number or position of the discrete element can be changed or changed. Therefore, all such modifications are intended to be included in the scope of the present invention. The order or sequence of any process or method steps can be changed or reordered according to alternative embodiments. In the claims, any "device plus function" clause is intended to cover the structure of the execution function described herein, and is not only structurally equivalent but also equivalent structure. Without departing from the scope of the present invention, other substitutions, modifications, changes and omissions can be made in the design, operating conditions and arrangement of the exemplary embodiments. Therefore, the invention is not limited to a specific embodiment, but extends to numerous modifications still falling within the scope of the appended claims.

[0154] Furthermore, in order to provide a concise description of exemplary embodiments, all features of an actual embodiment may not be described, i.e., those features that are not relevant to the best mode presently contemplated for carrying out the invention or those features that are not relevant to implementing the invention.

[0155] It should be understood that in the development of any actual implementation, as in any engineering or design project, numerous implementation-specific decisions may be made. Such a development effort may be complex and time-consuming, but for those of ordinary skill having the benefit of this disclosure, the development effort will be a routine task of design, fabrication, and production without undue experimentation.

[0156] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A point cloud scene flow estimation method under feature self-supervision, characterized by: include, Acquire two sets of point cloud data at adjacent moments by using a laser scanner; The two sets of point cloud data are used as the input of the point feature extraction network and need to be processed by three layers of Setconv++ to obtain the feature matrix; Calculate the feature cosine similarity matrix between the two sets of point clouds, construct soft corresponding points, and obtain the initial point cloud scene flow; The model is trained using the self-supervised loss term to obtain the point feature model and the relevant parameters in the optimal transmission problem; Using a refinement process based on a random walk algorithm, the corrected point cloud scene flow is obtained.

2. The method for estimating scene flow from point cloud under feature self-supervision as claimed in claim 1, characterized in that: The processing method of the three Setconv++ layers is the same, except that each layer concatenates the result of the previous layer and the coordinate difference between the center point and the points in the neighborhood in the feature dimension; Among them, each Setconv++ layer processing steps include: After three rounds of 2D convolution, instance normalization, and Leaky ReLU activation, feature matrices at three different levels are obtained; The feature matrices of three different levels are concatenated in the feature dimension to obtain concatenated multi-layer features; Through one-dimensional convolution and instance normalization, the concatenated multi-layer features are reduced in dimension to obtain fused multi-layer features. Add the fused multi-layer features and the third-layer features in the feature dimension and perform maximum pooling in the neighborhood to obtain the most significant features; Concatenate the most significant features with the concatenated multi-layer features in the feature dimension; The output features of the Setconv++ layer are obtained by using two-dimensional convolution, instance normalization, and maximum pooling within the neighborhood.

3. The method for estimating scene flow from point cloud under feature self-supervision as claimed in claim 1, characterized in that: The transmission cost constructed by the cosine similarity between point features is used to construct weights, and then the weights are used to construct soft corresponding points; The calculation formula of the weight is: Among them, w ij represents the weight of each j-th point in the target point cloud in the process of constructing the soft corresponding point of the i-th point in the source point cloud at the next moment, T * is the optimal transmission matrix, index t is the k point in the target point cloud t The set of index values ​​corresponding to the points; The calculation formula of the soft corresponding point is: Among them, is the soft corresponding point of the i-th point in the source point cloud at the next moment, represents the jth point in the target point cloud.

4. The method for estimating scene flow from point cloud under feature self-supervision according to any one of claims 1 to 3, characterized in that: The steps of training the model using the self-supervised loss term to obtain the point feature model and the relevant parameters in the optimal transmission problem include: The confidence of the constructed soft corresponding points is calculated using the cosine similarity matrix and the weights used in the construction process of the soft corresponding points. According to the confidence of the soft corresponding point, the confidence loss term, the chamfer distance loss term, the influence weight of the neighbor point on the center point, and the influence weight of the center point on the overall flow smoothness loss term are calculated respectively; By influencing the weights, the flow smoothing loss term is calculated; Construct the overall training loss based on the chamfer distance loss, confidence loss and flow smoothness loss; Back-propagation is performed through the overall training loss to obtain the training point feature model and the relevant learnable parameters in the optimal transmission problem.

5. The method for estimating scene flow of point cloud under feature self-supervision as claimed in claim 4, characterized in that: The soft corresponding point confidence formula is: u i =max(∑ j∈index w ij S ij ,0) Among them, u i represents the confidence of the soft corresponding point of the i-th point in the source point cloud, S ij Represents the cosine similarity between the i-th point in the source point cloud and the j-th point in the target point cloud; The chamfer distance loss term L dist The formula is: Among them, m represents the number of points in the point cloud, PC t+1 represents a point cloud composed of the soft corresponding points of each point in the source point cloud at the next moment, PC t+1 Represents the point cloud at the next moment.

6. The method for estimating scene flow from point cloud under feature self-supervision as claimed in claim 5, characterized in that: The confidence loss term L dist The formula is: Where m represents the number of points in the point cloud, u i Represents the confidence of the soft corresponding point of the i-th point in the source point cloud.

7. The method for estimating scene flow from point cloud under feature self-supervision as claimed in claim 6, characterized in that: The influence weight includes the weight x of the flow smoothing loss at each point in the overall flow smoothing loss term and the weight x of the flow smoothing loss at each point in the overall flow smoothing loss term. The weight x of the flow smoothing loss at each point in the overall flow smoothing loss term is: x=softmax(u i ),i∈1,2,...,m The weight of the flow smoothing loss at each point in the overall flow smoothing loss term for: Among them, N(i) represents the neighboring point corresponding to the i-th point in the source point cloud.

8. The method for estimating point cloud scene flow under feature self-supervision as claimed in claim 7, characterized in that: The flow smoothing loss term L smooth_new for: Among them, k s Indicates the number of selected neighboring points, index s Represents the index of the neighboring points, and flow represents the preliminary estimated point cloud scene flow.

9. The method for estimating scene flow from point cloud under feature self-supervision as claimed in claim 7, characterized in that: The overall training loss is: L total =λL dist +μL conf +νL smooth_new Among them, λ, μ and γ are hyperparameters.

10. The method for estimating scene flow of point cloud under feature self-supervision according to claim 9, characterized in that: The corrected point cloud scene flow Flow rw for: Flow rw =P3·Flow (∞) +P1Flow alter +P2(valid·Flow (∞) +(1-valid)Flow alter ) Among them, P1 represents the mask matrix of the flow with a confidence in the interval [0,0.2] in the preliminary estimated point cloud scene flow, P2 represents the mask matrix of the flow with a confidence in the interval [0.2,0.8] in the preliminary estimated point cloud scene flow, and P3 represents the mask matrix of the flow with a confidence in the interval [0.8,1] in the preliminary estimated point cloud scene flow. (∞) Represents the point cloud scene flow after fine-tuning, Flow alter Represents the reconstructed point cloud scene flow, valid represents the proportion of the combination, and is obtained by setting the values ​​in the confidence u in the interval [0,0.2] to 0.