Unsupervised point cloud shape corresponding method and system based on conditional diffusion model
Through the point cloud shape correspondence method based on the conditional diffusion model, the point cloud encoder and pseudo-label generator training transformer are used to gradually optimize the point cloud shape correspondence, solving the problem of large-scale displacement matching, and achieving high-precision unsupervised point cloud shape correspondence.
Patent Information
- Application Number
- CN202510380422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
The existing unsupervised point cloud shape correspondence methods are difficult to accurately match non-rigid objects with large-scale displacements, and it is difficult to effectively guide the direction of multi-step optimization under unsupervised conditions.
Using a method based on the conditional diffusion model, the structure-aware point embedding is extracted through the point cloud encoder, and the conditional diffusion model of the transformer is trained using a reliable pseudo-label generator. Combining the initial flow and local costs, the point cloud shape correspondence is gradually optimized, and the conditional diffusion model is used for prediction.
It realizes the precise shape correspondence of non-rigid objects with large displacement under unsupervised conditions, improves matching accuracy and robustness, and demonstrates its superiority and generalization capabilities on multiple data sets.
Smart Images

Figure CN120339654A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically, to an unsupervised point cloud shape correspondence method and system based on a conditional diffusion model. Background Art
[0002] Point cloud shape correspondence aims to identify the dense mapping between two non-rigid point clouds with deformable shapes. Shape correspondence is crucial for various practical applications, such as articulated motion transfer and shape editing. However, non-rigid objects exhibit significant motion variations. These motions lead to large displacements between corresponding parts of different shapes. In addition, as the original representation in three-dimensional space, the sparsity and irregularity of point clouds bring serious local noise, resulting in neighborhood perturbations.
[0003] To address the above challenges, various point cloud shape correspondence methods have been developed. To reduce resource burdens, researchers have increasingly focused on unsupervised methods that utilize unlabeled data for model training and inference. Current unsupervised point cloud shape correspondence methods achieve matching results by directly and independently calculating the similarity scores between point embeddings. This one-step matching method shows improved performance in regions with slight motions. However, the one-step matching mode may be difficult to accurately match corresponding points with large displacements and can only provide an approximate estimate.
[0004] By studying previous point-based shape correspondence methods, we found that to achieve more accurate shape correspondence, two key problems need to be solved: 1) How to handle large motion displacements in shape pairs? The accuracy bottleneck of most existing methods lies in the inability to accurately predict the correspondences of parts with large motion displacements, resulting in the accumulation of a large number of errors. Therefore, we should design a coarse-to-fine correspondence prediction method to replace the one-step method paradigm and use multi-step optimization to gradually converge to the accurate corresponding points. 2) How to guide the direction of multi-step optimization under unsupervised conditions? In the absence of ground truth labels to accurately guide the optimization direction, it is particularly crucial to design a selection strategy to identify high-quality landmark points as the guidance for the multi-step optimization direction. Summary of the Invention
[0005] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide an unsupervised point cloud shape correspondence method and system based on a conditional diffusion model.
[0006] In a first aspect, the purpose of the present invention can be achieved by the following technical solution: An unsupervised point cloud shape correspondence method based on a conditional diffusion model, the method comprising the following steps:
[0007] Obtain unsupervised point cloud data, input the unsupervised point cloud data into a pre-established point cloud encoder, and output encoded point cloud data;
[0008] Based on the similarity of point cloud features of unsupervised point cloud data, a reliable pre-established pseudo-label generator trains a pre-established Transformer-based conditional diffusion model by filtering reliable pseudo-labels, and outputs a trained Transformer-based conditional diffusion model;
[0009] Input the encoded point cloud data into the pre-established Transformer-based conditional diffusion model for prediction to obtain an unsupervised point cloud shape correspondence result.
[0010] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the pre-established point cloud encoder adopts a Transformer architecture, and uses the alignment between the input invariance of the Transformer and the disorder of the point cloud distribution to extract structure-aware point embeddings, and the architecture includes four Transformer blocks.
[0011] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the pre-established Transformer-based conditional diffusion model uses the initial flow X init and the local cost C as conditions. For any point s i ∈S, use the similarity matrix M to identify the point t j ∈T with the highest similarity as the initial corresponding point. The initial flow x i ∈X init measures the displacement of t j relative to s i , that is, x i =t j -s i . The local cost c i ∈C aggregates the similarity and relative displacement of the k nearest neighbors of t j in T, and is calculated as follows:
[0012]
[0013] where m ik ∈M, N T (t j ) represents the k nearest neighbors of t j in the target point cloud T. The condition K is composed of the concatenation of the initial flow X init and the local cost C: K = Concat(X init , C).
[0014] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the pre-established Transformer-based conditional diffusion model diffuses the initial noise flow according to the following formula:
[0015]
[0016] Among them, β t is the predefined variance at time step t. During time step t, the noise stream X t and the conditional K are input into the Transformer-based conditional diffusion model to obtain the flow residual Then, the initial flow X init and the flow residual are added together to generate the denoised noise stream X t-1 at the next time step t - 1, where
[0017] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the structure of the pre-established Transformer-based conditional diffusion model:
[0018] The noise stream X t and the conditional K are concatenated and input into the conditional diffusion model. The time step t is introduced using the adaLN-Zero method, and scaling and shifting are applied before and after the Transformer block and the feed-forward network FFN. The per-dimension scales α1, α2, γ1, γ2 and the shift parameters β1, β2 are regressed from the embedding vector of t, and the calculation method is as follows:
[0019] (α1, α2, γ1, γ2), (β1, β2) = MLP(Embed(t))
[0020] where MLP is a multi-layer perceptron and Embed is the encoding of the time step.
[0021] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the pre-established reliable pseudo-labels utilize the information in the similarity matrix to generate reliable pseudo-labels, model the uncertainty of each point through the variance value of the similarity vector. For each source point s i , the total variance σ 2 (s i ) is calculated as:
[0022] σ 2 (s i ) = Tr(Σ(m i ))
[0023] where Σ(m i ) is the covariance matrix of m i ∈M, Tr(·) represents the trace of the matrix, and the variance σ 2 is used to weight the denoising loss.
[0024] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the pre-established reliable pseudo-labels minimize the similarity cost by using a regularization term through a similarity matrix to obtain an optimal transport plan Process:
[0025] Given the target point representation Map the source point representation to the target point, denoted as X ot , and the formula is as follows:
[0026]
[0027] where H(x) = -∑ ij X ij logX ij is the entropy function, ∈ is a coefficient, and the Sinkhorn algorithm is used to obtain an approximate solution ∈ for the optimal transport in subsequent calculations, where ∈ > 0. By performing a soft-argmax operation on the approximate solution, the pseudo-point flow from the source point cloud to the target point cloud is identified Then, a recycling consistency matching strategy is used to filter out the point pairs that mutually predict the best K c corresponding point pairs to implement the training of the pre-established transformer-based conditional diffusion model
[0028] In a second aspect, to achieve the above object, the present invention discloses an unsupervised point cloud shape correspondence system based on a conditional diffusion model, including:
[0029] A data encoding module, configured to obtain unsupervised point cloud data, input the unsupervised point cloud data into a pre-established point cloud encoder, and output encoded point cloud data;
[0030] A model training module, configured to train a pre-established transformer-based conditional diffusion model through a pre-established reliable pseudo-label generator by filtering reliable pseudo-labels based on the similarity of the point cloud features of the unsupervised point cloud data, and output a trained transformer-based conditional diffusion model;
[0031] A prediction module, configured to input the encoded point cloud data into a pre-established transformer-based conditional diffusion model for prediction to obtain an unsupervised point cloud shape correspondence result
[0032] In another aspect of the present invention, to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, it adopts an unsupervised point cloud shape correspondence method based on a conditional diffusion model as described above
[0033] In another aspect of the present invention, in order to achieve the above object, a computer-readable storage medium is disclosed. A computer program is stored in the computer-readable storage medium. When the computer program is loaded and executed by a processor, an unsupervised point cloud shape correspondence method based on a conditional diffusion model as described above is adopted.
[0034] Advantages of the present invention:
[0035] The present invention conditions on the initial correspondence and local structure information through a conditional diffusion model based on a transformer, and improves the initial rough prediction through a denoising function. A reliable pseudo-label generator filters reliable pseudo-labels according to the similarity of point cloud features for training the conditional diffusion model. A large number of experiments conducted on the SHREC'19 and TOSCA benchmarks demonstrate the superiority of DiffCorr. Cross-dataset experiments on SURREAL and SMAL further demonstrate its excellent generalization ability. Description of the drawings
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;
[0037] Figure 1 is a schematic diagram of the method flow of the present invention;
[0038] Figure 2 is a schematic diagram of the framework of the unsupervised point cloud shape correspondence method based on the conditional diffusion model of the present invention;
[0039] Figure 3 is a schematic diagram of the system structure of the present invention;
[0040] Figure 4 is a verification diagram related to the recognition effect of the present invention. Detailed implementation manners
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0042] Embodiment 1:
[0043] As Figure 1As shown, an unsupervised point cloud shape correspondence method based on a conditional diffusion model, the method comprising the following steps:
[0044] S101: Obtain unsupervised point cloud data, input the unsupervised point cloud data into a pre-established point cloud encoder, and output encoded point cloud data;
[0045] The pre-established point cloud encoder adopts a Transformer architecture, and uses the alignment between the input invariance of the Transformer and the disorder of the point cloud distribution to extract structure-aware point embeddings. Our point cloud encoder architecture is similar to HSTR and includes four Transformer blocks. For the shape correspondence task, we limit the receptive field of each point within the k-NN neighborhood in the attention mechanism. This strategy enhances local structure modeling while ensuring the ability of remote perception.
[0046] S102: Based on the similarity of the point cloud features of the unsupervised point cloud data, a pre-established reliable pseudo-label generator trains a pre-established Transformer-based conditional diffusion model by filtering reliable pseudo-labels, and outputs a trained Transformer-based conditional diffusion model;
[0047] The pre-established Transformer-based conditional diffusion model uses the initial flow X init and the local cost C as conditions. For any point s i ∈S, the most similar point t j ∈T can be identified as the initial corresponding point using the similarity matrix M. The initial flow x i ∈X init measures the displacement of t j relative to s i in 3D space, i.e., x i = t j - s i . The local cost c i ∈C aggregates the similarity and relative displacement of the k nearest neighbors of t j in T. The specific calculation is as follows:
[0048]
[0049] where m ik ∈M, N T (t j ) represents the k nearest neighbors of t j in the target point cloud T. Finally, our condition K is composed of the concatenation of the initial flow X init and the local cost C, i.e., K = Concat(X init , C).
[0050] During training, diffuse the initial noise flow according to the following formula:
[0051]
[0052] wherein, β t is the predefined variance at time step t. During the inference process, it is replaced with random Gaussian noise. At time step t, the noise stream X t and the condition K are input into the Transformer-based conditional diffusion model to obtain the flow residual Then, the initial flow X init and the flow residual are added together to generate the denoised noise stream X t-1 at the next time step t - 1, that is,
[0053] The structure of the conditional diffusion model. The noise stream X t and the condition K are concatenated and input into the conditional diffusion model. In addition, we introduce the time step t using the adaLN-Zero method, where scaling and shifting are applied before and after the Transformer block and the feed-forward network (FFN). Specifically, we regress the per-dimension scales (α1, α2, γ1, γ2) and the shift parameters (β1, β2) from the embedding vector of t. The calculation method is as follows:
[0054] (α1, α2, γ1, γ2), (β1, β2) = MLP(Embed(t))
[0055] where MLP is a multi-layer perceptron and Embed is the encoding of the time step.
[0056] Pre-established reliable pseudo-label generator: Another major challenge in the unsupervised point cloud shape correspondence task is the lack of annotation labels. To facilitate the training of the conditional diffusion model, we designed a reliable pseudo-label generator that makes full use of the information in the similarity matrix to generate reliable pseudo-labels. First, we model the uncertainty of each point through the variance value of the similarity vector. For each source point s i , the total variance σ 2 (s i ) can be calculated as:
[0057] σ 2 (s i ) = Tr(Σ(m i ))
[0058] where Σ(m i ) is the covariance matrix of m i ∈M, and Tr(·) represents the trace of the matrix. High variance indicates diverse or diffuse patterns, suggesting unreliable predictions. This uncertainty (variance) σ2 For weighting the denoising loss.
[0059] Due to significant noise and many-to-one matching in the similarity matrix directly computed from point embeddings, our initial goal was to model point cloud matching as an optimal transport problem to resolve the many-to-one mismatch. Given the target point representation We are interested in mapping the source point representation to the target points. We denote this mapping as X ot , and its similarity cost matrix is By minimizing the similarity cost with a regularization term, the optimal transport plan
[0060]
[0061] where H(x) = -∑ ij X ij log X ij is the entropy function and ∈ is a coefficient. However, we found that the hard labels from the optimal transport solution (∈ = 0) unexpectedly led to suboptimal or even degraded performance, possibly because solving the optimal solution amplified the noise in the similarity matrix. Therefore, we use the Sinkhorn algorithm to obtain an approximate solution of the optimal transport (∈ > 0) for subsequent calculations. By performing a soft-argmax operation on the approximate solution, the pseudo-point flow from the source point cloud to the target point cloud is identified Finally, a cyclic consistency matching strategy is designed to filter out the pairs of points that mutually predict as the best K c corresponding pairs This will supervise the training of the diffusion model.
[0062] S103: Input the encoded point cloud data into a pre-established conditional diffusion model based on a transformer for prediction to obtain unsupervised point cloud shape correspondence results.
[0063] Embodiment 2: Second aspect, as Figure 3 shown, to achieve the above object, the present invention discloses an unsupervised point cloud shape correspondence system based on a conditional diffusion model, including:
[0064] A data encoding module 11, configured to obtain unsupervised point cloud data, input the unsupervised point cloud data into a pre-established point cloud encoder, and output encoded point cloud data;
[0065] A model training module 12, configured to, based on the similarity of the point cloud features of the unsupervised point cloud data, train a pre-established reliable pseudo-label generator to filter reliable pseudo-labels and then train a pre-established conditional diffusion model based on a transformer, and output a trained conditional diffusion model based on a transformer;
[0066] A prediction module 13 for inputting the encoded point cloud data into a pre-established transformer-based conditional diffusion model for prediction to obtain an unsupervised point cloud shape correspondence result.
[0067] As Figure 4 shown, SMAL / TOSCA and SURREAL / SHREC’19 are dataset names, Source is the source point cloud, DPC, SE-ORNet, HSTR, and Ours are the displays of the matching correspondence results of three state-of-the-art methods and the method of the present invention respectively, and Ground Truth refers to the standard matching result for comparing the effectiveness of the methods.
[0068] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in the computer storage medium to implement the above method.
[0069] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above-mentioned method. The storage medium can be any combination of one or more computer-readable media. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device.
[0070] In the description of this specification, the description referring to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0071] The above shows and describes the basic principles, main features, and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.
Claims
1. An unsupervised point cloud shape correspondence method based on conditional diffusion models, characterized in that, The method includes the following steps: Obtain unsupervised point cloud data, input the unsupervised point cloud data into a pre-established point cloud encoder, and output encoded point cloud data; Based on the similarity of the point cloud features of the unsupervised point cloud data, a pre-established reliable pseudo-label generator trains a pre-established transformer-based conditional diffusion model by filtering reliable pseudo-labels, and outputs a trained transformer-based conditional diffusion model; Input the encoded point cloud data into a pre-established transformer-based conditional diffusion model for prediction to obtain an unsupervised point cloud shape correspondence result.
2. The unsupervised point cloud shape correspondence method based on a conditional diffusion model according to claim 1, wherein The pre-established point cloud encoder adopts a Transformer architecture and uses the alignment between the input invariance of the Transformer and the disorder of the point cloud distribution to extract structure-aware point embeddings. The architecture includes multiple Transformer blocks.
3. An unsupervised point cloud shape correspondence method based on a conditional diffusion model according to claim 1, characterized in that, The pre - established transformer - based conditional diffusion model uses the initial flow X init and the local cost C as conditions. For any point s i ∈S, the point t j ∈T with the highest similarity is identified as the initial corresponding point using the similarity matrix M. The initial flow x i ∈X init measures the displacement of t j relative to s i , that is, x i =t j -s i . The local cost c i ∈C aggregates the similarities and relative displacements of the k nearest neighbors of t j in T and is calculated as follows: Among them, m ik ∈M,N T (t j ) represents the target point cloud T in t j The k nearest neighbors of the initial flow X init and the local cost C cascade composition: K = Concat (X init ,C).
4. An unsupervised point cloud shape correspondence method based on a conditional diffusion model according to claim 3, characterized in that The pre-established transformer-based conditional diffusion model diffuses the initial noise flow according to the following formula: wherein, β t is the predefined variance at time step t, and within time step t, the noise stream X t and the conditional K are input into the Transformer-based conditional diffusion model to obtain the flow residual Then, the initial flow X init and the flow residual are added together to generate the denoised noise stream X t-1 at the next time step t - 1, wherein, 5. The unsupervised point cloud shape correspondence method based on a conditional diffusion model according to claim 4, wherein The structure of the pre-established transformer-based conditional diffusion model: Noise flow X t and condition K are concatenated and input into the conditional diffusion model. The time step t is introduced using the adaLN-Zero method, and scaling and shifting are applied before and after the Transformer block and the feed-forward network FFN. The per-dimension scales α1, α2, γ1, γ2, and the shift parameters β1, β2 are regressed from the embedding vector of t, and the calculation method is as follows: (α1,α2,γ1,γ2),(β1,β2)=MLP(Embed(t)) where MLP is a multi-layer perceptron and Embed is the encoding of the time step.
6. The unsupervised point cloud shape correspondence method based on a conditional diffusion model according to claim 1, wherein, The pre-established reliable pseudo-labels utilize the information in the similarity matrix to generate reliable pseudo-labels, model the uncertainty of each point through the variance value of the similarity vector, and for each source point s i , the total variance σ 2 (s i ) is calculated as: σ 2 (s i )=Tr(Σ(m i )) where, Σ(m i ) is the covariance matrix of m i ∈M, Tr(·) represents the trace of a matrix, and the variance σ 2 is used to weight the denoising loss.
7. A method for unsupervised point cloud shape correspondence based on conditional diffusion model according to claim 6, characterized in that, The pre-established reliable pseudo-labels minimize the similarity cost by using a regularization term with the similarity matrix to obtain an optimal transport plan Process: Given target point representation Map the source point representation to the target point, denoted as X ot , and the formula is as follows: where \(H(x)=-\sum\) ij X ij \(\log X\) ij is the entropy function, \(\in\) is the coefficient, and the Sinkhorn algorithm is used to obtain an approximate solution \(\in\) of the optimal transport for subsequent calculations, where \(\in>0\). By performing a soft-argmax operation on the approximate solution, the pseudo-point flow from the source point cloud to the target point cloud is identified Then, a cycle-consistency matching strategy is used to filter out the pairs of points that mutually predict as the best K c corresponding point pairs to implement the training of a pre-established transformer-based conditional diffusion model 8. An unsupervised point cloud shape correspondence system based on a conditional diffusion model, characterized in that, It includes: A data encoding module for obtaining unsupervised point cloud data, inputting the unsupervised point cloud data into a pre-established point cloud encoder, and outputting encoded point cloud data; A model training module for training a pre-established transformer-based conditional diffusion model by a pre-established reliable pseudo-label generator filtering reliable pseudo-labels based on the similarity of the point cloud features of the unsupervised point cloud data, and outputting a trained transformer-based conditional diffusion model; A prediction module for inputting the encoded point cloud data into a pre-established transformer-based conditional diffusion model for prediction to obtain an unsupervised point cloud shape correspondence result.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it adopts an unsupervised point cloud shape correspondence method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program therein, characterized in that, When the computer program is loaded and executed by the processor, it adopts an unsupervised point cloud shape correspondence method according to any one of claims 1 to 7.