An identity recognition method and system based on 3D tooth point cloud multi-feature fusion

By constructing an identity recognition method based on multi-feature fusion of 3D tooth point clouds and utilizing color and structural feature extraction and feature fusion technology, the problems of computational complexity and equipment differences in existing methods are solved, achieving efficient and accurate identity recognition.

CN119600648BActive Publication Date: 2025-09-05UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411637433.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-09-05
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

Existing identity recognition methods based on 3D tooth point clouds have high computational complexity, low recognition efficiency, and are affected by device differences, resulting in insufficient recognition accuracy.

Method used

An identity recognition method based on multi-feature fusion of 3D tooth point cloud is adopted. Through the color feature extraction module, structural feature extraction module, feature fusion module, local sampling feature extraction module, feature aggregation module, coarse correspondence prediction module, point cloud decoding module and fine correspondence prediction module, a fast end-to-end identity recognition network is constructed, and the point correspondence relationship is established using the key points at the neighborhood level and the points at the single point level of the tooth point cloud.

Benefits of technology

It improves the accuracy and efficiency of identity recognition, enhances the stability and robustness of recognition, and reduces the impact of device differences on recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600648B_ABST
    Figure CN119600648B_ABST
Patent Text Reader

Abstract

The present invention discloses an identity recognition method and system based on 3D tooth point cloud multi-feature fusion. The method comprises: collecting a 3D tooth point cloud of an identity to be recognized; inputting the 3D tooth point cloud of the identity to be recognized and the 3D tooth point cloud of each registered identity in an identity database into a trained tooth multi-feature fusion identity recognition backbone network, outputting a registration result between the 3D tooth point cloud of the identity to be recognized and the 3D tooth point cloud of each registered identity; the backbone network comprises: a color feature extraction module, a structural feature extraction module, a feature fusion module, a local sampling feature extraction module, a feature aggregation module, a coarse correspondence prediction module, a point cloud decoding module, a fine correspondence prediction module, and a prediction registration module; and the identity information of the 3D tooth point cloud of the registered identity that is optimally registered with the 3D tooth point cloud of the identity to be recognized is used as the identity identifier of the 3D tooth point cloud of the identity to be recognized. The present invention can perform identity recognition based on 3D tooth point clouds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of identity recognition technology, and in particular to an identity recognition method and system based on 3D tooth point cloud multi-feature fusion. Background Art

[0002] The 3D tooth point cloud data collected by intraoral scanning instruments can accurately reflect the subtle differences and uniqueness of teeth, including geometric and structural features such as dental arch shape, interdental arrangement, tooth unevenness, and tooth-gum boundary. Compared with recognition using 2D images, 3D point cloud recognition can obtain richer tooth anatomical structure features, is less susceptible to factors such as lighting, angle, and occlusion, and has stronger anti-interference capabilities and higher accuracy. In addition, with the development of digital technology, the acquisition of 3D oral data is becoming increasingly convenient and popular, laying a solid foundation for the construction of a large-scale 3D tooth point cloud identity database, which in turn can promote the continuous improvement of identity recognition technology based on 3D tooth point clouds.

[0003] Identity recognition technologies based on 3D tooth point clouds require processing the collected 3D tooth point clouds to extract representative tooth features. These features should reflect the uniqueness of the teeth and form the basis for identification. Currently, mainstream 3D tooth point cloud-based identity recognition methods often use iterative algorithms to directly perform 3D-3D registration of the tooth point clouds, and then perform identity recognition based on the registration results. However, the registration process involved in these methods requires multiple rounds of iteration to find the optimal matching relationship between the two point clouds. This iterative process is often computationally complex, as it requires continuous creation of matching point pairs and pose adjustments in 3D space to minimize the alignment error between the two point clouds. This process involves numerous matrix operations and optimization iterations, resulting in low computational efficiency, especially with large point cloud data volumes. Furthermore, these mainstream methods typically directly utilize spatial distance differences between point clouds for registration. When tooth point clouds are collected from different devices, this often leads to significant deviations in the coordinate systems of the point clouds, complicating registration and affecting the accuracy of the recognition results. Furthermore, these mainstream methods rely solely on the spatial coordinate information of the tooth point clouds. The existing methods have obvious deficiencies in the robustness and accuracy of registration in actual identity recognition tasks. Summary of the Invention

[0004] The present invention provides an identity recognition method and system based on 3D tooth point cloud multi-feature fusion to solve the problems existing in the above-mentioned prior art. The technical solution is as follows:

[0005] On the one hand, an identity recognition method based on multi-feature fusion of 3D tooth point cloud is provided, comprising:

[0006] S1, collect the 3D tooth point cloud of the identity to be identified;

[0007] S2. Input the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity in the identity database into the trained identity recognition backbone network with multi-feature fusion of teeth, and output the registration result between the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity;

[0008] The tooth multi-feature fusion identity recognition backbone network includes: a color feature extraction module, a structural feature extraction module, a feature fusion module, a local sampling feature extraction module, a feature aggregation module, a coarse correspondence prediction module, a point cloud decoding module, a fine correspondence prediction module, and a prediction registration module;

[0009] S3. Using the identity information of the 3D tooth point cloud of the registered identity that is optimally aligned with the 3D tooth point cloud of the identity to be identified as the identity identifier of the 3D tooth point cloud of the identity to be identified.

[0010] In another aspect, an identity recognition system based on 3D tooth point cloud multi-feature fusion is provided, the system comprising:

[0011] The acquisition module is used to collect the 3D tooth point cloud of the identity to be identified;

[0012] A registration module is configured to input the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity in the identity database into the trained tooth multi-feature fusion identity recognition backbone network, and output a registration result between the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity;

[0013] The tooth multi-feature fusion identity recognition backbone network includes: a color feature extraction module, a structural feature extraction module, a feature fusion module, a local sampling feature extraction module, a feature aggregation module, a coarse correspondence prediction module, a point cloud decoding module, a fine correspondence prediction module, and a prediction registration module;

[0014] The identification module is configured to use the identity information of the 3D tooth point cloud of the registered identity that is optimally aligned with the 3D tooth point cloud of the identity to be identified as the identity identifier of the 3D tooth point cloud of the identity to be identified.

[0015] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned identity recognition method based on multi-feature fusion of 3D tooth point cloud.

[0016] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned identity recognition method based on multi-feature fusion of 3D tooth point cloud.

[0017] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0018] 1) A color feature extraction module and a structural feature extraction module for tooth point clouds were designed. These modules can extract the effective color feature representation and structural feature representation of a known point based on the color information and spatial coordinate information of the known point and its neighboring points. The extracted color features and structural features are then fused to enrich the feature information that can be used in the establishment of corresponding point relationships in the tooth point cloud before the identification stage from multiple dimensions, thereby improving the accuracy of identification.

[0019] 2) A fast end-to-end identity recognition network based on tooth point cloud was constructed. By using the key points at the neighborhood level and the points at the single point level of the tooth point cloud in a coarse-to-fine manner to establish point correspondences, the network can find matching point pairs more accurately, thereby improving the prediction efficiency of matching point pairs while ensuring the prediction accuracy of matching point pairs, thereby improving the efficiency of the identity recognition task.

[0020] 3) Before establishing corresponding point relationships within the network, a designed local sampling feature extraction module extracts features from the tooth point cloud. These extracted features serve as the basis for establishing matching point pairs. These extracted features are robust and do not change due to significant changes in the tooth point cloud's 3D coordinate system caused by changes in the acquisition equipment. This improves the stability of identity recognition performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 This is a flow chart of an identity recognition method based on 3D tooth point cloud multi-feature fusion provided by an embodiment of the present invention;

[0023] Figure 2 This is a schematic diagram of the backbone network structure for tooth multi-feature fusion identity recognition provided by an embodiment of the present invention;

[0024] Figure 3 This is an overall block diagram of the identity recognition method based on multi-feature fusion of teeth provided by an embodiment of the present invention;

[0025] Figure 4 Schematic diagram of the color feature extraction module provided by an embodiment of the present invention;

[0026] Figure 5 is a color frequency histogram provided by an embodiment of the present invention;

[0027] Figure 6 Schematic diagram of the structure of the structural feature extraction module provided by an embodiment of the present invention;

[0028] Figure 7 is a structural frequency histogram provided by an embodiment of the present invention;

[0029] Figure 8 Schematic diagram of the structure of a local sampling feature extraction module provided by an embodiment of the present invention;

[0030] Figure 9 This is a schematic diagram of the structure of a feature aggregation module provided by an embodiment of the present invention;

[0031] Figure 10 2 is a schematic diagram of the structure of a coarse correspondence prediction module provided by an embodiment of the present invention;

[0032] Figure 11 Schematic diagram of the structure of a point cloud decoding module provided by an embodiment of the present invention;

[0033] Figure 12 2 is a schematic diagram of the structure of a precise correspondence prediction module provided by an embodiment of the present invention;

[0034] Figure 13 Schematic diagram of the prediction and registration module structure provided by an embodiment of the present invention;

[0035] Figure 14 This is a block diagram of an identity recognition system based on 3D tooth point cloud multi-feature fusion provided by an embodiment of the present invention;

[0036] Figure 15 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0037] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0038] The embodiment of the present invention provides an identity recognition method based on 3D tooth point cloud multi-feature fusion, which can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of an identity recognition method based on 3D tooth point cloud multi-feature fusion is shown. The processing flow of this method may include the following steps:

[0039] S1, collect the 3D tooth point cloud of the identity to be identified;

[0040] The embodiment of the present invention collects the 3D tooth point cloud of the identity to be identified through an intraoral scanner.

[0041] S2. Input the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity in the identity database into the trained identity recognition backbone network with multi-feature fusion of teeth, and output the registration result between the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity;

[0042] like Figure 2 As shown, the tooth multi-feature fusion identity recognition backbone network includes: color feature extraction module, structural feature extraction module, feature fusion module, local sampling feature extraction module, feature aggregation module, coarse correspondence prediction module, point cloud decoding module, fine correspondence prediction module, and prediction registration module;

[0043] like Figure 3 As shown, the training process of the tooth multi-feature fusion identity recognition backbone network includes:

[0044] 1) Tooth point cloud data collection.

[0045] The embodiment of the present invention collects 3D tooth point cloud data of multiple individuals through an intraoral scanner, and records the identity information of the collected individuals in text form.

[0046] 2) Combining manual adjustments with the optimization-based registration algorithm ICP, we obtain a transformation matrix RT that can register two tooth point clouds scanned from the same person to completely overlap, and calculate the true matching point pairs between the two tooth point cloud matching pairs.

[0047] The specific processing process of step 2) is prior art and will not be described in detail here.

[0048] 3) Dataset division.

[0049] Consider the tooth point cloud matching pairs obtained in step 2 as a single sample. Divide all collected and constructed samples into a training set and a validation set in a 4:1 ratio to form a complete dataset. The algorithm will be trained on the training set and validated on the validation set. The model with the best performance on the validation set will be selected as the final model.

[0050] 4) Construction of identity database.

[0051] After training, any one of the tooth point clouds in all sample pairs is taken out and stored in the identity database for identity information registration, and the remaining tooth point clouds are used as test set data to test the performance of the model.

[0052] like Figure 4 As shown, the color feature extraction module is used to extract the color feature information of the 3D tooth point cloud. The specific extraction process includes:

[0053] Assume that the 3D tooth point cloud is P∈R N×6 , where N represents the number of points in P, and the number 6 indicates that each point in P has 3D spatial coordinate information and 3D color information;

[0054] Define any point in P as the center point P center and with the preset R color (such as 0.5mm) as the sphere radius, and define other points within this range as neighboring points P neighbor_1 ~P neighbor_n , the center point P center Perform difference calculation with the RGB color channel values ​​of each adjacent point: R diff =R center -R neighbor , G diff =G center -G neighbor , B diff =B center -B neighbor The calculated color difference is divided into certain intervals and the R diff , G diff and B diff The result (for example, divided into M color = 33 intervals, of which the intervals 1-11 are for statistical R diff The result of the 12th to 22nd interval is to count G diff The result of the 23rd to 33rd interval is to count B diff These intervals represent different levels of color difference, and the sum of the statistical results of all intervals of the three channels constitutes the center point P center Initial color frequency histogram ICH(P center ),like Figure 5 As shown, combined with the actual values ​​of the tooth color channel difference intervals, the embodiment of the present invention stipulates that the statistical value range of the three types of intervals is between -100 and 100;

[0055] Then, taking each adjacent point as the center, repeat the above operation to obtain the color frequency histogram ICH(P neighbor_1 )~ICH(P neighbor_n );

[0056] According to the relationship between each neighboring point and the center point P centerThe distance d is used, and 1 / d is used as the weight. The color frequency histogram results of all neighboring points are added to the center point P. center On the color frequency histogram, the center point P is obtained center Final color frequency histogram

[0057] The final color frequency histogram FCH(P center ) is converted to an M color dimensional vector, as the center point P center Color feature information;

[0058] Take all points in P as the center point and perform the above operation to obtain the color feature information of all points in the point cloud, which is recorded as The 3 indicates that the first three columns of information of the feature are the three-dimensional spatial coordinates of the point.

[0059] like Figure 6 As shown, the structural feature extraction module is used to extract the structural feature information of the 3D tooth point cloud. The specific extraction process includes:

[0060] Calculate the normal vector of each point in the point cloud based on the three-dimensional spatial coordinate information of each point in the point cloud and its surrounding points;

[0061] Define any point in P as the center point P center and with the preset R struct (such as 0.5mm) as the sphere radius, and define other points within this range as neighboring points P neighbor_1 ~P struct_n , describing the center point P center and the geometric features of each neighboring point, where the center point P is described as follows center and P neighbor_1 The geometric structure characteristics of the two points (the center point P center and other neighboring points have similar geometrical characteristics): Let vector d = P center -P neighbor_1 , represents the center point P center and P neighbor_1 The connecting vector between two points, the center point P center The normal vector is defined as P center , P neighbor_1 The normal vector is defined as P neighbor_1 , define α to represent n neighbor_1 The angle between and d is calculated as Measures the projection relationship of the normal vector of the neighborhood point in the direction of the target point; define φ to represent the center point P center Normal vector n centerThe angle between d and d is calculated as Reflects the neighborhood point P neighbor_1 Relative to the center point P center Normal vector n center The projection angle of n center and n neighbor_1 The angle between them is calculated as follows: θ = arccos(n center ·n neighbor_1 ), which measures the degree of alignment of the normal vectors of two points. These parameters together describe the center point P center and neighboring point P neighbor_1 The geometrical structural characteristics between

[0062] Construct the center point P center The initial structure frequency histogram ISH(P center ),like Figure 7 As shown, the initial structure frequency histogram ISH (P center )Total M struct intervals, used to count the center point P respectively center and P neighbor_1 The results of the three parameters α, φ and θ are used as feature information. The value range of α is [-1, 1]. When α = 1, it means that the normal vector of the neighboring point is completely parallel to the normal vector of the center point and has the same direction; when α = -1, it means that they are in opposite directions; when α = 0, it means that the two are perpendicular; the value range of φ is also [-1, 1]. When φ = 1, the normal vector of the center point is consistent with the direction of the line between the center point and the neighboring point; when φ = -1, the normal vector of the center point is opposite to the direction of the line between the center point and the neighboring point; when φ = 0, the normal vector of the center point is perpendicular to the line between the center point and the neighboring point; the value range of θ is [-π, π]. When θ = 0, it means there is no rotation; when θ = π or θ = -π, it means a rotation of 180 degrees, reflecting a larger angular offset;

[0063] For the center point P center The other neighboring points of are then used to construct the structural frequency histogram, and then added to the center point P in a weighted manner. center The initial structure frequency histogram FSH(P center ) on the center point P center Final structural frequency histogram

[0064] The center point P center The structural frequency histogram FSH(P center ) is converted to an M struct dimensional vector, as the center point P center Structural feature information;

[0065] Take all points in P as the center point and perform the above operation to obtain the structural feature information of all points in the point cloud, which is recorded as The 3 indicates that the first three columns of information of the feature are the three-dimensional spatial coordinates of the point.

[0066] The feature fusion module extracts M from each point color dimensional color feature information and M struct The 3D structural feature information is spliced ​​with the 3D spatial coordinate information of each point, and the output is a point cloud containing 3+M color +M struct The fusion result of dimensional feature information

[0067] like Figure 8 As shown, the local sampling feature extraction module first performs the feature fusion module output Perform a downsampling operation to obtain a set of sparse points as neighborhood-level key points, reducing the number of points in the point cloud;

[0068] In order to calculate the features of these key points, a neighborhood with K as the neighborhood radius is established for each key point, which is used to search for original points close to these key points in the point cloud before downsampling. Then, according to the distance relationship between the searched original points and the corresponding key points, the features of these original points are aggregated to the corresponding key points as the local feature information F' of the corresponding key points, and the key points with local feature information (P', F') are output, where P' represents the three-dimensional spatial coordinates of the key points.

[0069] The features extracted by the local sampling feature extraction module are highly stable. Even if the acquisition device is changed, resulting in a change in the initial spatial position of the model, these features will not change.

[0070] like Figure 9 As shown in the figure, the feature aggregation module relies on the attention mechanism to strengthen and aggregate the local feature information in the neighborhood-level key points, integrate the global and local context information of each point in the input key points, and enable the features of the neighborhood-level key points of the two to-be-registered point clouds to interact, thereby improving the subsequent point cloud registration accuracy. It includes two self-attention modules and one cross-attention module. The processing process is as follows:

[0071] First, the features of the input neighborhood-level key points are enhanced by the self-attention module. The self-attention mechanism contained in the self-attention module generates the query matrix Q, key matrix K and value matrix V by linearly transforming the input feature matrix F':

[0072] Q=F'W Q,K=F'W K ,V=F'W V

[0073] Where W Q 、W K 、W V is a learnable linear transformation matrix;

[0074] Next, the attention score matrix A is obtained by calculating the dot product of the query matrix and the key matrix and normalizing it:

[0075]

[0076] where d k is the size in the key matrix, which is used for normalization to prevent the value from being too large. The attention score matrix A is used to perform weighted summation on the value matrix V to obtain the aggregated feature matrix F":

[0077] F″=A·V

[0078] In this way, the features of each point not only contain its own information, but also aggregate the global context information related to it;

[0079] Subsequently, the result is input into the cross attention module to integrate the feature information between key points at different point cloud neighborhood levels. The feature matrices of the two point cloud key points processed by the self-attention module are F' X 'and F' Y ', the cross attention module then performs feature aggregation through the following steps:

[0080] Q X =F″ X W Q ,K Y =F″ Y W K ,V Y =F″ Y W V

[0081] Calculate F″ X and F″ Y The attention score matrix A between:

[0082]

[0083] Using the attention score matrix A X Y pair matrix V Y Perform weighted summation to obtain the enhanced feature matrix F″′ X :

[0084] F″′ X =AX Y.V Y

[0085] Calculate F″ X and F″ Y The cross attention is F″′ Y :

[0086] F″′ Y =A Y X·V X

[0087] Through the cross-attention module, the features of the two key points are mutually enhanced;

[0088] After the cross attention module, the feature representation capability is further enhanced by a self-attention calculation. Through the feature aggregation module, the features of each point in the output neighborhood-level key points not only integrate the geometric information of its neighborhood, but also fuse the features from the global context, and also interact with the key points of the point cloud to be registered with it. Finally, the key points of the registered point cloud X after feature aggregation are output. and the point cloud of the identity to be identified after feature aggregation Y Key Points

[0089] The features output by the feature aggregation module are more robust to the subsequent matching point establishment task.

[0090] like Figure 10 As shown, the coarse correspondence prediction module obtains the key points of the registered point cloud X after feature aggregation by calculating the inner product of the point cloud feature vector and the point cloud of the identity to be identified after feature aggregation Y Key Points A similarity matrix S' between them, each element in the similarity matrix S' represents the similarity between two key points;

[0091] After obtaining the similarity matrix S', a normalization algorithm is used to normalize the similarity matrix (the purpose of normalization is to adjust these similarity values ​​to a uniform range (usually 0 to 1) to ensure that the similarity values ​​between different point clouds can be compared on the same scale. In this way, the similarity evaluation between all point pairs is performed under the same standard, making the difference between high and low values ​​more reasonable and not causing the matching process to be biased towards a specific range of values). A matching probability matrix is ​​obtained, which represents the matching confidence of each pair of points between the key points of the neighborhood of the two point clouds;

[0092] Then, by setting a confidence threshold, key point pairs with higher confidence are screened out from the matching probability matrix as rough correspondence predictions;

[0093] Finally, a neighborhood-level matching point pair containing multiple coarse correspondences is output. This coarse correspondence set will be further processed in the subsequent fine correspondence prediction module to obtain more accurate point pair matching results.

[0094] like Figure 11 As shown, the point cloud decoding module decodes and recovers the rough neighborhood key point correspondences, refines them to the single point level, and uses the KP convolution decoder to convert the key points of the registered point cloud X that has undergone feature aggregation into And the identity point cloud to be identified after feature aggregation Y Key Points Decode to the same resolution as the point cloud before downsampling to obtain the decoded point cloud and This process involves a series of convolution operations to gradually restore the spatial distribution and geometric features similar to those of the point cloud before downsampling. The KP convolution decoder will gradually propagate the features of the neighborhood key points to the restored points. These propagated features will be combined with the geometric information of the point cloud before downsampling to generate a more fine-grained feature representation. The decoding process also involves a point-to-neighborhood level key point grouping strategy. The point cloud grouping module is used to group each point in the decoded point cloud according to its neighborhood level key points, and the points in the decoded point cloud are assigned to the nearest neighborhood level key points, and the grouping results of the registered point cloud X are output. And the grouping result of the identity point cloud Y to be identified For subsequent refined registration, the subsequent fine correspondence prediction module uses the grouping results obtained from the point cloud decoding module and the key point matching results obtained from the coarse correspondence prediction module to further refine the matching point pairs and obtain single-point level matching point pairs.

[0095] like Figure 12 As shown, the precise correspondence prediction module includes: similarity matrix calculation, normalization processing and confidence screening steps. First, a set of matching point pairs is selected from the neighborhood level matching point pairs, and then the point cloud group constructed by these two points is selected from the grouping results: Grouping results of registered point cloud X and point cloud of the identity to be identified Y The grouping results Next, the similarity matrix is ​​calculated for these two point clouds. The similarity matrix By calculating the inner product of all point feature vectors between two point cloud groups, each element in the matrix represents the similarity between the two points; Afterwards, the results are normalized to obtain a matching probability matrix between the two point cloud groups, where the value of each element in the matching probability matrix represents the matching confidence of any two points between the two point cloud groups. Afterwards, the matching point pairs between the two point cloud groups are screened out by setting a confidence threshold. Finally, all neighborhood-level matching point pairs and their corresponding point cloud grouping results are used to perform the above processing, and all matching point pair screening results are merged together to obtain single-point-level matching point pairs as the output result of the precise correspondence prediction module.

[0096] like Figure 13 As shown, the prediction registration module calculates the optimal rotation and translation matrix based on the single-point level matching point pair obtained by the precise correspondence prediction module, and realizes the two tooth point clouds. and Specifically, the algorithm randomly selects a minimum number of point pairs, typically three, from the matched point pairs to estimate a candidate rigid transformation. This transformation matrix contains rotation and translation parameters used to align one point cloud to another. The estimated transformation is applied to all matched point pairs and evaluated to see if all point pairs remain consistent after the transformation. If the Euclidean distance between two points in a point pair is less than a predefined threshold after the transformation is applied, the feature point pair is considered an "inlier." The algorithm counts all inliers and repeats the above steps until the set number of inliers is met. In each iteration, the algorithm randomly selects a different point pair for transformation estimation. After multiple iterations, the transformation with the largest number of inliers is selected as the final estimated rigid transformation matrix. The final estimated rigid transformation matrix is ​​used to align the source point cloud to the target point cloud. This step achieves the final registration of the two point clouds, aligning them in the same coordinate system. In general, the function of this module is to use the corresponding point pairs output by the precise correspondence prediction module to find the optimal rigid transformation matrix between the two tooth point clouds, so as to achieve alignment between the two tooth point clouds and lay the foundation for subsequent identity recognition functions.

[0097] The main reason why the embodiments of the present invention can improve recognition efficiency and ensure recognition accuracy is that: in the corresponding point establishment stage, the point correspondence process from coarse to fine can quickly and accurately establish matching point pairs between the two tooth point clouds, and the two tooth point clouds only need to establish matching point pairs once during the entire registration process. Subsequently, the predictive registration module uses the information of these matching point pairs to calculate the rotation and translation matrices, thereby achieving the registration of the two tooth point clouds and finally obtaining the final identity recognition result. However, the registration methods involved in existing optimization-based identity recognition methods usually need to go through multiple rounds of iterative optimization processes, and each round of iterative optimization process requires re-establishing matching point pairs, and there are often erroneous matching point pairs in the established matching point pairs. Therefore, the recognition efficiency of existing optimization-based identity recognition methods is very poor.

[0098] The tooth multi-feature fusion identity recognition backbone network, the total loss function L in the training process is the coarse corresponding prediction loss L c And the corresponding prediction loss L f The weighted sum of:

[0099] L=L c +λL f

[0100] Among them, λ is the weight parameter used to balance the two loss terms. Through this combination, the model can be optimized at both coarse and fine scales, ensuring that the model can learn the precise correspondence between point clouds from coarse to fine. At the coarse scale, the model is supervised to learn the reliable correspondence between the key points obtained by downsampling, while at the fine scale, the model further optimizes the matching accuracy of each point at the single point level. Through the combination of these two parts, the final model can achieve high-precision point cloud registration;

[0101] The coarse corresponding prediction loss L c It is used to supervise the matching results of neighborhood-level key points. Its core idea is to use the local overlap ratio as a weighting scheme to supervise the coarse correspondence. Given a pair of neighborhood-level key points P' X (i') and P' Y (j'), their corresponding neighborhood point grouping results are and calculate Zhongyu The proportion of points with corresponding relationships is:

[0102]

[0103] in, Is a known quantity, which represents the real transformation matrix that can completely align the X point cloud and Y point cloud input into the network, τ P represents the distance threshold;

[0104] calculate Zhongyu The proportion of points with corresponding relationships is:

[0105]

[0106] According to the above two formulas, we get the weighted matrix W':

[0107] W'(i',j')=min(r(i',j'),r(j',i')),i'≤n'∧j'≤m'

[0108] Among them, n' and m' respectively represent the total number of neighborhood-level key points of the two point clouds, and each element in W'(i', j') represents the weighted value of the key point pair (i', j'). This weighted matrix is ​​used to reflect the geometric overlap between the key point pairs. The larger the weight, the more reliable the matching key point pair. c When , the cross entropy loss is used to measure the difference between the similarity matrix S' output by the model and the weight matrix W':

[0109]

[0110] Among them, S'(i',j') is the confidence score of the key point pair (i',j') recorded in the similarity matrix, which indicates the similarity between the corresponding key point pairs. This loss function supervises the model learning to predict more accurate coarse correspondences by minimizing the difference between the model's predicted point coarse matching results and the weighted matrix W'.

[0111] The precision corresponds to the prediction loss L f Used to supervise fine matching at the single point level, it optimizes based on coarse correspondences:

[0112]

[0113] in, and They represent the two point groups formed by the lth neighborhood level key point matching point pair (i', j'), It is a binary matrix that represents the correct point pair matching relationship between the two point groups. For the matching point pairs, its value is 1, otherwise it is 0. The precision correspondence prediction loss L f Cross entropy loss is also used to measure the confidence matrix of the model output With binary matrix The differences between:

[0114]

[0115] in, Represents the similarity matrix between the points in the two point groups formed by the I-th neighborhood-level key point matching point pair (i', j') predicted by the model. This loss function guides the model to learn more accurate single-point level point matching relationships by minimizing the error at the single-point level.

[0116] S3. Using the identity information of the 3D tooth point cloud of the registered identity that is best aligned with the 3D tooth point cloud of the identity to be identified (ie, with the lowest alignment error) as the identity identifier of the 3D tooth point cloud of the identity to be identified.

[0117] like Figure 14 As shown, an embodiment of the present invention further provides an identity recognition system based on 3D tooth point cloud multi-feature fusion, the system comprising:

[0118] The acquisition module 1410 is used to acquire a 3D tooth point cloud of the identity to be identified;

[0119] A registration module 1420 is configured to input the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity in the identity database into the trained tooth multi-feature fusion identity recognition backbone network, and output a registration result between the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity;

[0120] The tooth multi-feature fusion identity recognition backbone network includes: a color feature extraction module, a structural feature extraction module, a feature fusion module, a local sampling feature extraction module, a feature aggregation module, a coarse correspondence prediction module, a point cloud decoding module, a fine correspondence prediction module, and a prediction registration module;

[0121] The identification module 1430 is configured to use the identity information of the 3D tooth point cloud of the registered identity that is optimally aligned with the 3D tooth point cloud of the identity to be identified as the identity identifier of the 3D tooth point cloud of the identity to be identified.

[0122] An embodiment of the present invention provides an identity recognition system based on 3D tooth point cloud multi-feature fusion, and its functional structure corresponds to an identity recognition method based on 3D tooth point cloud multi-feature fusion provided by an embodiment of the present invention, which will not be repeated here.

[0123] Figure 15It is a structural diagram of an electronic device 1500 provided in an embodiment of the present invention. The electronic device 1500 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 1501 and one or more memories 1502, wherein the memory 1502 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 1501 to implement the steps of the above-mentioned identity recognition method based on 3D tooth point cloud multi-feature fusion.

[0124] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory device containing instructions. The instructions are executable by a processor in a terminal to implement the above-described method for identifying an individual based on 3D tooth point cloud multi-feature fusion. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0125] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0126] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An identity recognition method based on 3D tooth point cloud multi-feature fusion, characterized in that: The method comprises: S1, collect the 3D tooth point cloud of the identity to be identified; S2. Input the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity in the identity database into the trained identity recognition backbone network with multi-feature fusion of teeth, and output the registration result between the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity; The tooth multi-feature fusion identity recognition backbone network includes: a color feature extraction module, a structural feature extraction module, a feature fusion module, a local sampling feature extraction module, a feature aggregation module, a coarse correspondence prediction module, a point cloud decoding module, a fine correspondence prediction module, and a prediction registration module; S3, using the identity information of the 3D tooth point cloud of the registered identity that is optimally aligned with the 3D tooth point cloud of the identity to be identified as the identity identifier of the 3D tooth point cloud of the identity to be identified; The color feature extraction module is used to extract the color feature information of the 3D tooth point cloud. The specific extraction process includes: Assume that the 3D tooth point cloud is P∈R N×6 , where N represents the number of points in P, and the number 6 indicates that each point in P has 3D spatial coordinate information and 3D color information; Define any point in P as the center point P center and with the preset R color As the radius of the sphere, define other points within this range as neighboring points P neighbor_1 ~P neighbor_n , the center point P center Perform difference calculation with the RGB color channel values ​​of each adjacent point: R diff =R center -R neighbor , G diff =G center -G neighbor , B diff =B center -B neighbor The calculated color difference is divided into certain intervals and the R diff , G diff and B diff These intervals represent different levels of color difference, and the sum of the statistical results of all intervals of the three channels constitutes the center point P center Initial color frequency histogram ICH(P center ); Then, taking each adjacent point as the center, repeat the above operation to obtain the color frequency histogram ICH(P neighbor_1 )~ICH(P neighbor_n ); According to the relationship between each neighboring point and the center point P center The distance d is used, and 1 / d is used as the weight. The color frequency histogram results of all neighboring points are added to the center point P. center On the color frequency histogram, the center point P is obtained center Final color frequency histogram The final color frequency histogram FCH(P center ) is converted to an M color dimensional vector, as the center point P center Color feature information; Take all points in P as the center point and perform the above operation to obtain the color feature information of all points in the point cloud, which is recorded as The 3 indicates that the first three columns of information of the feature are the three-dimensional spatial coordinates of the point.

2. The method according to claim 1, characterized in that The structural feature extraction module is used to extract structural feature information of the 3D tooth point cloud. The specific extraction process includes: Calculate the normal vector of each point in the point cloud based on the three-dimensional spatial coordinate information of each point in the point cloud and its surrounding points; Define any point in P as the center point P center and with the preset R struct As the radius of the sphere, define other points within this range as neighboring points P neighbor_1 ~P neighbor_n , describing the center point P center and the geometric features of each neighboring point, where the center point P is described as follows center and P neighbor_1 Geometric structure characteristics of two points: let vector d = P center -P neighbor_1 , represents the center point P center and P neighbor_1 The connecting vector between two points, the center point P center The normal vector is defined as n center , P neighbor_1 The normal vector is defined as n neighbor_1 , define α to represent n neighbor_1 The angle between and d is calculated as Measures the projection relationship of the normal vector of the neighborhood point in the direction of the target point; define φ to represent the center point P center Normal vector n center The angle between d and d is calculated as Reflects the neighborhood point P neighbor_1 Relative to the center point P center Normal vector n center The projection angle of n center and n neighbor_1 The angle between them is calculated as follows: θ = arccos(n center ·n neighbor_1 ), which measures the degree of alignment of the normal vectors of two points. These parameters together describe the center point P center and neighboring point P neighbor_1 The geometrical structural characteristics between Construct the center point P center The initial structure frequency histogram ISH(P center ), the initial structure frequency histogram ISH (P center )Total M struct intervals, used to count the center point P respectively center and P neighbor_1 The results of the three parameters α, φ and θ are used as feature information. The value range of α is [-1, 1]. When α = 1, it means that the normal vector of the neighboring point is completely parallel to the normal vector of the center point and has the same direction; when α = -1, it means that they are in opposite directions; when α = 0, it means that the two are perpendicular; the value range of φ is also [-1, 1]. When φ = 1, the normal vector of the center point is consistent with the direction of the line between the center point and the neighboring point; when φ = -1, the normal vector of the center point is opposite to the direction of the line between the center point and the neighboring point; when φ = 0, the normal vector of the center point is perpendicular to the line between the center point and the neighboring point; the value range of θ is [-π, π]. When θ = 0, it means there is no rotation; when θ = π or θ = -π, it means a rotation of 180 degrees, reflecting a larger angular offset; For the center point P center The other neighboring points of are then used to construct the structural frequency histogram, and then added to the center point P in a weighted manner. center The initial structure frequency histogram FSH(P center ) on the center point P center Final structural frequency histogram The center point P center The structural frequency histogram FSH(P center ) is converted to an M struct dimensional vector, as the center point P center Structural feature information; Take all points in P as the center point and perform the above operation to obtain the structural feature information of all points in the point cloud, which is recorded as The 3 indicates that the first three columns of information of the feature are the three-dimensional spatial coordinates of the point.

3. The method according to claim 2, characterized in that The feature fusion module extracts M from each point color dimensional color feature information and M struct The 3D structural feature information is spliced ​​with the 3D spatial coordinate information of each point, and the output is a point cloud containing 3+M color +M struct The fusion result of dimensional feature information The local sampling feature extraction module first performs the feature fusion module output Perform a downsampling operation to obtain a set of sparse points as neighborhood-level key points, reducing the number of points in the point cloud; A neighborhood is established for each key point with a neighborhood radius of K, which is used to search for original points close to these key points in the point cloud before downsampling. Then, according to the distance relationship between the searched original points and the corresponding key points, the features of these original points are aggregated to the corresponding key points as the local feature information F' of the corresponding key points, and the key points with local feature information (P', F') are output, where P' represents the three-dimensional spatial coordinates of the key points.

4. The method according to claim 3, characterized in that The feature aggregation module relies on the attention mechanism to enhance and aggregate local feature information in neighborhood-level key points, integrate the global and local context information of each point in the input key points, and enable interaction between the features of the neighborhood-level key points of the two to-be-registered point clouds, thereby improving the subsequent point cloud registration accuracy. It includes two self-attention modules and one cross-attention module. The processing process is as follows: First, the features of the input neighborhood-level key points are enhanced by the self-attention module. The self-attention mechanism contained in the self-attention module generates the query matrix Q, key matrix K and value matrix V by linearly transforming the input feature matrix F': Q=F'W Q ,K=F'W K ,V=F'W V Where W Q 、W K 、W V is a learnable linear transformation matrix; Next, the attention score matrix A is obtained by calculating the dot product of the query matrix and the key matrix and normalizing it: where d k is the size in the key matrix, which is used for normalization to prevent the value from being too large. The attention score matrix A is used to perform weighted summation on the value matrix V to obtain the aggregated feature matrix F": F”=A·V In this way, the features of each point not only contain its own information, but also aggregate the global context information related to it; Subsequently, the result is input into the cross-attention module to integrate the feature information between key points at different point cloud neighborhood levels. The feature matrices of the two point cloud key points processed by the self-attention module are F'X' and F'Y' respectively. The cross-attention module then performs feature aggregation through the following steps: Q X =F″ X W Q ,K Y =F″ Y W K ,V Y =F″ Y W V Calculate the attention score matrix A between F'X' and F'Y': Using the attention score matrix A X Y pair matrix V Y Perform weighted summation to obtain the enhanced feature matrix F″′ X : F″′ X =A X Y·V Y Calculate F″ X and F″ Y The cross attention is F″′ Y : F″′ Y =A Y X·V X Through the cross-attention module, the features of the two key points are mutually enhanced; After the cross attention module, the feature representation capability is further enhanced by a self-attention calculation. Through the feature aggregation module, the features of each point in the output neighborhood-level key points not only integrate the geometric information of its neighborhood, but also fuse the features from the global context, and also interact with the key points of the point cloud to be registered with it. Finally, the key points of the registered point cloud X after feature aggregation are output. And the key points of the identity point cloud Y to be identified after feature aggregation 5. The method according to claim 4, characterized in that The coarse correspondence prediction module obtains the key points of the registered point cloud X after feature aggregation by calculating the inner product of the point cloud feature vector And the key points of the identity point cloud Y to be identified after feature aggregation A similarity matrix S' between them, each element in the similarity matrix S' represents the similarity between two key points; After obtaining the similarity matrix S', a normalization algorithm is used to normalize the similarity matrix and solve it to obtain a matching probability matrix, which represents the matching confidence of each pair of points between the key points in the neighborhood of two point clouds; Then, by setting a confidence threshold, key point pairs with higher confidence are screened out from the matching probability matrix as rough correspondence predictions; Finally, a neighborhood-level matching point pair containing multiple coarse correspondences is output. This coarse correspondence set will be further processed in the subsequent fine correspondence prediction module to obtain more accurate point pair matching results.

6. The method according to claim 5, characterized in that The point cloud decoding module decodes and recovers the rough neighborhood key point correspondences to the single point level, and uses the KP convolution decoder to convert the key points of the registered point cloud X after feature aggregation into And the key points of the identity point cloud Y to be identified after feature aggregation Decode to the same resolution as the point cloud before downsampling to obtain the decoded point cloud and This process involves a series of convolution operations to gradually restore the spatial distribution and geometric features similar to those of the point cloud before downsampling. The KP convolution decoder will gradually propagate the features of the neighborhood key points to the restored points. These propagated features will be combined with the geometric information of the point cloud before downsampling to generate a more fine-grained feature representation. The decoding process also involves a point-to-neighborhood level key point grouping strategy. The point cloud grouping module is used to group each point in the decoded point cloud according to its neighborhood level key points, and the points in the decoded point cloud are assigned to the nearest neighborhood level key points, and the grouping results of the registered point cloud X are output. And the grouping result of the identity point cloud Y to be identified For subsequent refined registration, the subsequent fine correspondence prediction module uses the grouping results obtained from the point cloud decoding module and the key point matching results obtained from the coarse correspondence prediction module to further refine the matching point pairs and obtain single-point level matching point pairs.

7. The method according to claim 6, characterized in that The precise correspondence prediction module includes similarity matrix calculation, normalization processing and confidence screening steps. First, a set of matching point pairs is selected from the neighborhood level matching point pairs, and then the point cloud group constructed by these two points is selected from the grouping results: Grouping results of registered point cloud X And the grouping result of the identity point cloud Y to be identified Next, the similarity matrix is ​​calculated for these two point clouds. The similarity matrix By calculating the inner product of all point feature vectors between two point cloud groups, each element in the matrix represents the similarity between the two points; Afterwards, the results are normalized to obtain a matching probability matrix between the two point cloud groups, where the value of each element in the matching probability matrix represents the matching confidence of any two points between the two point cloud groups. Afterwards, the matching point pairs between the two point cloud groups are screened out by setting a confidence threshold. Finally, all neighborhood-level matching point pairs and their corresponding point cloud grouping results are used to perform the above processing, and all matching point pair screening results are merged together to obtain single-point-level matching point pairs as the output result of the precise correspondence prediction module.

8. The method according to any one of claims 1 to 7, characterized in that The tooth multi-feature fusion identity recognition backbone network, the total loss function L in the training process is the coarse corresponding prediction loss L c And the corresponding prediction loss L f The weighted sum of: L=L c +λL f Among them, λ is the weight parameter used to balance the two loss terms. Through this combination, the model can be optimized at both coarse and fine scales, ensuring that the model can learn the precise correspondence between point clouds from coarse to fine. At the coarse scale, the model is supervised to learn the reliable correspondence between the key points obtained by downsampling, while at the fine scale, the model further optimizes the matching accuracy of each point at the single point level. Through the combination of these two parts, the final model can achieve high-precision point cloud registration; The coarse corresponding prediction loss L c It is used to supervise the matching results of neighborhood-level key points. Its core idea is to use the local overlap ratio as a weighting scheme to supervise the coarse correspondence. Given a pair of neighborhood-level key points P' X (i') and P' Y (k'), their corresponding neighborhood point grouping results are and calculate Zhongyu The proportion of points with corresponding relationships is: in, Is a known quantity, which represents the real transformation matrix that can completely align the X point cloud and Y point cloud input into the network, τ P represents the distance threshold; calculate Zhongyu The proportion of points with corresponding relationships is: According to the above two formulas, the weighted matrix W' is obtained: W'(i',j')=min(r(i',j'),r(j',i')),i'≤n'∧j'≤m' Among them, n' and m' respectively represent the total number of neighborhood-level key points of the two point clouds, and each element in W'(i', j') represents the weighted value of the key point pair (i', j'). This weighted matrix is ​​used to reflect the geometric overlap between the key point pairs. The larger the weight, the more reliable the matching key point pair. c When , the cross entropy loss is used to measure the difference between the similarity matrix S' output by the model and the weight matrix W': Among them, S'(i',j') is the confidence score of the key point pair (i',j') recorded in the similarity matrix, which indicates the similarity between the corresponding key point pairs. This loss function supervises the model learning to predict more accurate coarse correspondences by minimizing the difference between the model's predicted point coarse matching results and the weighted matrix W'. The precision corresponds to the prediction loss L f Used to supervise fine matching at the single point level, it optimizes based on coarse correspondences: in, and They represent the two point groups formed by the lth neighborhood level key point matching point pair (i', j'), It is a binary matrix that represents the correct point pair matching relationship between the two point groups. For the matching point pairs, its value is 1, otherwise it is 0. The precision correspondence prediction loss L f Cross entropy loss is also used to measure the confidence matrix of the model output With binary matrix The differences between: in, Represents the similarity matrix between the points in the two point groups formed by the lth neighborhood-level key point matching point pair (i', j') predicted by the model. This loss function guides the model to learn more accurate single-point level point matching relationships by minimizing the error at the single-point level.

9. An identity recognition system based on 3D tooth point cloud multi-feature fusion, characterized by: The system comprises: The acquisition module is used to collect the 3D tooth point cloud of the identity to be identified; A registration module is configured to input the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity in the identity database into the trained tooth multi-feature fusion identity recognition backbone network, and output a registration result between the 3D tooth point cloud of the identity to be identified and the 3D tooth point cloud of each registered identity; The tooth multi-feature fusion identity recognition backbone network includes: a color feature extraction module, a structural feature extraction module, a feature fusion module, a local sampling feature extraction module, a feature aggregation module, a coarse correspondence prediction module, a point cloud decoding module, a fine correspondence prediction module, and a prediction registration module; an identification module, configured to use the identity information of the 3D tooth point cloud of the registered identity that is optimally aligned with the 3D tooth point cloud of the identity to be identified as the identity identifier of the 3D tooth point cloud of the identity to be identified; The color feature extraction module is used to extract the color feature information of the 3D tooth point cloud. The specific extraction process includes: Assume that the 3D tooth point cloud is P∈R N×6 , where N represents the number of points in P, and the number 6 indicates that each point in P has 3D spatial coordinate information and 3D color information; Define any point in P as the center point P center and with the preset R color As the radius of the sphere, define other points within this range as neighboring points P neighbor_1 ~P neighbor_n , the center point P center Perform difference calculation with the RGB color channel values ​​of each adjacent point: R diff =R center -R neighbor , G diff =G center -G neighbor , B diff =B center -B neighbor The calculated color difference is divided into certain intervals and the R diff , G diff and B diff These intervals represent different levels of color difference, and the sum of the statistical results of all intervals of the three channels constitutes the center point P center Initial color frequency histogram ICH(P center ); Then, taking each adjacent point as the center, repeat the above operation to obtain the color frequency histogram ICH(P neighbor_1 )~ICH(P neighbor_n ); According to the relationship between each neighboring point and the center point P center The distance d is used, and 1 / d is used as the weight. The color frequency histogram results of all neighboring points are added to the center point P. center On the color frequency histogram, the center point P is obtained center Final color frequency histogram The final color frequency histogram FCH(P center ) is converted to an M color dimensional vector, as the center point P center Color feature information; Take all points in P as the center point and perform the above operation to obtain the color feature information of all points in the point cloud, which is recorded as The 3 indicates that the first three columns of information of the feature are the three-dimensional spatial coordinates of the point.

Citation Information

Patent Citations

  • Underwater image enhancement method based on color correction and detail enhancement

    CN110689587A

  • CBCT and laser scanning point cloud data tooth registration method based on supervoxels

    CN112200843A