Multi-scale global local feature interaction-based palate wrinkle point cloud recognition system
The palatal wrinkle point cloud recognition system based on multi-scale global and local feature interaction solves the problem of information loss when converting three-dimensional palatal wrinkle data into two-dimensional image data, realizes efficient palatal wrinkle recognition, and improves recognition accuracy and feature extraction efficiency.
Patent Information
- Application Number
- CN202510939104.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-03
AI Technical Summary
When existing palatal wrinkle recognition technology converts three-dimensional data into two-dimensional image data, it cannot fully retain the three-dimensional structural information of the palatal wrinkle. In addition, traditional algorithms are greatly affected by the quality of palatal wrinkle images, making it difficult to extract representative and discriminative features. They have high computational complexity and are easily plagued by the curse of dimensionality problem.
A palatal wrinkle point cloud recognition system based on multi-scale global-local feature interaction is adopted. By constructing a downsampling module, a feature extraction network and a recognition head, the three-dimensional palatal wrinkle point cloud data is sampled using a point cloud data sampling module guided by a space-filling curve. Combined with the multi-scale feature extraction and feature enhancement modules, the interaction and enhancement of global and local features are achieved.
It significantly enhances the ability to distinguish palatal wrinkle morphology, improves recognition accuracy, solves the problems of information loss and high computational complexity in traditional methods, and improves the network's feature extraction efficiency and recognition accuracy of palatal wrinkle point clouds.
Smart Images

Figure CN120748010A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biometric recognition, and in particular relates to a palatal wrinkle point cloud recognition system based on multi-scale global-local feature interaction. Background Art
[0002] Palatal wrinkle is a biometric feature that can be used for identity authentication in special scenarios. Existing methods use traditional image processing technology to perform three-dimensional palatal wrinkle recognition. Three-dimensional palatal wrinkle recognition is achieved by identifying a sequence of two-dimensional palatal wrinkle slice images. Traditional methods such as cyclic spectrum analysis and fractional Fourier transform are mainly used to extract key features of palatal wrinkles. These schemes have the following main problems: on the one hand, the three-dimensional structural information of the palatal wrinkle cannot be fully retained during the conversion of three-dimensional data into two-dimensional image data; on the other hand, traditional algorithms are greatly affected by the quality of palatal wrinkle images, and it is difficult to extract representative and discriminative features. When processing high-dimensional feature spaces, the computational complexity increases accordingly, and it is easily plagued by the dimensionality curse problem.
[0003] The current status of research on palatal wrinkles both domestically and internationally. Research on palatal wrinkle recognition can be broadly categorized into three directions: palatal wrinkle recognition based on traditional morphological analysis, palatal wrinkle recognition based on digital image processing, and palatal wrinkle recognition based on deep learning. Palatal wrinkle recognition based on traditional morphological analysis extracts morphological features, such as shape, size, and curvature, through manual observation and analysis of palatal wrinkle images, and then performs palatal wrinkle matching. Palatal wrinkle recognition based on digital image processing extracts palatal wrinkle image features using various operators or frequency domain analysis and transform domain feature extraction techniques, and then uses support vector machines (SVMs) or K-nearest neighbor (KNN) classifiers for matching. Palatal wrinkle recognition based on deep learning uses convolutional neural networks to automatically extract multi-level abstract features from the original image, including low-level edges, mid-level textures, and high-level semantic information. Fully connected layers (linear layers) are then used to further process the extracted multi-level features, mapping the high-dimensional features into a low-dimensional space for classification or matching.
[0004] Palatal wrinkle identification based on traditional morphological analysis: In the early stages of palatal wrinkle identification technology, research primarily relied on manual analysis. Scholars have systematically analyzed the morphological structure, measurement, and directional distribution of palatal wrinkles, establishing a variety of morphologically based palatal wrinkle classification methods. As early as 1889, Allen et al. first defined palatal wrinkles as a series of transverse or curved wrinkles on the human palate (hard palate). The shape, size, number, position, and arrangement of these wrinkles vary from individual to individual. In 1955, Lysell proposed a relatively systematic method for palatal wrinkle morphology classification. Since then, many scholars have proposed different palatal wrinkle morphology classification methods. Among them, the classification scheme proposed by Thomas, Kapali et al. has gained widespread recognition and application in palatal wrinkle research and practical applications due to its scientific and practical application. However, these methods have many limitations. For example: 1. Classification results rely heavily on the subjective judgment of the annotator, which can easily lead to significant discrepancies between different annotators, thus affecting the consistency and reliability of classification. 2. Manual annotation is time-consuming and labor-intensive, making it difficult to ensure sufficient accuracy and real-time performance in large-scale sample processing. 3. During palatal wrinkle authentication, newly collected samples often require manual annotation by experts, which not only increases the workload but also reduces application efficiency. In summary, morphological palatal wrinkle recognition methods have significant shortcomings in terms of accuracy, annotation consistency, and large-scale application.
[0005] Palatal wrinkle recognition based on digital image processing: With the continuous advancement of digital image processing technology, more and more scholars have begun applying this technology to palatal wrinkle recognition research. In 2010, Bernitz et al. conducted an in-depth study of palatal wrinkle samples based on the principle of affine transformation. In the absence of biomarkers such as DNA, palatal wrinkle features demonstrate unique advantages in identity recognition. In 2015, Pan Fei et al. further demonstrated through image analysis the stability, variability, and universality of palatal wrinkles, providing a new perspective for research in this field. That same year, Taneva et al. proposed a new method: they first cast the palatal wrinkle into a plaster model, then used a scanner to convert it into a three-dimensional digital image, and finally extracted image features for matching. Although this method provided new research insights, the matching results fell short of expectations and were generally lackluster. In 2016, Wu Xiuping et al. developed a recognition system based on two-dimensional palatal wrinkle images for forensic identity verification. That same year, Assiri designed and developed a three-dimensional palatal wrinkle biometric recognition system. In 2017, Wu Xiuping et al. used the palatal wrinkle model analysis method to evaluate the individual uniqueness of palatal wrinkle morphology and the stability of palatal wrinkle morphology before and after orthodontic treatment. Their research explored the application of palatal wrinkle morphology in individual identification in forensic science, especially how palatal wrinkle morphology plays a role in forensic identity identification. Studies have shown that although orthodontic treatment can cause certain changes in palatal wrinkle morphology, the extent of such changes is limited, and the palatal wrinkle still maintains a certain uniqueness. In 2021, Zhang et al. used the equidistant slicing method to reduce the dimensionality of three-dimensional palatal wrinkle data, converting complex three-dimensional palatal wrinkle data into two-dimensional image data, thereby reducing the computational complexity, and introduced a cyclic spectrum feature blocking scheme to optimize the feature dimensionality reduction process and achieved significant recognition results. In 2023, Shang Guan et al. proposed a palatal wrinkle recognition method based on the Fractional Fourier Transform (FRFT). This method fully utilizes the advantages of FRFT in extracting time-domain and frequency-domain features of signals. By performing multi-order fractional Fourier transforms on palatal wrinkle images, a comprehensive description of their texture and morphological details is achieved, thereby more accurately reflecting the internal structure of the palatal wrinkle area. This study provides a new technical path for palatal wrinkle recognition technology, lays a solid foundation for practical applications in the field of forensic identity authentication, and promotes the further development of palatal wrinkle recognition technology. However, such methods have poor robustness in the face of image noise, and recognition accuracy is difficult to guarantee. Therefore, palatal wrinkle recognition methods that rely on digital image processing technology still need to be improved in terms of stability and generalization ability.
[0006] Palatal wrinkle recognition based on deep learning technology: Due to the success of deep learning in the field of image recognition, more and more researchers have applied it to palatal wrinkle recognition and proposed a variety of methods. In 2021, Luo Qiang extracted the one-dimensional signal of the maxillary edge curve of the palatal wrinkle, and for the first time used a pre-trained convolutional neural network to construct a three-dimensional palatal wrinkle digital recognition system. He also conducted experiments to analyze in detail the performance differences of the models in terms of recognition accuracy, robustness and feature extraction. Yang Tingyu proposed a palatal wrinkle recognition network with multi-channel feature fusion, which uses three independent channels to extract the texture and spatial features of the palatal wrinkle respectively, and integrates these features through a fusion strategy to form a more comprehensive feature representation. In the same year, Yang Tingyu designed a new palatal wrinkle cross-domain recognition network, whose core architecture is a multi-scale feature pyramid based on context perception. This method integrates a learnable adaptively optimized Gabor filter, which enables the network to efficiently extract the texture and structural information of the palatal wrinkle at different scales, providing a new direction for research in related fields. In 2024, Wang Shuai proposed a three-dimensional palatal wrinkle recognition network based on multi-scale feature fusion. By drawing on the research ideas of gait recognition, he designed a palatal wrinkle recognition network.
[0007] However, this type of method mainly relies on two-dimensional image data and cannot fully utilize the three-dimensional geometric information of palatal wrinkles, resulting in information loss during the recognition process. Summary of the Invention
[0008] In order to solve the technical problem in the existing technology that some tiny palatal wrinkle details may be lost due to factors such as viewing angle selection and image resolution when three-dimensional palatal wrinkle data is converted into two-dimensional image data, the present invention provides a palatal wrinkle point cloud recognition system based on multi-scale global-local feature interaction. Three-dimensional palatal wrinkle point cloud data is used for palatal wrinkle recognition. The multi-scale strategy can model the point cloud data from multiple levels, mine richer feature information, and significantly enhance the network's ability to distinguish palatal wrinkle morphology.
[0009] To achieve the above objectives, the technical solution adopted by the present invention is: a palatal wrinkle point cloud recognition system based on multi-scale global-local feature interaction, with the following specific steps:
[0010] Step 1: Construct a palatal wrinkle point cloud recognition network consisting of a downsampling module, a feature extraction network and a recognition head; the downsampling module is used to downsample the original point cloud data in order to effectively reduce the amount of data while ensuring the retention of important information; the feature extraction network is used to efficiently represent the downsampled point cloud data in order to learn the rich and multi-dimensional geometric information of the palatal wrinkle point cloud; the recognition head is mainly responsible for processing and analyzing the extracted features to output the final recognition results.
[0011] Step 2: Sample the three-dimensional palatal wrinkle point cloud data through the point cloud data sampling module guided by the space-filling curve, map the points in the three-dimensional space to the one-dimensional space through Z-order curve encoding, traverse the converted data in order, and then uniformly sample it to reduce the data volume, screen out sampling points that are more sensitive to the palatal wrinkle geometry and local correlation, and input the downsampled palatal wrinkle point cloud data into the multi-scale feature extraction backbone network.
[0012] Step 3: A hierarchical multi-scale feature extraction scheme is used to extract features from point cloud data. Multiple groups of feature aggregation modules are designed in the feature extraction network. Considering that the point cloud data is gradually downsampled in the process of continuous feature aggregation, a multi-scale feature extraction module and a feature enhancement module are designed in each feature aggregation network. In the multi-scale feature extraction module, three different scales are used to extract neighborhood local features of different receptive fields. Global feature information can be obtained by aggregating local features. Information interaction of global and local features can be achieved by weighted feature learning of global features and local features of different scales. In the feature enhancement module, the correlation between different feature channels is learned to improve the network's ability to learn key features of salient points and enhance the extracted features. This design can ensure the network's feature extraction efficiency for sparse point clouds, thereby solving the problem of insufficient expression of local geometric features.
[0013] The multi-scale feature extraction module of each feature aggregation network includes a multi-scale local feature extraction stage and a global-local feature interaction stage. In the multi-scale local feature extraction stage, three KNN neighborhoods of different scales are constructed for each sampling point. The scale range is determined by controlling the K value to extract neighborhood local features of different receptive fields, namely the semantic feature information and geometric position feature information of each sampling point and its neighborhood points; in the global-local feature interaction stage, the combined features of the three different scale neighborhood features are regarded as global features, and weighted learning is performed on them and the neighborhood features at each scale to achieve the purpose of establishing the connection between global feature information and local feature information and enhancing the dependency and expression ability between features.
[0014] The specific operation is as follows: In the multi-scale local feature extraction stage, a multi-scale KNN neighborhood (three scales) is constructed for each sampling point. , , ), to ensure the diversity of the receptive field. At the sampling point Each neighborhood of First, by combining its semantic feature information and geometric position feature information Neighborhood fusion features ,
[0015] (1)
[0016] in, express The semantic feature information of express The geometric position feature information of , the graph convolution in the neighborhood is formulated as:
[0017] (2)
[0018] in, yes The local fusion features of is a 1×1 convolution operation, Indicates the Neighborhood points and The relationship between them is defined as:
[0019] (3)
[0020] in, Represents the neighborhood range The semantic feature information of Represents the neighborhood range Specifically, since the dimensions of the input features and corresponding point sets are (N, C) and (N, 3), The dimension is , and defines different output dimensions of graph convolution for different scale neighborhoods , Used to construct sampling points The semantic and geometric relationship between the points and the neighborhood. Use the maximum pooling operation to aggregate local information to the sampling point , given sampling points Aggregate local features at each scale , connecting them through a cascade operation to generate Multi-scale local features , which is expressed as:
[0021] (4)
[0022] in, Indicates sampling point Three local features at three different scales.
[0023] After the three-dimensional palatal wrinkle point cloud data undergoes multi-scale local feature extraction, the second stage of feature extraction, that is, the global-local feature interaction stage, is carried out. As a new feature learning mechanism, global-local feature interaction establishes a connection between global feature information and local feature information, enhances the dependency and expression ability between features, and thus significantly improves the network's ability to learn and recognize complex geometric features in the palatal wrinkle point cloud. Global-local feature interaction combines the features output from the multi-scale local feature extraction stage with the features output from the multi-scale local feature extraction stage. As input, Indicates the number of sampling points, express The characteristic dimension of .
[0024] First, Projected into two different feature spaces to generate Query (hereinafter referred to as Q matrix) and Key matrix (hereinafter referred to as K matrix):
[0025] (5)
[0026] in, It is a learnable weight matrix, and then calculates the attention map of the global and local feature interaction stage Formulated as:
[0027] (6)
[0028] in, represents the query matrix and key matrix, is a learnable position encoding matrix;
[0029] Next, the neighborhood feature map at each scale is regarded as a branch, and the output neighborhood feature set is obtained by calculating the weighted sum of all input sets. This can effectively utilize the global information of all points including sampling points and neighborhood points, rather than just extracting the location information and semantic information of the sampling points themselves. This design can improve the learning ability of features and reduce the local information loss caused by the pooling operation. The output feature can be expressed as:
[0030] (7)
[0031] in, Represents the value matrix, which is the neighborhood feature map at each scale, represents the maximum pooling operation;
[0032] Connect at every scale , to obtain global-local interaction features :
[0033] (8)
[0034] in, Generated by neighborhood feature maps of different scales.
[0035] Global-local interaction features It is sent to the LBR module consisting of a linear layer, a batch normalization layer, and a RELU activation function layer, and the obtained features are added to the multi-scale features to obtain the final output features. :
[0036] (9).
[0037] In the feature enhancement module of each feature aggregation network, the output features from the multi-scale feature extraction module are run in parallel under the channel feature enhancement mechanism and the spatial feature enhancement mechanism. Finally, the channel feature enhancement mechanism and the spatial feature enhancement mechanism are jointly applied to the input features to obtain the final output features. .
[0038] In the channel feature enhancement mechanism, point cloud features are aggregated through maximum pooling and average pooling operations, and then forwarded to and , the specific formula is as follows:
[0039] (10)
[0040] (11)
[0041] in, represents the maximum pooling function, represents the average pooling function.
[0042] The output of the two convolutional layers The sum is then passed through the sigmoid function to learn the weight of each feature channel. The channel feature enhancement mechanism is defined as:
[0043] (12)
[0044] In the spatial feature enhancement mechanism, features are input into the MLP to change the channel dimension, and information about each spatial position is aggregated through the batch normalization layer and the pooling layer. The spatial position attention weight is obtained through the sigmoid function. The spatial feature enhancement mechanism is defined as:
[0045] (13)
[0046] Apply the channel feature enhancement mechanism and the spatial feature enhancement mechanism together to the input feature to obtain the output feature for:
[0047] (14).
[0048] Step 4: The features after feature aggregation processing are sent to the recognition head, and after processing, the point cloud recognition results can be output.
[0049] The present invention proposes a palatal wrinkle point cloud recognition system based on the interaction of multi-scale global and local features. To extract representative structural information from the original point cloud, the present invention designs a point cloud data sampling module guided by a space-filling curve. The space-filling curve is in the form of a Z-shape to achieve effective downsampling while retaining key geometric features. Secondly, to enable the network to learn richer geometric features, the present invention proposes a multi-scale feature extraction module for capturing multi-scale local information and achieving efficient fusion between local and global features. The present invention guides the network to learn the importance of salient points in the channel and spatial dimensions by constructing a feature enhancement module, thereby enhancing the expression effect of key features. The proposed method significantly outperforms the current state-of-the-art network models in palatal wrinkle recognition tasks, especially in terms of recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a structural block diagram of the present invention.
[0051] Figure 2 This is the Z-order sampling flow chart.
[0052] Figure 3 This is the structural diagram of the multi-scale feature extraction module.
[0053] Figure 4 A structural diagram of the interaction between global and local features.
[0054] Figure 5 This is the structural diagram of the feature enhancement module.
[0055] Figure 6 The histogram of the network's recognition accuracy under different numbers of sampling points.
[0056] Figure 7 This is a comparison chart of the error rates of experimental networks using different samples.
[0057] Figure 8 A comparison chart of the recognition accuracy of different networks based on a three-dimensional palatal wrinkle point cloud.
[0058] Figure 9 The three-dimensional palatal wrinkle data collected under four conditions are shown in Figure 2. Figure 9 (a) is the sample collection diagram T when the saliva is not cleaned. Figure 9 (b) is the sample S collected after rinsing the mouth. Figure 9 (c) is sample B collected under standard conditions, Figure 9 (d) is the sample picture Q without teeth collected during the collection.
[0059] Figure 10 There are two forms of palatal wrinkle data. Figure 10 (a) is a two-dimensional palatal wrinkle image data diagram, Figure 10 (b) is the three-dimensional palatal wrinkle point cloud data map.
[0060] Figure 11 The three-dimensional palatal wrinkle data diagrams after cropping in four cases are shown below. Figure 11 (a) is the point cloud data of T after cropping (sample collected without cleaning saliva), Figure 11 (b) is the point cloud data of S after cropping (sample collected after rinsing the mouth), Figure 11 (c) is the point cloud data of B after cropping (sample collected under standard conditions), Figure 11 (d) is the point cloud data of Q after cropping (teeth were not collected during acquisition).
[0061] Figure 12 The three-dimensional palatal wrinkle point cloud data diagrams in four cases are shown. Figure 12 (a) is the point cloud data of T, Figure 12 (b) is the point cloud data of S, Figure 12 (c) is the point cloud data of B, Figure 12 (d) is the point cloud data diagram of Q.
[0062] Figure 13 The three-dimensional palatal wrinkle point cloud data graphs under four different sample point numbers are shown. Figure 13 (a) is the point cloud data with 10,000 sample points. Figure 13 (b) is the point cloud data diagram with 8192 sample points. Figure 13 (c) is the point cloud data diagram with 4096 sample points. Figure 13 (d) is the point cloud data diagram with 2048 sample points. DETAILED DESCRIPTION
[0063] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0064] The accuracy of palatal wrinkle recognition is defined as the ratio of the number of correctly classified positive samples to the total number of samples, as shown in (15):
[0065] (15)
[0066] in, represents the total number of samples, (True Positive) means that the model correctly predicts the positive class samples as positive.
[0067] The equal error rate (EER) is a commonly used metric for evaluating the performance of a classifier or recognition system. The lower the EER, the better the performance of the classifier or recognition system. The calculation of the EER requires calculating the false acceptance rate (FAR) and the false rejection rate (FRR) at different classification thresholds. By adjusting these thresholds, a point can be found where the FAR and FRR are equal. The corresponding error rate at this point is the EER. The specific calculation formulas for FAR and FRR are shown in (16) and (17):
[0068] (16)
[0069] (17)
[0070] Among them, FP (False Positive) refers to the incorrect prediction of a sample that is actually a negative class as a positive class, TN (True Negative) refers to the model correctly identifying a negative class sample as a negative class; and FN (False Negative) refers to the incorrect prediction of a sample that is actually a positive class as a negative class.
[0071] The 3Shape TRIOS 5 oral scanner is used as the core acquisition device. After completing the palatal wrinkle area scan, the scan data is saved or exported to formats such as STL files.
[0072] To systematically evaluate and reduce the impact of these interfering factors, four different sampling schemes were designed to optimize data collection quality and improve the reliability of palatal rugae identification. These four scenarios included samples collected without saliva cleaning (T), samples collected after rinsing (S), samples collected under standard conditions (B), and samples collected without teeth but only including the incisor papillae (Q).
[0073] Figure 9 The three-dimensional palatal wrinkle data collected under four different conditions were observed. During the collection and processing of the three-dimensional palatal wrinkle data, it was found that the original data not only contained the morphological information of the palatal wrinkle, but also included a large amount of tooth data. However, too much tooth information may lead to additional non-palatal wrinkle characteristic differences between different samples, thereby affecting the accuracy of subsequent data analysis and model construction. Therefore, in order to ensure the purity of the palatal wrinkle data and improve the reliability of the data, the present invention introduces a data clipping method to eliminate redundant tooth parts and retain only the palatal wrinkle data in the central area.
[0074] like Figure 10As shown in the figure, compared with the two-dimensional palatal wrinkle image data, the three-dimensional palatal wrinkle point cloud data has obvious advantages in accuracy, morphological description, spatial information integrity and subsequent analysis.
[0075] CloudCompare software was used to process the three-dimensional palatal wrinkle data obtained under four different acquisition conditions. For each read three-dimensional data, polygonal region cropping was used to accurately remove non-palatal wrinkle areas and obtain palatal wrinkle data with regular morphology and high integrity. This method can not only effectively remove redundant information, but also ensure the integrity of the palatal wrinkle structure and the consistency of the data.
[0076] During the cropping process, an appropriate cropping region was selected based on the area of the palatal wrinkle. The vertex positions of the polygonal cropping region were set as follows: the midpoints of the two front teeth, the midpoints of the second and third teeth on the left, the midpoint of the sixth tooth on the left, the midpoints of the second and third teeth on the right, and the midpoint of the sixth tooth on the right. These key points ensured that the cropping region fully encompassed the palatal wrinkle information while effectively eliminating unnecessary tooth data.
[0077] The specific process for cropping is as follows: First, the collected 3D palatal wrinkle data is read using CloudCompare software. Then, the polygon tool is used to manually select the key points mentioned above and draw the polygonal boundary of the cropping area. Subsequently, the selected area is confirmed and cropped, retaining only the palatal wrinkle data within the polygon and removing any excess tooth information. Finally, the cropped data is saved in .obj format and named according to a naming convention that combines a label with the corresponding acquisition abbreviation to ensure standardized data management.
[0078] Figure 11 The three-dimensional palatal wrinkle data after cropping under four different acquisition conditions are shown. The palatal wrinkle data obtained by this method is more refined and removes redundant information that may affect the analysis results.
[0079] Next, the four cropped obj files corresponding to each volunteer were loaded into the MeshLab environment. To ensure the correct orientation of the data, "double-sided display" mode was selected to check the data for any undeleted redundant information, thereby ensuring data integrity and accuracy. If any undeleted redundant information was found, it was deleted to ensure the accuracy of subsequent analysis.
[0080] During data preprocessing, the Poisson-Disk Sampling method was used to process the loaded 3D palatal wrinkle data. 20,000 sampling points were set for each palatal wrinkle point cloud. After sampling, the four resulting point cloud data sets were named to facilitate subsequent data management and analysis. The naming method consisted of the volunteer number plus the suffixes T, S, B, and Q, corresponding to data collected under different conditions. All data were saved in the ply format, which preserves not only the 3D coordinate information of the point cloud but also the normal vector information, ensuring data integrity and richness.
[0081] In order to improve the versatility and operability of the data, the data in ply format is converted into txt format by using the data conversion method. The txt format can store the coordinate information of the point cloud in text form, making it convenient for subsequent algorithm processing and data analysis. Figure 12 The following diagrams illustrate the 3D palatal wrinkle point cloud data for four different sampling conditions, visually demonstrating the distribution characteristics of the palatal wrinkle point cloud after sampling. The above processing was performed on palatal wrinkle data from 180 individuals, including 70 males and 110 females, with an average age of 23 to 35 years. Thirty individuals had a history of orthodontic treatment. Four sets of data were collected for each individual, resulting in a total of 720 3D palatal wrinkle point cloud data samples.
[0082] Through visual analysis of different sampling densities, the changes in the morphological characteristics of the three-dimensional palatal wrinkle point cloud data at 10,000, 8,192, 4,096, and 2,048 data points are demonstrated. Figure 13 The distribution of palatal wrinkle point cloud data under different sampling densities is intuitively displayed. High-density point clouds can more accurately describe the palatal wrinkle structure, while low-density point clouds simplify the detailed information while retaining the overall shape.
[0083] In order to construct a standardized 3D palatal wrinkle point cloud dataset, two key processing steps were performed on the collected raw data:
[0084] (1) Data cropping. The original collected data not only contains the palatal wrinkle structure, but may also contain too much tooth information. The interference of tooth data may lead to non-palatal wrinkle characteristic differences between different samples. Therefore, CloudCompare software was used to remove redundant information from the three-dimensional palatal wrinkle data under four conditions. The polygonal region cropping method was adopted in the software, and key cropping points (such as the midpoint of the incisor, the midpoint between the second and third teeth, the midpoint of the sixth tooth, etc.) were selected according to the characteristics of the palatal wrinkle data to ensure that the cropping area can fully retain the palatal wrinkle information while effectively removing redundant tooth parts. Finally, the cropped data was stored in obj format to form standardized palatal wrinkle point cloud data.
[0085] (2) Data sampling. In order to further optimize data quality and improve data consistency and comparability, the cropped palatal wrinkle data were uniformly sampled using MeshLab software using Poisson distance. This method can evenly distribute point cloud data while retaining the geometric details of the palatal wrinkle, making the sampled data more stable in morphological analysis. The number of sampling points for each palatal wrinkle data was set to 20,000 to ensure data integrity and experimental consistency. The sampled data were named with the volunteer number plus the suffixes T, S, B, and Q, and saved in ply format, and then converted to txt format to meet different data analysis requirements.
[0086] like Figure 1 As shown, the present invention proposes a palatal wrinkle point cloud recognition network with multi-scale global-local feature interaction. This method uses three-dimensional palatal wrinkle point cloud data containing coordinate position information and normal vector information as input. Before entering the network, the palatal wrinkle point cloud data is first sampled using a space-filling curve-guided point cloud data sampling module (referred to as Z-order sampling) to extract key feature points, reducing the data volume while ensuring the preservation of important information. Subsequently, the downsampled palatal wrinkle point cloud data is input into a multi-scale feature extraction backbone network, which consists of three feature aggregation networks. Each feature aggregation network contains two modules: a multi-scale feature extraction module and a feature enhancement module. In the multi-scale local feature extraction stage, different neighborhood K values are constructed for each point to extract feature information from different local ranges in the palatal wrinkle point cloud. Finally, the feature information from different local ranges is aggregated into a global feature. In the global-local feature interaction stage, the multi-scale local features extracted by the multi-scale feature extraction module are linked to the global features, enabling interactivity between global features and local neighborhood information, thereby improving feature expression and the discriminative ability of the model. On this basis, the feature enhancement module automatically learns attention weights for each channel and space, allowing the network to focus on significant feature areas in the palatal wrinkle point cloud when extracting features. This further optimizes feature extraction, ensures the network's efficiency in extracting features from sparse point clouds, and addresses the issue of insufficient representation of local geometric features. After the input data passes through three feature aggregation networks, the output features of each feature aggregation network are subjected to a maximum pooling operation, which are then concatenated to generate multi-level feature information. Finally, a multi-layer perceptron (MLP) is used for palatal wrinkle point cloud recognition. The MLP consists of two fully connected layers with batch normalization and RELU activation functions.
[0087] This paper introduces a space-filling curve-guided point cloud data sampling module (Z-order sampling module for short) to sample 3D palatal wrinkle point cloud data. As a space-filling curve, the Z-order curve has a strong ability to preserve spatial locality, mapping points in 3D space to 1D space while preserving the spatial adjacency between the original points as much as possible. Its introduction is based on the following two considerations:
[0088] (1) Taking both global and local considerations into account: Using equally spaced sampling in the order after Z-order encoding can not only cover the entire point cloud space (reflecting the global structure), but also ensure that there is a certain semantic correlation between the selected points (reflecting local consistency), thereby achieving dual modeling of the global structure and local features of the point cloud.
[0089] (2) Improving semantic focusing ability: Z-order sampling guides the sampling points to be distributed in semantically significant areas, so that the model can pay more attention to key areas in the subsequent feature extraction process, thereby improving the accuracy of classification or recognition tasks.
[0090] like Figure 2 As shown in Figure 1, the principle of Z-order sampling is to use a continuous curve to pass through all points in space, with each point corresponding to a position code. After sequential encoding, the three-dimensional position coordinates of each point in the sub-point cloud are mapped to a one-dimensional feature space. The position of the original point is well preserved due to the properties of the Z-order curve. After encoding and sorting the points, equally spaced sampling is performed. The final sampled point set can effectively represent the global structural characteristics and local semantic relevance of the original point set.
[0091] After the Z-order sampling module, the data volume of the 3D palatal wrinkle point cloud is reduced to N / 4 points, such as Figure 3 As shown, the present invention introduces a multi-scale feature extraction module during the feature extraction phase. By extracting features at different scales and interacting between local and global features, the module fully recovers and enhances the semantic information of 3D palatal wrinkle point cloud data. The multi-scale feature extraction module is divided into two stages: multi-scale local feature extraction and global-local feature interaction.
[0092] The specific operation of multi-scale local feature extraction is as follows: In the multi-scale local feature extraction stage, a multi-scale KNN neighborhood is constructed for each sampling point (three scales in the experiment). , , ), to ensure the diversity of the receptive field. At the sampling point Each neighborhood of First, the coordinate position information and semantic feature information of the point Combine them to generate fusion features :
[0093] (18)
[0094] in, and Respectively The semantic feature information and geometric position feature information of the given fusion neighborhood feature , the graph convolution in the neighborhood is formulated as:
[0095] (19)
[0096] in, yes Aggregate local features, is a 1×1 convolution operation, Indicates the Neighborhood points and The relationship between them is defined as:
[0097] (20)
[0098] in, and Respectively represent the neighborhood range semantic feature information and geometric position feature information.
[0099] Specifically, in Figure 3 In , since the dimensions of the input features and corresponding point sets are (N, C) and (N, 3) respectively, The dimension is , defines different output dimensions of graph convolution for different scale neighborhoods , Used to construct sampling points The semantic and geometric relationship between the points and the neighborhood. Then, the maximum pooling operation is used to aggregate the local information to the sampling point. , given sampling points Aggregate local features at each scale , connecting them through a cascade operation to generate Multi-scale local features , which is expressed as:
[0100] (twenty one)
[0101] in, Indicates sampling point Three local features at three different scales.
[0102] After the three-dimensional palatal wrinkle point cloud data undergoes multi-scale local feature extraction, the second stage of feature extraction is carried out, that is, global-local feature interaction. As a new feature learning mechanism, global-local feature interaction establishes a connection between global feature information and local feature information, enhances the dependency and expression ability between features, and thus significantly improves the network's ability to learn and recognize complex geometric features in palatal wrinkle point clouds. Global-local feature interaction is as follows Figure 4 As shown, the aggregated features extracted from multi-scale local features are As input, is the number of sampling points, express The characteristic dimension of Project into two different feature spaces to generate Q, K matrices:
[0103] (twenty two)
[0104] in, is a learnable weight matrix.
[0105] Then, the attention map of the global and local feature interaction stage is calculated Formulated as:
[0106] (twenty three)
[0107] in, denote the query matrix and the key matrix, and is a learnable positional encoding matrix.
[0108] The neighborhood feature map at each scale is regarded as a branch, and the output neighborhood feature set is obtained by calculating the weighted sum of all input sets. The output feature is expressed as:
[0109] (twenty four)
[0110] in, Represents the value matrix, which is the neighborhood feature map at each scale, Represents the max pooling operation.
[0111] Finally, connect at each scale , to obtain global-local interaction features :
[0112] (25)
[0113] in, It is generated from neighborhood feature maps of different scales.
[0114] Then the global-local interaction feature The final output feature is obtained by adding the multi-scale features through the LBR module (including linear layer, batch normalization layer, RELU activation function) :
[0115] (26)
[0116] like Figure 5 As shown in Figure 2, the feature enhancement module runs the output features from the multi-scale feature extraction module in parallel under the channel feature enhancement mechanism and the spatial feature enhancement mechanism, and enhances the network's ability to extract features by capturing the channel spatial features of the most important points. In the channel feature enhancement mechanism, the point cloud features are aggregated through maximum pooling and average pooling operations and forwarded to and The design of the two convolutional layers is to first reduce the feature dimension and then increase the feature dimension to better extract features. The formula is as follows:
[0117] (27)
[0118] (28)
[0119] in, and Represent the maximum pooling and average pooling functions respectively.
[0120] Then the output of the two convolutional layers is The sum is then passed through the sigmoid function to learn the weight of each feature channel. The channel feature enhancement mechanism is defined as:
[0121] (29)
[0122] In the spatial feature enhancement mechanism, features are input into the MLP to change the channel dimension, and information about each spatial position is aggregated through the batch normalization layer and the pooling layer. The spatial position attention weight is obtained through the sigmoid function. The spatial feature enhancement mechanism is defined as:
[0123] (30)
[0124] Applying the channel feature enhancement mechanism and the spatial feature enhancement mechanism together to the input feature can obtain the final output feature for:
[0125] (31)
[0126] Experimental analysis: During training, the input of the network of the present invention is palatal wrinkle point cloud data of N=8192 points, each point has C=6-dimensional features, including (Euclidean coordinate information (x, y, z), normal vector information (nx, ny, nz)). The proposed network is trained using PyTorch. The optimizer in the network uses the Adam optimizer, and batch normalization is performed on all layers except the last classification layer. The initial learning setting is 0.01. The network is trained for a total of 600 epochs on a single Nvidia GeForce RTX3090 GPU with a batch size of 4.
[0127] Dataset: The dataset constructed in the present invention contains palatal wrinkle information of 180 individuals. Four sets of data are collected for each individual, namely, three-dimensional palatal wrinkle point cloud data with saliva (T), three-dimensional palatal wrinkle point cloud data after rinsing the mouth (S), three-dimensional palatal wrinkle point cloud data after rinsing and drying the mouth (B), and three-dimensional palatal wrinkle point cloud data with the incisor mastoid as the reference point (Q). A total of 720 three-dimensional palatal wrinkle point cloud data samples are obtained.
[0128] To verify the role of the multi-scale feature extraction module in improving network performance, a series of ablation experiments were conducted using three-dimensional palatal wrinkle point cloud data (B) after rinsing and drying the mouth. The experiments focused on analyzing the specific impact of different scale numbers on network recognition performance. The scale number refers to the number of branches in the multi-scale feature extraction module. In the experiments, the present invention adjusted the number of scales in the feature extractor and observed changes in the network's recognition accuracy and equal error rate. The experimental results, shown in Table 1, show that when the multi-scale feature extraction module extracts features at three scales, the network's recognition accuracy reaches the highest level, and the equal error rate decreases significantly, demonstrating superior performance compared to single-scale feature extraction.
[0129]
[0130] When the multi-scale feature extraction module extracts features at three scales, the network is able to simultaneously learn multi-level information from different scales. This is particularly true when processing palatal wrinkle point cloud data with complex local geometry. A single scale often fails to capture all key features, potentially leading to incomplete feature extraction or missing information. The multi-scale strategy, however, models point cloud data at multiple levels, mining richer feature information and significantly enhancing the network's ability to discriminate palatal wrinkle morphology. This demonstrates that, in practical applications, the use of a multi-scale strategy can help improve generalization capabilities in point cloud data recognition tasks.
[0131] To evaluate the performance of the proposed method under varying point cloud densities, we conducted comparative experiments with two mainstream methods at varying sampling point numbers: EB-LG, a typical local-global feature learning model using error feature back-projection, and DGCNN, a dynamic graph convolutional network, a classic deep learning model for processing point cloud and graph-structured data.
[0132] From Table 2 and Figure 6 It can be seen that in the five cases of sampling points of 2048, 4096, 8192, 9000, and 10000, the method proposed in this invention outperforms the other two comparison methods in terms of accuracy, showing a significant performance advantage and achieving the best experimental results. This experimental result shows that the method of this invention has strong feature extraction capabilities and good generalization performance, can adapt to point cloud data input at different sampling densities, and avoid performance fluctuations caused by changes in the number of points. It can be further observed from Table 2 that all three networks perform best when the number of points is 8192, so the number of points is set to 8192 in the subsequent experimental stage.
[0133]
[0134] In order to verify the effectiveness of the proposed module, the three-dimensional palatal wrinkle point cloud data (B) after rinsing and drying the mouth was used as a test, and relevant ablation experiments were designed and carried out. Specifically, it includes independent analysis and combined analysis of the multi-scale feature extraction module, Z-order sampling module and feature enhancement module. The baseline is the baseline model. The multi-scale feature extraction module is added to the baseline model and represented as Model1. The Z-order sampling module is added to the baseline model and represented as Model2. The feature enhancement module is added to the baseline model and represented as Model3. The multi-scale feature extraction module and feature enhancement module are added to the baseline model and represented as Model4. The Z-order sampling module and feature enhancement module are added to the baseline model and represented as Model5. The Z-order sampling module and multi-scale feature extraction module are added to the baseline model and represented as Model6. The three modules are added to the baseline model and represented as Ours. The experimental results are shown in Table 3.
[0135] As shown in the second row of Table 3, Model 1 improves the performance of the baseline model from 85.56% to 87.78%. This demonstrates that the multi-scale feature extraction module effectively extracts multi-scale local information and interactively integrates this information with global information, thereby enhancing the network's feature representation capabilities. As shown in the third row of Table 3, Model 2 improves the performance of the baseline model from 85.56% to 87.22%, demonstrating that this module preserves important structural information during sampling. As shown in the fourth row of Table 3, Model 3 improves the performance of the baseline model from 85.56% to 89.44%, further demonstrating the module's effectiveness in enhancing feature discriminability. The effectiveness of the modules was verified by combining any two modules. As shown in the fifth row of Table 3, Model 4 achieves an accuracy of 93.89%. Combining the Z-order sampling module with the feature enhancement module also improves the performance of the model. As shown in the sixth row of Table 3, Model 5 achieves an accuracy of 91.11%. Finally, when all three modules are added to the network, the accuracy reaches the highest, as shown in row 8 of Table 3, where Ours’ accuracy is 94.44%. This shows that the proposed modules all improve the performance of the network.
[0136]
[0137] according to Figure 7 The experimental results in Table 4 show that the proposed method achieved lower error rates than the other two methods (EB-LG and DGCNN) for all four test samples, demonstrating its wide applicability and reliability in 3D palatal wrinkle point cloud recognition. Table 4 also shows that the proposed method achieved higher recognition accuracy than the other two methods (EB-LG and DGCNN) for all four test samples. When tested on 3D palatal wrinkle point cloud data (B) after rinsing and drying the mouth, the proposed method achieved an accuracy of 94.44%, significantly higher than the results of EB-LG and DGCNN. When tested on 3D palatal wrinkle point cloud data (T) containing saliva, the proposed method achieved an accuracy of 86.11%, exceeding the results of the other two methods, demonstrating that the network maintains good performance even when recognizing noisy data. When tested on 3D palatal wrinkle point cloud data (Q) using the incisor mastoid as the reference point, the proposed method achieved an accuracy of 57.22%, significantly outperforming the other methods, demonstrating that the proposed method can achieve good recognition results even with missing data. Experimental analysis demonstrates that the proposed network has strong generalization capabilities.
[0138]
[0139] The three-dimensional palatal wrinkle point cloud data after rinsing the mouth (S), the three-dimensional palatal wrinkle point cloud data with the incisor mastoid as the reference point (Q), and the three-dimensional palatal wrinkle point cloud data with saliva (T) were used as training, and the three-dimensional palatal wrinkle point cloud data after rinsing and drying the mouth (B) was used as testing. The evaluation indicators used in the experiment were accuracy (Acc) and equal error rate (EER). In order to make a clear comparison, the input data type of each method ((x, y, z) represents Euclidean coordinate information, (nx, ny, nz) represents normal vector information) and the corresponding number of input data points are displayed. The method of the present invention is then compared with other point cloud methods based on deep learning. The comparison results are shown in Tables 5 and Figure 8 As shown in the figure, specifically, for the palatal wrinkle point cloud dataset, the proposed method achieved the best accuracy of 94.11% among all the benchmark methods. In terms of the EER indicator, the proposed method also achieved the lowest value.
[0140]
[0141] The present invention proposes a palatal wrinkle point cloud recognition network based on the interaction of multi-scale global and local features. First, in order to extract representative structural information from the original point cloud, a Z-order sampling module is designed to achieve effective downsampling while retaining key geometric features. Secondly, in order to enable the network to learn richer geometric features, a multi-scale feature extraction module is proposed to capture multi-scale local information and achieve efficient fusion between local features and global features. By constructing a feature enhancement module, the network is guided to learn the importance of salient points in the channel and spatial dimensions, thereby enhancing the expression effect of key features. Through systematic analysis and verification in multiple experiments, the proposed method has achieved significant performance improvement in the three-dimensional palatal wrinkle point cloud recognition task, proving its effectiveness and superiority.
[0142] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of the present invention.
Claims
1. A palatal wrinkle point cloud recognition system based on multi-scale global-local feature interaction, characterized by: The specific steps are as follows: Step 1: Construct a palatal wrinkle point cloud recognition network consisting of a downsampling module, a feature extraction network, and a recognition head; Step 2: Sampling the 3D palatal wrinkle point cloud data through a point cloud data sampling module guided by a space-filling curve, and inputting the downsampled palatal wrinkle point cloud data into a multi-scale feature extraction backbone network; Step 3: In the feature extraction network, multiple groups of feature aggregation modules are introduced, and a multi-scale feature extraction module and a feature enhancement module are set in each feature aggregation module; The multi-scale feature extraction module consists of two parts: multi-scale local feature extraction and global-local feature interaction. In the multi-scale local feature extraction part, three KNN neighborhoods of different scales are constructed for each sampling point to extract neighborhood local features of different receptive fields, namely the semantic feature information and geometric position feature information of each sampling point and its neighborhood points. In the global-local feature interaction part, the combined features of the three different scale neighborhood features are regarded as global features, and weighted learning is performed on the global features and the neighborhood features at each scale. In the feature enhancement module, the output features from the multi-scale feature extraction module are run in parallel under the channel feature enhancement mechanism and the spatial feature enhancement mechanism. In the channel feature enhancement mechanism, the point cloud features are aggregated through maximum pooling and average pooling operations. Finally, the channel feature enhancement mechanism and the spatial feature enhancement mechanism are jointly applied to the input features to obtain the final output features. ; Step 4: The features after feature aggregation processing are sent to the recognition head, and the point cloud recognition results are output after processing by the recognition head.
2. The palatal wrinkle point cloud recognition system based on multi-scale global-local feature interaction according to claim 1 is characterized in that: The downsampling module is used to downsample the original point cloud data; the feature extraction network is used to perform feature representation on the downsampled point cloud data; and the recognition head processes and analyzes the extracted features to output the final recognition result.
3. The palatal wrinkle point cloud recognition system based on multi-scale global-local feature interaction according to claim 1 is characterized in that: The multi-scale local feature extraction stage constructs three KNN neighborhoods of different scales for each sampling point. Each neighborhood of By combining its semantic feature information and geometric position feature information To generate neighborhood fusion features : (1) The graph convolution in the neighborhood is formulated as: (2) in, yes The local fusion features of The convolution operation has a kernel size of 1×1. Indicates the Neighborhood points and The relationship between them is as follows: (3) in, and Respectively represent the semantic feature information and geometric position feature information of the neighborhood range, and define different output dimensions of graph convolution for different scale neighborhoods , Used to construct sampling points The semantic and geometric relationship between the points and the neighborhood, using the maximum pooling operation to fuse the local information to the sampling point , given sampling points Local fusion features at each scale , generated by cascading operations Global characteristics of , specifically expressed as: (4) in, Indicates sampling point Three local features at three different scales; Global-local feature interaction combines the global features output from the multi-scale local feature extraction stage As input, Indicates the number of sampling points, express The characteristic dimension of Projecting into two different feature spaces to generate the Query matrix and the Key matrix, which are abbreviated as Q and K respectively during the calculation process, namely: (5) in, is a learnable weight matrix that calculates the attention map of the global and local feature interaction stage , the specific formula is: (6) in, represents the query matrix and key matrix, is the learnable position encoding matrix, express The characteristic dimension of The features after output processing are: (7) in, Represents the value matrix, which is the neighborhood feature map at each scale, represents the maximum pooling operation; Connect at every scale , to obtain global-local interaction features : (8) in, Generated by neighborhood feature maps of different scales; Afterwards, the global-local interaction features It is sent to the LBR module consisting of a linear layer, a batch normalization layer, and a RELU activation function layer, and the obtained features are added to the multi-scale features to obtain the final output features. : (9)。 4. The palatal wrinkle point cloud recognition system based on multi-scale global-local feature interaction according to claim 1, characterized in that: The feature enhancement module runs the output features from the multi-scale feature extraction module in parallel under the channel feature enhancement mechanism and the spatial feature enhancement mechanism. In the channel feature enhancement mechanism, the point cloud features are aggregated by the maximum pooling and average pooling operations, and then forwarded to and , the specific formula is as follows: (10) (11) in, and Represent the maximum pooling and average pooling functions respectively; The output of the two convolutional layers The sum is then passed through the sigmoid function to learn the weight of each feature channel. The channel feature enhancement mechanism is defined as: (12) In the spatial feature enhancement mechanism, features are input into the MLP to change the channel dimension, and information about each spatial position is aggregated through the batch normalization layer and the pooling layer. The spatial position attention weight is obtained through the sigmoid function. The spatial feature enhancement mechanism is defined as: (13) Apply the channel feature enhancement mechanism and the spatial feature enhancement mechanism together to the input feature to obtain the output feature for: (14)。