A multi-stream feature fusion system for three-dimensional palate collapse identification

CN118366145BActive Publication Date: 2026-09-25STOMATOLOGICAL HOSPITAL OF SHANXI MEDICAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311534495.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2026-09-25
Estimated Expiration
2043-12-12

AI Technical Summary

Technical Problem

由于腭皱图像的复杂性和变化性,很难通过手动选择特征点或选择合适的纹理分析方法来准确提取和描述腭皱特征

Benefits of technology

[0037]本发明与现有技术相比,具体有益效果体现在:本发明提出了一种多流特征融合网络,使用三条支路来分别提取不同的腭皱特征,使用经过Gabor滤波器处理的细节来增强腭皱图像进行约束的细节特征提取支路,一个残差特征流的每个残差块的全局特征和一个稠密连接支路提取的空间信息融合后对细节特征进行补充,提取了多种腭皱特征的三条流支路融合在一起,得到用于腭皱识别的丰富特征信息,提取了多种腭皱特征进行融合以用于高效腭皱识别,本发明将腭皱图像转换为二进制码,减少了存储空间,降低了匹配时间。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118366145B_ABST
    Figure CN118366145B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of biometric feature recognition image processing, and specifically discloses a multi-stream feature fusion network for three-dimensional palate wrinkle recognition; the specific technical scheme is as follows: the multi-stream feature fusion network comprises parallelly distributed detail feature stream branches, residual feature stream branches and dense feature stream branches; the detail feature stream branches are convolutional neural networks constrained by detail feature maps subjected to Gabor filtering; the detail feature stream branches are used for learning palate wrinkle detail texture features; the residual feature stream branches are based on global feature aggregation blocks and are used for learning global features; the dense feature stream branches are based on spatial features; the feature connection modules are used to fuse any two feature stream branches; an end-to-end network is formed; the three feature stream branches of various palate wrinkle features are fused together to obtain feature information for palate wrinkle recognition, so as to realize efficient palate wrinkle recognition; the palate wrinkle image is converted into binary code, thereby reducing the storage space and the matching time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biometric recognition image processing technology, specifically relating to a multi-stream feature fusion network for three-dimensional palatal wrinkle recognition. Background Technology

[0002] In the digital society, individual identity authentication is becoming increasingly important for information security. However, some traditional methods, such as keys and passwords, have significant drawbacks, such as being easily forgotten or vulnerable to acquisition by criminals. In recent years, biometric identification has been recognized as one of the most important and effective personal identity authentication solutions. Biometrics is an effective technology that uses physiological or behavioral characteristics of the human body for identity verification. Typically, biometric identification systems have two modes of identification: recognition and verification. Recognition is a one-to-many comparison, while verification is a one-to-one comparison. With the continuous development of technology, biometric identification technology plays an increasingly important role in security and individual identity verification. Traditional identification methods such as fingerprint recognition and facial recognition have been widely adopted, but in certain special circumstances, these methods may have limitations, such as damaged fingerprints, obscured faces, or altered facial features.

[0003] Palatal crease recognition, as an emerging biometric identification technology, possesses characteristics such as universality, uniqueness, stability, and difficulty in forgery, thus attracting increasing attention from researchers. Palatal creases, covered by muscle tissue and mucosa, are located in the anterior part of the hard palate in the human oral cavity. They are irregular, asymmetrical mucosal ridges extending posteriorly from the incisor papillae and laterally from the mid-palatal suture, with 3-7 creases on each side, bounded by the mid-palatal suture. Their unique and stable shape is not easily affected by external environmental interference. By analyzing the morphological and textural features of palatal creases, accurate individual identification and identity verification can be achieved.

[0004] However, traditional palatal wrinkle feature recognition methods typically rely on manual feature extraction and analysis, which is subject to subjectivity and significant errors. To further improve the accuracy and efficiency of palatal wrinkle feature recognition, introducing deep learning technology has become a viable solution. Deep learning technology, with its excellent feature learning and pattern recognition capabilities, has achieved remarkable results in fields such as image, speech, and natural language processing. By using deep neural networks, high-level abstraction and effective recognition of palatal wrinkle features can be achieved through training on large-scale datasets and automatic feature extraction. However, current research on deep learning-based palatal wrinkle feature recognition methods is relatively limited and requires further exploration and improvement.

[0005] Morphological analysis methods were among the earliest applied to palatal crease feature recognition. Common morphological analysis methods include: feature point-based methods, which manually select or automatically detect feature points of the palatal crease in the image, such as intersections and corners, to construct feature vectors describing the morphology of the palatal crease, and then use these feature vectors for palatal crease matching and authentication; and texture analysis-based methods, which extract and analyze texture features, such as gray-level co-occurrence matrix and wavelet transform, to obtain the texture features of the palatal crease image, and then use these texture features for palatal crease matching and recognition. While these morphological analysis methods are simple and intuitive, they suffer from strong subjectivity and limited effectiveness. Due to the complexity and variability of palatal crease images, it is difficult to accurately extract and describe palatal crease features by manually selecting feature points or choosing appropriate texture analysis methods. Traditional palatal crease feature recognition methods have certain limitations and subjectivity in terms of manual operation and feature extraction, and cannot fully explore and utilize the high-level feature information in palatal crease images. Summary of the Invention

[0006] To address the technical problems existing in the prior art, this invention provides a multi-stream feature fusion network for three-dimensional palatal wrinkle recognition, which can quickly, accurately, and stably identify palatal wrinkle features and obtain rich feature information for palatal wrinkle recognition.

[0007] To achieve the above objectives, the technical solution adopted in this invention is as follows: a multi-flow feature fusion system for three-dimensional palatal wrinkle recognition, comprising parallel distributed detail feature flow branches, residual feature flow branches, and dense feature flow branches. The detail feature flow branches are convolutional neural networks constrained by detail feature maps processed by Gabor filtering. The detail feature flow branches are used to learn palatal wrinkle detail texture features. The residual feature flow branches are based on global feature aggregation blocks and are used to learn global features. The dense feature flow branches are based on spatial features. A feature connection module is used to fuse any two feature flow branches among the detail feature flow branches, residual feature flow branches, and dense flow feature flow branches to form an end-to-end network.

[0008] In the detail flow branch, there are 5 convolutional layers. Gabor filters are used to process the palatal wrinkle image, capturing edge and texture information. The detail flow encoded features in the edge and texture information are then processed... The loss is constrained, and the constrained features are concatenated with the features of the detail flow encoding part to obtain the final detail flow features.

[0009] In the detail flow branch, the outputs of the 5 convolutional layers are converted through both encoding and decoding. The dimension concatenates the encoded and decoded data to transform the palatal wrinkle detail image obtained from the Gabor filter into... Dimension, through the loss function Apply constraints:

[0010]

[0011] in, For image Both encoding and decoding are converted to Features in dimensionality, For image Detailed features after Gabor filter processing The distance, i.e., the Manhattan distance.

[0012] In the residual feature flow branch, the residual units of the residual feature flow branch have Conv, BN, and ReLU layer connections. The input of the residual unit is addition. Each residual unit... Defined as:

[0013] ;

[0014] in, For the first The input of each unit, For the first The output of each unit The output obtained after the input passes through the residual unit is the series of operations within the residual unit.

[0015] The global pooling descriptor obtains the output from each residual unit. , The size is ,have Each feature map, after processing, results in a global average pooling descriptor size that is... Shrink to The specific processing procedure is as follows:

[0016]

[0017] It is the location in FM The value at that position, FM is the output obtained after pooling. and These refer to the coordinates of different positions in the obtained feature map. and It is the height and width of the feature map, i.e. The residual blocks 1, 2, and 3 yielded The features are generated in dimensions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 ... , , Features refer to Figure 1 The output obtained from the third branch after global average pooling (GAP) is... It is the number of channels in the feature map. These refer to the number of channels in the three different outputs obtained after global average pooling at different positions in the branch of the graph. Connect to represent residual flow branches The global feature aggregation block is specifically represented as follows:

[0018]

[0019] Finally, output As one of the inputs to the feature connection module, the residual feature flow branch contains three residual units, and the global feature concatenation of each residual unit serves as the input to the feature connection module.

[0020] In dense characteristic flow branches, the first in the dense block Convolutional layers receive The representation uses all the feature maps from the previous layer as input:

[0021]

[0022] In the formula, Indicates the previous layer The connection of feature maps, It is a composite function.

[0023] A bottleneck layer is used within each dense block of a dense feature flow branch. The bottleneck layer consists of a A convolutional layer and a The convolutional layers consist of;

[0024] Transition layers are used between consecutive dense blocks to reduce the feature map dimensionality and spatial size. The transition layer consists of a 1×1 convolutional layer and a... The average pooling layer composition;

[0025] The dense feature flow branch contains four dense blocks, denoted as dense block 1, dense block 2, dense block 3 and dense block 4, respectively. The spatial features of the last dense block are used as the input of the feature connection module.

[0026] Each produce Feature map of the number of elements, the first Each convolutional layer has Input feature map, where, The number of channels for dense block input, a hyperparameter. This represents the growth rate.

[0027] In the feature fusion module, information from the residual feature flow branch and the dense feature flow branch is first concatenated in the feature connection module. The concatenated feature information is then combined with detail features through another feature connection module. In the detail feature flow branch, the final feature vector dimension is 128. The dimension of the concatenated global residual features and spatial dense features is set to 128, and the final feature dimension output by the feature connection module is 256. The 256-dimensional feature is used as the final encoding of the sample. By calculating the Euclidean distance between different sample encodings, it is determined whether two sample encodings belong to the same class, thereby performing matching and recognition.

[0028] For images and images Assuming the extracted features are represented as and The loss function is then defined as:

[0029]

[0030] in, For matching tags, when the image and When a true match is formed, ,otherwise, ; for and The distance between them is calculated using Euclidean distance; This is the marginal threshold;

[0031] Quantization loss is defined as:

[0032]

[0033] in, Representing an image The code, assuming it has One image, The distance;

[0034] The optimization objective is:

[0035]

[0036] in, and These are parameters that determine the weights between the various losses. Let be the total number of samples. The total loss for all samples is calculated by summing the results. In the formula, for A pair The total loss obtained by summation is the sum of the losses between all pairs; For each sample The total quantization loss obtained after summation is, i.e. indivual The total quantization loss obtained after summation is .

[0037] Compared with existing technologies, the specific beneficial effects of this invention are as follows: This invention proposes a multi-stream feature fusion network that uses three branches to extract different palatal wrinkle features. A detail feature extraction branch uses details processed by a Gabor filter to enhance the palatal wrinkle image under constraint. The global features of each residual block in a residual feature stream and the spatial information extracted by a dense connection branch are fused to supplement the detail features. The three stream branches that extract multiple palatal wrinkle features are fused together to obtain rich feature information for palatal wrinkle recognition. This invention extracts and fuses multiple palatal wrinkle features for efficient palatal wrinkle recognition. Furthermore, this invention converts the palatal wrinkle image into binary code, reducing storage space and matching time. Attached Figure Description

[0038] Figure 1 This is a flowchart illustrating the processing of palatal wrinkle images according to the present invention.

[0039] Figure 2 This is a structural diagram of the residual cells used in the residual feature flow.

[0040] Figure 3 This is a diagram showing the connectivity of convolutional layers within a dense block.

[0041] Figure 4 This is a three-dimensional image of the palate folds captured by a scanner.

[0042] Figure 5 Forty images of equidistant cross-sectional curves of the palatal folds.

[0043] Figure 6 The image after being stitched together to classify odd and even numbers.

[0044] Figure 7 This is a distance distribution map obtained from the palate wrinkle recognition algorithm. Figure 7 (a) is a distribution diagram of the T domain. Figure 7 (b) is a distribution map of domain B. Figure 7 (c) is a distribution diagram of the S-domain. Figure 7 (d) is the distribution diagram of the Q domain.

[0045] Figure 8 The diagram shows the network performance under different growth rates. Detailed Implementation

[0046] To make the technical problems to be solved, the technical solutions, and the beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0047] like Figure 1 As shown, a multi-stream feature fusion system for 3D palatal wrinkle recognition includes parallel arrangement of detail feature stream branches, residual feature stream branches, and dense feature stream branches. The detail feature stream branches are CNNs constrained by detail maps processed by Gabor filters. The detail stream branches are used to learn palatal wrinkle detail texture features. The residual feature stream branches are based on global feature aggregation blocks and are used to learn global features. The dense feature stream branches are based on spatial features. A feature connection module is used to fuse any two of the detail feature stream branches, residual feature stream branches, and dense feature stream branches to form an end-to-end network called GtRDNet.

[0048] In the detail flow branch, there are 5 convolutional layers. In order to make the features extracted by this branch focus on the detailed texture features of the palatal wrinkles, a Gabor filter is used to process the palatal wrinkle image to capture the edge and texture information in the image. The detail flow encoded features in the edge and texture information are constrained by L1 loss, which is the loss represented by L1 distance. The constrained features are concatenated with the detail flow encoded features to obtain the final detail flow features.

[0049] In the detail flow branch, the outputs of the 5 convolutional layers are converted through both encoding and decoding. Dimension, concatenating the encoded and decoded data, transforms the jaw wrinkle detail image obtained from Gabor filter processing into... Dimension, through the loss function Apply constraints:

[0050]

[0051] in, For image Both encoding and decoding are converted to Features in dimensionality, For image Detailed features after Gabor filter processing yes Distance, also known as Manhattan distance.

[0052] like Figure 2As shown, in CNNs, multiple convolutional layers are stacked together to form a deeper neural network, and deeper layers can learn more discriminative features from the input image. However, as the number of network layers increases, the gradient gradually diminishes during backpropagation, making it difficult to optimize deep networks. To address this gradient vanishing problem, a residual learning technique using residual blocks is proposed in the CNN architecture. Residual networks introduce "skip connections," meaning cross-layer connections are introduced into the network, allowing gradients to propagate faster. By directly adding the input to the output of a middle layer of the network, residual networks can effectively pass gradients, making the training of deep networks more efficient and stable. In addition to solving the gradient problem, residual networks can also better fit nonlinear mappings and improve the network's representational power.

[0053] In the residual feature flow branch, the residual units of the residual feature flow branch have typical Conv, BN, and ReLU layer connections. The input of the residual unit is addition, and each residual unit... Defined as:

[0054]

[0055] in, Indicates the first The input of each unit, Indicates the first The output of each unit This represents the output obtained after the input passes through the residual unit, which is a series of operations within the residual unit.

[0056] The global pooling descriptor obtains the output from each residual unit. , The size is ,have Each feature map, after processing, results in a global average pooling descriptor size that is... Shrink to The specific processing procedure is as follows:

[0057]

[0058] yes Middle position The value at that position, FM is the output obtained after pooling. and These refer to the coordinates of different positions in the obtained feature map. and It is the height and width of the feature map, i.e. The residual blocks 1, 2, and 3 yielded The features are generated in dimensions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 ... , , Features refer to Figure 1 The output obtained from the third branch after global average pooling (GAP) is... It is the number of channels in the feature map. These refer to the number of channels in the three different outputs obtained after global average pooling at different positions in the branch of the graph. Connect to represent residual flow branches The global feature aggregation block is specifically represented as follows:

[0059]

[0060] Finally, output As one of the inputs to FCM, the residual flow branch contains three residual units, and the global features of each residual unit are concatenated as the input to FCM.

[0061] FCM stands for Feature Connection Module.

[0062] like Figure 4 As shown, in ResNet, a fixed number of output feature maps are generated from the transformation layer within the residual block. Then, the feature maps from the first layer are added to the feature maps from the second layer using an identity function, and so on. Therefore, this network improves gradient flow and feature propagation using a summation function. The core of DenseNet (Dense Convolutional Network) is dense connectivity, which concatenates the feature maps of all preceding layers with those of subsequent layers. This densely connected architecture ensures that the input of each layer contains feature information from all preceding layers, thus promoting information transfer and feature reuse. Dense connectivity improves information flow and gradients by connecting the feature maps of preceding layers with those of the current layer using channel-dimensional concatenation operations.

[0063] like Figure 3 As shown, in a dense flow branch, the first [unit] in the dense block Convolutional layers receive The representation uses all the feature maps from the previous layer as input:

[0064]

[0065] In the formula, Indicates the previous layer The connection of feature maps, This is a composite function, defined using three consecutive operations (BN, ReLU, and convolution).

[0066] Bottleneck layer: To reduce the number of parameters and computational complexity of the model, a bottleneck layer is used within each dense block of a dense flow branch. The bottleneck layer consists of a... A convolutional layer and a It consists of convolutional layers to reduce the number of channels. The bottleneck layer can compress the features of the previous layers to a lower dimension, thereby reducing the number of parameters and computational load.

[0067] Transition Layer: Its function is to reduce the dimensionality and spatial size of the feature map by using transition layers between consecutive dense blocks. A transition layer consists of a... A convolutional layer and a The transition layer consists of average pooling layers to reduce the number of channels and size of the feature maps. By reducing the dimensionality of the data, the transition layer can control the size of the model, reduce the number of parameters, and reduce the complexity of the model. The stream contains four dense blocks, denoted as dense block 1, 2, 3, and 4, with the spatial features of the last dense block serving as the input to the FCM.

[0068] Growth Rate: The growth rate is a crucial hyperparameter in DenseNet, controlling the number of feature maps generated per layer. The growth rate determines the number of channels in the newly generated feature map within each dense block. Increasing the growth rate increases the information transfer rate of dense connections, improving network performance. In the network, each... produce The feature map of the number of elements. Therefore, the first... Each convolutional layer has Input feature map, where The number of channels input to the dense block. This hyperparameter... Known as the growth rate, The value is set to 32.

[0069] In the feature fusion module, the two streams of information from the residual feature stream branch and the dense feature stream branch are first concatenated in the feature connection module. The concatenated feature information is then combined with the features of the detail stream through an FCM. In the detail stream branch, the final feature vector dimension is 128. The dimension of the concatenated global residual features and spatial dense features is set to 128, and the final feature dimension of the FCM output is 256. The 256-dimensional feature is used as the final encoding of the sample. By calculating the Euclidean distance between different sample encodings, it is determined whether they belong to the same class, thereby performing matching and recognition.

[0070] For images and images Assuming the extracted features are represented as and The loss function is then defined as:

[0071]

[0072] in, It is a paired label, when the image and When a true match is formed, ,otherwise, ; express and The distance between them is calculated using Euclidean distance; It is a marginal threshold, set to 180;

[0073] Quantization loss is the loss incurred due to converting features into binary code using a sign function, and is specifically defined as:

[0074]

[0075] in, Representing an image The code assumes there are N images. yes Distance, also known as Euclidean distance.

[0076] The overall loss of the entire network is then the weighted sum of the three losses mentioned earlier:

[0077]

[0078] in, and These are parameters that determine the weights between the various losses. The total number of samples, and The labels of the samples are used to calculate the total loss of all samples by summing them.

[0079] The fusion network structure of this invention was applied to the recognition of a palatal wrinkle database, and a palatal wrinkle database was established. A large palatal wrinkle database was collected using 3shape, and 13,600 palatal wrinkle images of 85 individuals were collected under different conditions. This invention performed a series of preprocessing operations on the palatal wrinkle images, including cropping, enhancement, noise reduction and stitching.

[0080] Simultaneously, the fusion network structure of this invention is applied to the palatal wrinkle recognition algorithm (GtRDNet), which converts palatal wrinkle images into binary codes, improving the efficiency of feature matching. To extract sufficient palatal wrinkle features for better recognition, the network of this invention uses three branches to extract different palatal wrinkle features respectively.

[0081] like Figures 4 to 6As shown, the palatal fold database collected data from 85 individuals. Samples collected before saliva removal were defined as T, samples collected after not rinsing were defined as S, samples collected under standard conditions (rinsing and drying) were defined as B, and samples where no teeth were collected and only the incisor papillae were included were defined as Q. To facilitate the extraction of palatal fold features by the deep learning network, this invention slices each 3D palatal fold image equidistantly, resulting in 40 slice images, for a total of 13,600 images. This achieves the transformation of 3D images into 2D slice images. Then, this invention performs a series of preprocessing operations on the palatal fold images, including cropping, enhancement, and noise reduction. To enable the network to extract complete palatal fold features, this invention classifies each complete palatal fold slice image into odd and even sorting, and then stitches the 20 odd-numbered images and the 20 even-numbered images together into a single large image. In this way, the network can extract complete palatal fold features at once, and the odd-even sorting also expands the dataset.

[0082] In the experiments, this invention trains the network using supervised loss. During testing, palatal wrinkle recognition and verification are performed on datasets under different conditions. For each category, an odd half of the palatal wrinkle images are used as the training set, and the remaining half as the test set to evaluate the network's performance. For palatal wrinkle recognition, after feature extraction, each test image is matched against all training images to find the most similar one. If they belong to the same object, the recognition is considered successful. For palatal wrinkle verification, test images are also matched against training images to calculate the Hamming distance. The false acceptance rate can then be calculated. False rejection rate and average error rate The experiment was implemented using the TensorFlow framework on an NVIDIA GTX 2080 graphics processor with 11GB of memory. The base learning rate was set to 0.001, and the Adam optimizer and stochastic gradient descent method were employed. .

[0083] like Figure 7 As shown, the present invention uses TO, SO, BO and TE, SE, BE, and QE (Even, E) were used as the training set, and QO, SE, BE, and QE (Even, E) were used as the test set. During testing, this invention used TO, SO, BO, and QO as benchmarks in the database, while TE, SE, BE, and QE were used as input samples for the network to be recognized. The final results are shown in Table 1. It can be seen that the network achieves the highest accuracy in the T domain, with a recognition accuracy reaching [percentage missing]. The average error rate is lowest in domain B, and the recognition accuracy is [missing information]. The effect across domains is somewhat different from that within the same domain.

[0084] Table 1. Accuracy on the Palate Wrinkle dataset and average error rate

[0085]

[0086] To further demonstrate the network's recognition capabilities more intuitively, the distributions of fake and real sample matching on different datasets are shown as follows: Figure 7 As shown in the diagram, it is clear that the distributions of fake and real sample matches are mostly separable, with a small overlap. These distribution plots fully demonstrate the good discriminative ability of the proposed method. In this embodiment, several experiments were conducted on the T database to demonstrate the main hyperparameters used to balance different loss weights. and The effects are shown in Table 2. The results indicate that when... At that time, the network performed best, with accuracy and mean error rate of [missing values]. and .

[0087] Table 2. Palmprint recognition accuracy with different parameters and average error rate

[0088]

[0089] In deep learning, max pooling primarily focuses on extracting the most salient features from an image or feature map, selecting the largest value from each pooling window as the pooling result. This preserves the strongest signal and edge information, making max pooling highly effective for detecting object location and texture details. Average pooling, on the other hand, focuses on extracting overall features from the image or feature map. It calculates the average pixel value within each pooling window as the pooling result, which is better at preserving overall background information, reducing noise, and mitigating overfitting. This embodiment conducted several experiments to demonstrate the impact of different uses of max pooling and average pooling on the results in detail flow. The results are shown in Table 3. The best results were achieved when average pooling was used in the first layer and max pooling was used in the other three layers.

[0090] Table 3. Impact of Average Pooling and Max Pooling on the Network

[0091]

[0092] like Figure 8 As shown, the growth rate is the number of feature maps added to the global state by each convolutional layer in a dense block of the network. In this invention, the performance on the benchmark dataset was evaluated using four different growth rate values: 4, 8, 16, and 32. The impact of performance on the growth rate variation is shown below. Figure 8As shown, for a growth rate of 32, the recognition effect of palatal wrinkles is superior. This embodiment conducts an ablation study to evaluate the effectiveness of the proposed network using different branches. This network consists of modules such as detail flow, Gabor detail constraint module, residual feature flow, and dense feature connection flow. The accuracy of each module is... The average error rate is shown in Table 4. The results show that the overall performance of the network is greatly improved after adding the Gabor detail constraint module. In addition, the fusion of the three branches shows better performance than the individual flow, demonstrating the effectiveness of the network on the palatal wrinkle dataset.

[0093] Table 4 Ablation studies of different parts

[0094]

[0095] This embodiment evaluates the performance of the network of the present invention compared with other CNNs on the palatal wrinkle dataset, and compares the network's recognition accuracy with other fine-tuned CNNs. The comparison is shown in Table 5. As can be seen from Table 5, the CNN architecture proposed in this embodiment can achieve an accuracy of 87.06%, which is better than other CNNs, demonstrating the superiority of the framework of this invention.

[0096] Table 5. Recognition accuracy of different methods

[0097]

[0098] This invention proposes a multi-stream fusion network called GtRDNet for accurate palatal wrinkle identification in vision-based environments. It combines detailed features constrained by detailed texture features processed by Gabor filters from one stream, global information from each residual block, and dense spatial information from another stream. The combination of these three streams explores accurate palatal wrinkle information in palatal wrinkle images. Through various parameter studies, ablation studies, and qualitative and quantitative analyses, it is shown that the proposed model outperforms other baseline models on the palatal wrinkle dataset of this invention.

[0099] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included within the scope of the present invention.

Claims

1. A multi-flow feature fusion system for three-dimensional palatal wrinkle recognition, characterized in that, It includes parallel distributed detail feature flow branches, residual feature flow branches, and dense feature flow branches. The detail feature flow branches are convolutional neural networks constrained by detail feature maps processed by Gabor filtering. The detail feature flow branches are used to learn the detailed texture features of the palatal wrinkles. The residual feature flow branches are based on global feature aggregation blocks and are used to learn global features. The dense feature flow branches are based on spatial features. The feature connection module is used to fuse any two of the detail feature flow branches, residual feature flow branches, and dense feature flow branches to form an end-to-end network. In the detail feature flow branch, there are 5 convolutional layers. Gabor filters are used to process the palatal wrinkle image to capture the edge and texture information in the image. The detail feature flow encoding part of the edge and texture information is constrained by L1 loss. The constrained features are concatenated with the detail feature flow encoding part to obtain the final detail features. In the residual feature flow branch, the residual unit has Conv, BN, and ReLU layer connections. The input of the residual unit is addition. Each residual unit... Defined as: ; in, For the first The input of each unit, For the first The output of each unit The output is obtained after the input passes through the residual unit; The global pooling descriptor obtains the output from each residual unit. , The size is ,have Each feature map, after processing, results in a global average pooling descriptor size that is... Shrink to The specific processing procedure is as follows: ; It is the location in FM The value at that position, FM is the output obtained after pooling. and These refer to the coordinates of different positions in the obtained feature map. and These are the height and width of the feature map, obtained from residual block 1, residual block 2, and residual block 3. The features are generated in dimensions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 ... , , Features This refers to the output obtained by global average pooling in the third branch of Figure 1. It is the number of channels in the feature map. These refer to the number of channels for the three different outputs obtained after global average pooling at different positions in the branch in the diagram. Connect to represent residual flow branches The global feature aggregation block is specifically represented as follows: ; Finally, output As one of the inputs to the feature connection module, the residual flow branch contains three residual units, and the global feature concatenation of each residual unit serves as the input to the feature connection module. A bottleneck layer is used within each dense block of a dense feature flow branch. The bottleneck layer consists of a A convolutional layer and a The convolutional layers consist of; Transition layers are used between consecutive dense blocks to reduce feature map dimensionality and spatial size. A transition layer consists of a... A convolutional layer and a The average pooling layer composition; The dense flow branch contains four dense blocks, denoted as dense block 1, dense block 2, dense block 3 and dense block 4, respectively. The spatial features of the last dense block are used as the input of the feature connection module. Each produce Feature map of the number of elements, the first Each convolutional layer has Input feature map, where, The number of channels for dense block input, a hyperparameter. This represents the growth rate.

2. The multi-flow feature fusion system for three-dimensional palatal wrinkle recognition according to claim 1, characterized in that, In the detail feature flow branch, the outputs of the 5 convolutional layers are converted through both encoding and decoding. The dimension concatenates the converted encoded and decoded features to transform the palatal wrinkle detail image obtained from the Gabor filter into... Dimension, through the loss function Apply constraints: ; in, For image Both encoding and decoding are converted to Features in dimensionality, For image Detailed features after Gabor filter processing The distance, i.e., the Manhattan distance.

3. A multi-flow feature fusion system for three-dimensional palatal wrinkle recognition according to claim 2, characterized in that, In dense characteristic flow branches, the first in the dense block Convolutional layers receive The representation uses all the feature maps from the previous layer as input: ; In the formula, Indicates the previous layer The connection of feature maps It is a composite function.

4. A multi-flow feature fusion system for three-dimensional palatal wrinkle recognition according to claim 3, characterized in that, In the feature fusion module, information from the residual feature flow branch and the dense feature flow branch is first concatenated in the feature connection module. The concatenated feature information is then combined with the output of the detail feature flow through another feature connection module. In the detail feature flow branch, the final feature vector dimension is 128. The dimension of the concatenated global residual features and spatial dense features is set to 128, and the final feature dimension output by the feature connection module is 256. The 256-dimensional feature is used as the final encoding of the sample. By calculating the Euclidean distance between different sample encodings, it is determined whether two sample encodings belong to the same class, thereby performing matching and recognition.

5. A multi-flow feature fusion system for three-dimensional palatal wrinkle recognition according to claim 4, characterized in that, For images and images Assuming the extracted features are represented as and The loss function is then defined as: ; in, For matching tags, when the image and When a true match is formed, ,otherwise, ; for and The distance between them is calculated using Euclidean distance; This is the marginal threshold; Quantization loss is defined as: ; in, Representing an image The code, assuming it has One image, The distance; The optimization objective is: ; in, and These are parameters that determine the weights between the various losses, where N is the total number of samples. For N pairs The total loss obtained by summation is the sum of the losses between all pairs; For each sample The total summation loss is obtained as follows: indivual The total quantization loss obtained after summation is .