Oral leukoplakia canceration progress recognition method based on deep learning

Through the dynamic feature calibration of the ConvNeXt backbone network and feature refining module, combined with the two-stage fine-tuning strategy, the problem of identification of oral leukoplakia carcinoma in the existing technology is solved, and high-precision and rapid oral leukoplakia carcinoma progress evaluation is achieved.

CN120471838APending Publication Date: 2025-08-12THE FIRST AFFILIATED HOSPITAL OF ANHUI MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510497226.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the transitional stage of oral leukoplakia cancer. Traditional convolutional neural networks lack the ability to capture subtle features, resulting in limited generalization performance of the model in clinical complex scenarios.

Method used

The ConvNeXt backbone network and feature refining module are used to generate optimized features through dynamic clustering and attention weighting, and combined with two-stage fine-tuning strategies and gradient descent algorithm training models to achieve a three-level intelligent assessment of the progression of oral leukoplakia carcinoma.

Benefits of technology

Accurate identification of early lesions and advanced cancers is achieved, the clinical applicability and recognition accuracy of the model are improved, the identification gap in the transition stage is filled, and a non-invasive and efficient cancer evaluation scheme is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471838A_ABST
    Figure CN120471838A_ABST
Patent Text Reader

Abstract

The invention discloses an oral leukoplakia canceration progress recognition method based on deep learning. The method comprises the steps of performing preprocessing of unified size zooming, pixel normalization and data enhancement on oral clinical photos; constructing a three-classification network architecture composed of a ConvNeXt backbone network, a dynamic feature refining module and a classification head, wherein the feature refining module realizes feature optimization through a learnable clustering center and a transformation matrix; performing model training by adopting a two-stage fine tuning strategy, fixing backbone network parameters in the first stage and only updating refining module and head parameters, and jointly optimizing all parameters in the second stage; and inputting the preprocessed to-be-diagnosed image into the trained network, and outputting a normal / white spot / white spot canceration classification result and a confidence score. According to the method, three-level intelligent evaluation of oral lesions of normal-precancerous-cancerous is realized for the first time, the classification accuracy of the test set reaches 90% or above, and the method has the characteristics of high clinical applicability and high reasoning efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence-assisted medical diagnosis technology, and in particular to a method for identifying the progression of oral leukoplakia canceration based on deep learning. Background Art

[0002] Oral leukoplakia is a key indicator of oral precancerous lesions. Its early identification and assessment of malignant progression are of great significance for clinical diagnosis and treatment. In the field of medical imaging diagnosis, oral endoscopy images and clinical photographs are the primary basis for screening leukoplakia. Traditional diagnostic methods rely heavily on pathological biopsy and physician experience, and have limitations such as high invasiveness, high subjective variability, and low identification rate of early lesions.

[0003] With the development of artificial intelligence (AI) technology, deep learning-based computer-aided diagnosis (CAD) systems have provided new insights into the identification of leukoplakia. However, existing technical solutions still face challenges. Most current research is limited to binary classification of oral cancer, failing to identify leukoplakia, a critical transitional stage of precancerous lesions. Furthermore, traditional convolutional neural networks are unable to capture the subtle features of leukoplakia lesions, limiting the model's generalization performance in complex clinical scenarios. Therefore, a highly accurate and efficient method for classifying oral leukoplakia cancer is urgently needed. Summary of the Invention

[0004] In order to overcome the shortcomings of the existing technology, the present invention provides a method for identifying the progression of oral leukoplakia canceration based on deep learning, aiming to improve the accuracy and clinical applicability of early cancer identification while reducing dependence on pathological biopsy and physician's subjective experience.

[0005] In order to achieve the above-mentioned purpose of the invention, the technical solutions adopted to solve the technical problems are as follows: A method for identifying the progression of oral leukoplakia cancer based on deep learning, comprising the following steps: Step 1: Model data preprocessing: The input image size of H×W is uniformly scaled to 224×224, followed by pixel value normalization and data augmentation. Step 2: Model network structure design: The model network structure includes a backbone network, a feature refinement module, and a head network. The backbone network adopts the ConvNeXt architecture and includes four hierarchical feature extraction stages to extract feature maps of different scales. The feature refinement module generates optimized features through dynamic clustering and attention weighting. The head network uses a fully connected layer to perform classification prediction on the optimized features. Step 3: Model evaluation system construction: Classification accuracy Acc and F1-score are used as indicators to evaluate model performance; Step 4, Model Training Phase: The backbone network parameters are initialized using pre-trained parameters from the large-scale public general-purpose image dataset ImageNet-21K. The parameters of the feature refinement module and the head network are randomly initialized. During the model training phase, a two-stage fine-tuning strategy and a gradient descent algorithm are used for iterative training to obtain the optimal solution. The optimal model parameters are determined based on the evaluation index values described in Step 3 on the validation set. Step 5, model inference stage: Use the optimal parameters in step 4 to load the model, input a single oral lesion image, and after preprocessing in step 1, pass it through the ConvNeXt backbone network, feature refinement module and head network in sequence to output a three-category probability vector. The maximum probability value is taken as the final prediction category, and the prediction results of normal, white spot or white spot canceration are output.

[0006] Furthermore, in step 2, the backbone network using the ConvNeXt architecture extracts features of different scales through stacked convolutional layers. This operation is mathematically expressed as: in, Indicates in The nonlinear transformation applied in the stage, Represents the backbone network The output of the stage, since ConvNeXt has four stages, ; For the subsequent classification process, only the output features of the last stage are flattened in space and input into the next part. The flattened features are: in, , Indicates the number of feature channels after flattening.

[0007] Furthermore, in step 2, the operation of the feature refinement module includes: constructing a learnable cluster center matrix , and the transformation matrix , ,in, Indicates the number of cluster centers; then the input features Normalization is performed and global features are calculated. This operation is mathematically expressed as: in, represents the normalized features, Represents global features, Representation layer normalization operation, then calculate the cosine similarity attention weight of global features and cluster centers , this operation is mathematically expressed as: in, Indicates the cluster center vectors, Represents the softmax function, using attention weights to transform the matrix Perform weighted fusion to cluster the normalized features and obtain optimized features , and finally the output features are obtained through layer normalization and MLP layer , this operation is mathematically expressed as: in, Indicates the attention weights, Indicates the A transformation matrix, Represents the MLP layer.

[0008] Preferably, in step 4, the random initialization method is: he_normal, lecun_uniform, glorot_normal, glorot_uniform or lecun_normal.

[0009] Preferably, in step 4, the two-stage fine-tuning strategy includes: In the first stage, the parameters of the backbone network ConvNeXt are kept unchanged and only the parameters of the feature refinement module and the head network are updated; In the second stage, all network structure parameters of the model are updated, and the learning rate decays by 0.1 every 7 epochs.

[0010] Preferably, in step 4, the gradient descent algorithm is: Adam, SGD, MSprop or Adadelta.

[0011] Due to the adoption of the above technical solution, the present invention has the following advantages and positive effects compared with the prior art: 1. The leukoplakia cancer progression recognition model of the present invention can accurately identify the subtle features of early lesions and the significant features of late cancer through the multi-level feature extraction of the ConvNeXt backbone network and the dynamic feature calibration of the feature refinement module, and has higher clinical applicability.

[0012] 2. In view of the complex and changeable characteristics of vitiligo lesions, the present invention innovatively designs a feature refinement module based on the attention mechanism. This module realizes dynamic recalibration of lesion features through learnable clustering centers and transformation matrices, enabling the model to adaptively focus on the image areas with the most diagnostic value.

[0013] 3. This invention realizes for the first time a three-level intelligent assessment of the progression of oral leukoplakia cancer, filling the gap in the existing technology in the lack of transition stage identification, providing key technical support for early clinical intervention, and has the advantages of high recognition accuracy, strong generalization ability, and fast reasoning speed.

[0014] 4. The deep learning technology solution based on the present invention can automatically and standardizedly complete the assessment of the malignant progression of oral leukoplakia. Compared with the traditional diagnostic method that relies on pathological biopsy and subjective judgment of physicians, the present invention provides a non-invasive, high-precision and repeatable oral leukoplakia cancer classification solution through advanced deep learning algorithms, which has the advantages of being non-invasive, efficient and highly repeatable. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings: Figure 1 This is a flow chart of a method for identifying the progression of oral leukoplakia canceration based on deep learning according to an embodiment of the present invention; Figure 2 is a schematic diagram of a network architecture according to an embodiment of the present invention; Figure 3 is a schematic diagram of a confusion matrix according to an embodiment of the present invention, showing the classification results of a test sample set; Figure 4 is a schematic diagram of the classification results of white spot samples according to an embodiment of the present invention; Figure 5 2. Schematic diagram of classification results of leukoplakia cancerous samples according to an embodiment of the present invention; Figure 6 2 is a schematic diagram of the normal sample classification result according to an embodiment of the present invention. DETAILED DESCRIPTION

[0016] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0017] Figure 1 This is a flow chart of a method for identifying the progression of oral leukoplakia cancer based on deep learning according to an embodiment of the present invention, which specifically includes the following steps: Step 1: Model data preprocessing: The input image size of H×W is uniformly scaled to 224×224, followed by pixel value normalization and data augmentation. Step 2: Model network structure design: The model network structure includes a backbone network, a feature refinement module, and a head network. The backbone network adopts the ConvNeXt architecture and includes four hierarchical feature extraction stages to extract feature maps of different scales. The feature refinement module generates optimized features through dynamic clustering and attention weighting. The head network uses a fully connected layer to perform classification prediction on the optimized features. Specifically, in step 2, the backbone network using the ConvNeXt architecture extracts features of different scales through stacked convolutional layers. This operation is mathematically expressed as: in, Indicates in The nonlinear transformation applied in the stage, Represents the backbone network The output of the stage, since ConvNeXt has four stages, ; For the subsequent classification process, only the output features of the last stage are flattened in space and input into the next part. The flattened features are: in, , Indicates the number of feature channels after flattening.

[0018] Furthermore, in step 2, the operation of the feature refinement module includes: constructing a learnable cluster center matrix , and the transformation matrix , ,in, Indicates the number of cluster centers; then the input features Normalization is performed and global features are calculated. This operation is mathematically expressed as: in, represents the normalized features, Represents global features, Representation layer normalization operation, then calculate the cosine similarity attention weight of global features and cluster centers , this operation is mathematically expressed as: in, Indicates the cluster center vectors, Represents the softmax function, using attention weights to transform the matrix Perform weighted fusion to cluster the normalized features and obtain optimized features , and finally the output features are obtained through layer normalization and MLP layer , this operation is mathematically expressed as: in, Indicates the attention weights, Indicates the A transformation matrix, Represents the MLP layer.

[0019] Step 3: Model evaluation system construction: Classification accuracy Acc and F1-score are used as indicators to evaluate model performance; Step 4, Model Training Phase: The backbone network parameters are initialized using pre-trained parameters from the large-scale public general-purpose image dataset ImageNet-21K. The parameters of the feature refinement module and the head network are randomly initialized. During the model training phase, a two-stage fine-tuning strategy and a gradient descent algorithm are used for iterative training to obtain the optimal solution. The optimal model parameters are determined based on the evaluation index values described in Step 3 on the validation set. Preferably, in step 4, the random initialization method is: he_normal, lecun_uniform, glorot_normal, glorot_uniform or lecun_normal.

[0020] Preferably, in step 4, the two-stage fine-tuning strategy includes: in the first stage, keeping the backbone network ConvNeXt parameters unchanged and only updating the feature refinement module and head network parameters; in the second stage, updating all network structure parameters of the model, and the learning rate decays by 0.1 every 7 epochs.

[0021] Preferably, in step 4, the gradient descent algorithm is: Adam, SGD, MSprop or Adadelta.

[0022] Step 5, model inference stage: Use the optimal parameters in step 4 to load the model, input a single oral lesion image, and after preprocessing in step 1, pass it through the ConvNeXt backbone network, feature refinement module and head network in sequence to output a three-category probability vector. The maximum probability value is taken as the final prediction category, and the prediction results of normal, white spot or white spot canceration are output.

[0023] Figure 2Figure 2 is a schematic diagram of the network architecture of an embodiment of the present invention. The input image is first passed through a ConvNeXt backbone network to extract feature maps. Next, the feature maps are flattened and fed into a feature refinement module to reconstruct and enhance the original features. Finally, the enhanced feature maps are fed into a probability generator to output the final prediction result.

[0024] Figure 3 The following is a diagram of the confusion matrix of an embodiment of the present invention, showing the classification results of the test sample set. The model classification accuracy for normal images reached 96.67%, the classification accuracy for vitiligo images reached 90.00%, and the classification accuracy for vitiligo cancer images reached 86.67%.

[0025] In this embodiment, the prediction results are as follows: Figure 4 、 Figure 5 、 Figure 6 As shown, Figure 4 、 Figure 5 and Figure 6 These are the prediction results of oral images of leukoplakia patients, leukoplakia cancerous patients, and normal people after model inference.

[0026] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for identifying the progression of oral leukoplakia cancer based on deep learning, characterized in that: The following steps are involved: Step 1: Model data preprocessing: The input image size of H×W is uniformly scaled to 224×224, followed by pixel value normalization and data augmentation. Step 2: Model network structure design: The model network structure includes a backbone network, a feature refinement module, and a head network. The backbone network adopts the ConvNeXt architecture and includes four hierarchical feature extraction stages to extract feature maps of different scales. The feature refinement module generates optimized features through dynamic clustering and attention weighting. The head network uses a fully connected layer to perform classification prediction on the optimized features. Step 3: Construct a model evaluation system: Acc, which reflects the overall classification accuracy, and F1-score, a comprehensive indicator that measures classification precision and recall, are used as model performance evaluation indicators. Step 4, Model Training Phase: The backbone network parameters are initialized using pre-trained parameters from the large-scale public general-purpose image dataset ImageNet-21K. The parameters of the feature refinement module and the head network are randomly initialized. During the model training phase, a two-stage fine-tuning strategy and a gradient descent algorithm are used for iterative training to obtain the optimal solution. The optimal model parameters are determined based on the evaluation index values described in Step 3 on the validation set. Step 5, model inference stage: Use the optimal parameters in step 4 to load the model, input a single oral lesion image, and after preprocessing in step 1, pass it through the ConvNeXt backbone network, feature refinement module and head network in sequence to output a three-category probability vector. The maximum probability value is taken as the final prediction category, and the prediction results of normal, white spot or white spot canceration are output.

2. A method for identifying the progression of oral leukoplakia canceration based on deep learning according to claim 1, characterized in that: In step 2, the backbone network using the ConvNeXt architecture extracts features of different scales through stacked convolutional layers. This operation is mathematically expressed as: in, Indicates in The nonlinear transformation applied in the stage, Represents the backbone network The output of the stage, since ConvNeXt has four stages, ; For the subsequent classification process, only the output features of the last stage are flattened in space and input into the next part. The flattened features are: in, , Indicates the number of feature channels after flattening.

3. The method for identifying the progression of oral leukoplakia canceration based on deep learning according to claim 1, characterized in that: In step 2, the operation of the feature refinement module includes: constructing a learnable cluster center matrix , and the transformation matrix , ,in, Indicates the number of cluster centers; then the input features Normalization is performed and global features are calculated. This operation is mathematically expressed as: in, represents the normalized features, Represents global features, Representation layer normalization operation, then calculate the cosine similarity attention weight of global features and cluster centers , this operation is mathematically expressed as: in, Indicates the cluster center vectors, Represents the softmax function, using attention weights to transform the matrix Perform weighted fusion to cluster the normalized features and obtain optimized features , and finally the output features are obtained through layer normalization and MLP layer , this operation is mathematically expressed as: in, Indicates the attention weights, Indicates the A transformation matrix, Represents the MLP layer.

4. The method for identifying the progression of oral leukoplakia canceration based on deep learning according to claim 1, characterized in that: In step 4, the random initialization method is: he_normal, lecun_uniform, glorot_normal, glorot_uniform or lecun_normal.

5. The method for identifying the progression of oral leukoplakia canceration based on deep learning according to claim 1, characterized in that: In step 4, the two-stage fine-tuning strategy includes: In the first stage, the parameters of the backbone network ConvNeXt are kept unchanged and only the parameters of the feature refinement module and the head network are updated; In the second stage, all network structure parameters of the model are updated, and the learning rate decays by 0.1 every 7 epochs.

6. The method for identifying the progression of oral leukoplakia canceration based on deep learning according to claim 1, characterized in that: In step 4, the gradient descent algorithm is Adam, SGD, MSprop or Adadelta.