Oral medical image classification method based on WGAN-GP fused CNN-Transform hybrid architecture

By using a hybrid architecture of WGAN-GP and CNN-Transformer to generate high-quality forged images and combining them with deep neural networks, the problems of insufficient generalization ability and noise interference in the classification of small-sample oral medical images are solved, thereby achieving accurate identification and improved classification accuracy of oral anatomical structures.

CN120932024APending Publication Date: 2025-11-11ANHUI MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511395568.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies lack generalization ability in small-sample medical image classification. Traditional generative adversarial networks lose high-frequency details and distort anatomical structures in the generated data. Furthermore, images captured by smartphones are easily affected by noise, resulting in poor classification accuracy and robustness.

Method used

We employ a hybrid architecture that combines WGAN-GP with CNN-Transformer to generate high-quality fake images through generative adversarial networks. By combining the global feature encoding of Transformer with the local feature extraction of CNN, we construct a deep neural image classification network to enhance data generation and classification capabilities.

Benefits of technology

It improves the classification accuracy and generalization ability of oral medical images, enabling precise identification of anatomical structures such as teeth, gums, oral mucosa, and jawbone, enhancing the stability and robustness to noise of the model, and improving the interpretability of the classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932024A_ABST
    Figure CN120932024A_ABST
Patent Text Reader

Abstract

The invention discloses an oral medical image classification method based on a WGAN-GP fused CNN-Transform hybrid architecture. The method comprises the following steps: 1, preprocessing a data set; 2, building an oral medical image generation network based on a generative adversarial network; (3) a CNN-Transform double-flow collaborative deep neural image classification network is built, and the CNN-Transform double-flow collaborative deep neural image classification network is built; 4, constructing a cross entropy loss function of the deep neural image classification network; and 5, training the deep neural image classification network through an Adam optimization algorithm, and continuously optimizing until the cross entropy loss function converges. According to the method, anatomical structures including teeth, gingiva, oral mucosa, jaw bone and the like in the oral medical image can be accurately identified, and a classification model adaptive to the oral medical image from smart phone equipment is developed according to rich local details, high global structure correlation and large contrast change of the anatomical structures; therefore, the accuracy, generalization ability and robustness of the classification model on the oral medical image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of oral medical image classification, specifically an oral medical image classification method based on a hybrid architecture of WGAN-GP and CNN-Transformer. Background Technology

[0002] The application of artificial intelligence in oral cancer diagnosis mainly relies on machine learning and deep learning. By inputting a dataset of cancer symptom images into a built model, and after training with machine learning, it can automatically identify the symptom characteristics of cancer from massive amounts of data, and ultimately arrive at a prediction result. If this artificial intelligence system is installed on mobile devices, it will provide doctors and patients with convenient and low-cost assisted diagnostic services, enabling automated screening and early detection of oral lesions.

[0003] Currently, this technology suffers from insufficient generalization ability of the model to clinical images due to limited sample size. In small-sample medical image classification scenarios, the limited number of available labeled medical image samples prevents the neural network model from being adequately trained, resulting in poor generalization ability. Traditional generative adversarial networks (GANs) suffer from severe loss of high-frequency details during data generation, making it difficult to preserve minute lesions. Directly generating images can cause abnormalities in anatomical features such as tooth alignment and mucosal texture, leading to anatomical distortion. Furthermore, this technology mainly relies on convolutional neural networks (CNNs) to extract local contextual information, which also introduces the problem of limited receptive field. In addition, due to factors such as uneven lighting, instrument occlusion, and individual differences in the clinical environment, oral lesion images taken by smartphones inevitably contain noise. Adding even small perturbations to the image can cause the classification model to make incorrect judgments, which may lead to serious consequences. Summary of the Invention

[0004] This invention addresses the shortcomings of existing technologies in preserving minute lesions in oral images, as well as the low accuracy and generalization of data classification. It proposes an oral medical image classification method based on a WGAN-GP fusion CNN-Transformer hybrid architecture. This method aims to accurately identify and classify oral medical images containing anatomical structures such as teeth, gums, oral mucosa, and jawbone, thereby improving classification accuracy, generalization ability, and robustness in oral medical images.

[0005] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The oral medical image classification method of the present invention, based on a WGAN-GP fusion CNN-Transformer hybrid architecture, is characterized by the following steps: Step 1: Obtain the dimension as Real oral medical images were obtained and normalized preprocessed to obtain a preprocessed oral medical image sequence. ,in, Indicates the first A dental image, This represents the total number of oral medical images. Indicates the height of an oral medical image. Indicates the width of the oral medicine image. This represents the number of channels in an oral medicine image; let... The true category label is denoted as ,when When, it means For an exception category, when When, it means The true normal category; Step 2: Construct a generative adversarial network-based oral medicine image generation network, including: a heatmap guide, a generator, and a scorer, and then... The images were processed to obtain a forged oral medicine image sequence. Thus obtain Authenticity rating and Authenticity rating Used to construct the gradient penalty loss function The oral medicine image generation network was trained to obtain the optimal oral medicine image generation model WGAN-GP, which was used to output the first... The optimal forged oral medicine image that preserves the spatial distribution features of teeth, gums, oral mucosa, and jawbone. Thus, the optimal set of forged oral medical images is obtained. ,in, Indicates the first A fake dental image; Step 3: Construct a CNN-Transformer dual-stream collaborative deep neural image classification network, including: a primary feature extraction module, a CNN local feature extraction module based on the ResNet50 main structure, a Transformer global feature encoding module, and a classification module, and perform... and Merged training sample set Any k-th training sample Processing is performed, and at the same time, results are obtained. Predicted normal category probability and predicting the probability of anomaly categories ;make The true category label is denoted as ,when When, it means For an exception category, when When, it means Normal category; Step 4: Construct the cross-entropy loss function of the deep neural image classification network using equation (2). : (2) In equation (2), express The number of training samples in the dataset; Step 5: Train the deep neural image classification network using the Adam optimization algorithm and calculate the cross-entropy loss function. To update all the parameters to be learned in the network, after passing through the cross-entropy loss function The process continues until convergence, thus obtaining the optimal oral medicine image classification model, which is used to perform the binary classification task of oral medicine images.

[0006] The oral medical image classification method based on a WGAN-GP fusion CNN-Transformer hybrid architecture described in this invention is characterized in that step 2 includes the following steps: Step 2.1, Heatmap Guide Processing is performed to obtain the first... A heat map ; Step 2.2, the generator processes the first... random noise vectors and Processing is performed to obtain the same as Fake oral images with the same probability distribution as the image data ; Step 2.3: The scorer respectively... and Downsampling was performed separately to obtain the true high-level abstract feature map of a single channel. And single-channel fake high-level abstract feature maps Then, they were directly flattened into Authenticity rating and Authenticity rating ; Step 2.4: Construct the gradient penalty loss function using equation (1) : (1) In equation (1), Indicates the scorer's... Expectations Indicates the scorer's... Expectations It is a gradient penalty term. It is the penalty coefficient; Step 2.5: Train the oral cavity image generation network based on generative adversarial network using the backpropagation algorithm, and calculate the gradient penalty loss function. Update all network parameters to be learned in the scorer and generator, and optimize the scorer and generator alternately until the game converges, thereby obtaining the optimal oral medicine image generation model WGAN-GP.

[0007] Furthermore, step 2.1 includes the following steps: Step 2.1.1: Utilize pre-trained... Model pair Feature extraction is performed to obtain the first... High-level feature map ; Step 2.1.2, along the channel dimension Perform mean pooling to generate the first... Single-channel characteristic response map ; Step 2.1.3: Use bicubic interpolation method to... Perform upsampling to obtain the first A heat map .

[0008] Furthermore, step 2.2 includes the following steps: Step 2.2.1, along the spatial dimension Perform global average pooling to obtain Global feature representation ; Step 2.2.2: Randomly sample the first... noise vectors , Let the dimension of the noise vector be . and After splicing along the first dimension, we get the th A joint input tensor ; Step 2.2.3, will The input generator performs stepwise feature amplification to generate the first... A fake oral disease image .

[0009] Furthermore, step 3 includes the following steps: Step 3.1: The primary feature extraction module uses initial convolutional layers and pooling layers to... Perform primary feature extraction to obtain the k-th primary feature map. ; Step 3.2: CNN Local Feature Extraction Module Based on ResNet50 Main Structure Local feature extraction is performed on the spatial distribution characteristics of teeth, gums, oral mucosa, and jawbone to obtain the k-th local feature map. ; Step 3.3: The Transformer global feature encoding module includes: a Transformer input reshaping unit, a Transformer encoding unit, a Transformer output reshaping unit, and a CNN output flattening unit, and performs... After processing, the k-th one-dimensional global association vector is obtained. and the k-th one-dimensional local detail vector ; Step 3.4: The classification module includes a feature fusion unit and two fully connected layers, and performs... and Processing is performed to obtain Predicted normal category probability and predicting the probability of anomaly categories .

[0010] Furthermore, step 3.3 includes the following steps: Step 3.3.1: The Transformer input reshaping unit will... Reconstructed into two-dimensional CNN local feature sequences ; Step 3.3.2: The Transformer encoding unit utilizes a multi-head attention mechanism and a feedforward neural network to... The process involves establishing feature dependencies between different regions of the teeth, gums, oral mucosa, and jawbone, and obtaining the k-th global feature sequence. ; Step 3.3.3, Transformer output reshaping unit pair After processing, the k-th one-dimensional global association vector is obtained. ; Step 3.3.4: Flattening the CNN output Flattening the spatial and channel dimensions yields the result... One-dimensional local detail vectors with the same dimensions .

[0011] Furthermore, step 3.4 includes the following steps: Step 3.4.1, will and After concatenation along the channel dimension, the k-th fused feature vector based on local spatial features and global relational features is obtained. ; Step 3.4.2, will The input is processed through a nonlinear transformation in the first fully connected layer, and the k-th intermediate feature vector is output. , This represents the number of neurons in the first fully connected layer. Step 3.4.3, will The input is processed in the second fully connected layer to obtain... Original classification score Then, it is input into the Softmax function for processing, and at the same time, it obtains... Predicted normal category probability and predicting the probability of anomaly categories .

[0012] The present invention provides an electronic device, including a memory and a processor, characterized in that the memory is used to store a program that supports the processor in executing the oral medicine image classification method, and the processor is configured to execute the program stored in the memory.

[0013] The present invention provides a computer-readable storage medium storing a computer program, characterized in that the computer program, when executed by a processor, performs the steps of the oral medicine image classification method.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention proposes a deep learning framework for data generation and intelligent classification of oral medical images, including: a lesion-focused generative adversarial network architecture and a CNN-Transformer dual-stream collaborative network oral lesion classification model. Based on the characteristics of rich local details, strong global structural correlation, large contrast variation, and susceptibility to metal artifacts, it can accurately identify the anatomical structure of oral images, taking into account both local details and global structure, thereby improving the ability to classify oral medical images.

[0015] 2. This invention enhances data through a lesion-focused generative adversarial network architecture, solving the problem of insufficient existing oral medical image data. It uses heatmap guidance to enable the model to focus on lesion areas, enhancing the fine control over image generation. The WGAN-GP architecture is adopted, and the gradient penalty term is introduced to constrain the gradient of the scorer, which effectively improves training stability, avoids gradient explosion, reduces dependence on network structure, and enhances generalization ability.

[0016] 3. This invention proposes a CNN-Transformer dual-stream collaborative network oral lesion classification model. This design strategy not only successfully makes up for the shortcomings of traditional single CNN models in reconstructing fine structures in sequential image tasks, but also overcomes the limitations of single Transformer models in terms of high computational complexity and difficulty in capturing local features, promoting the extraction of multi-dimensional features. The hierarchical fusion network used in the feature fusion part further improves the accuracy of classification.

[0017] 4. The present invention is based on a heatmap enhancement model generated by Grad-CAM, which helps medical personnel understand the areas of interest in classification decisions and enhances the interpretability of classification. Attached Figure Description

[0018] Figure 1 This is an organizational structure diagram of the present invention; Figure 2 This is a flowchart of the WGAN-GP process of the present invention; Figure 3 Flowchart for generating lesion-focused heatmaps; Figure 4 Image data generated by GAN; Figure 5 This is a flowchart of the CNN-Transformer deep learning diagnostic model. Detailed Implementation

[0019] In this example, to accurately classify oral medical images and improve the precision and generalization ability of lesion feature extraction, a classification method for oral medical images based on a WGAN-GP fusion CNN-Transformer hybrid architecture is proposed. WGAN-GP increases the number of labeled medical image samples, allowing the neural network model to be fully trained. Combining the long-range dependency capture capability of Transformer with the local feature extraction advantage of CNN, this method can accurately identify anatomical structures in oral medical images, including teeth, gums, oral mucosa, and jawbones. Considering the rich local details, strong global structural correlations, large contrast variations, and susceptibility to metal artifacts in oral medical images, a classification method adapted for oral medical images from smartphone devices is developed. This improves the generalization ability, accuracy, and robustness of the classification model on oral medical images, providing key technical support for clinical prevention and precision treatment. Specifically, for example... Figure 1 As shown, this oral medicine image classification method includes the following steps: Step 1: Obtain the dimension as Real oral medical images were obtained and normalized preprocessed to obtain a preprocessed oral medical image sequence. ,in, Indicates the first A dental image, This represents the total number of oral medical images. Indicates the height of an oral medical image. Indicates the width of the oral medicine image. This represents the number of channels in an oral medicine image; let... The true category label is denoted as ,when When, it means For an exception category, when When, it means The data is classified as true normal. The oral cavity image dataset underwent standardized preprocessing to unify image size and adapt to the standard input requirements of deep learning models. Then, the image pixel values ​​were normalized and scaled to the [0,1] range to reduce the gradient vanishing problem during training and ensure faster model convergence.

[0020] Step 2: Construct a generative adversarial network-based oral medicine image generation network, including: a heatmap guide, a generator, and a scorer, and then... The images were processed to obtain a forged oral medicine image sequence. Thus obtain Authenticity rating and Authenticity rating Used to construct the gradient penalty loss function And train the oral medicine image generation network to obtain the optimal oral medicine image generation model WGAN-GP, the structure of which is as follows: Figure 2 As shown, it is used to output the first... The optimal forged oral medicine image that preserves the spatial distribution features of teeth, gums, oral mucosa, and jawbone. Thus, the optimal set of forged oral medical images is obtained. ,in, Indicates the first A fake dental image.

[0021] Step 2.1, Heatmap Guide Processing is performed to obtain the first... A heat map ; Step 2.1.1: Utilize pre-trained... Model pair Feature extraction is performed to obtain the first... High-level feature map High-level feature map ; Step 2.1.2, along the channel dimension Perform mean pooling to generate the first... Single-channel characteristic response map Single-channel characteristic response map ; Step 2.1.3: Use bicubic interpolation method to... Perform upsampling to obtain the first A heat map .

[0022] Step 2.2, the generator processes the first... random noise vectors and Processing is performed to obtain the same as Fake oral images with the same probability distribution as the image data .

[0023] Step 2.2.1, along the spatial dimension Perform global average pooling to obtain Global feature representation Global feature representation of heatmap ; Step 2.2.2: Randomly sample the first... noise vectors , Let the dimension of the noise vector be . and After splicing along the first dimension, we get the th A joint input tensor , .

[0024] Step 2.2.3, will The input generator performs stepwise feature amplification to generate the first... A fake oral disease image Fake images of oral diseases The generator structure is as follows: Figure 3 As shown, the generated fake image is... Figure 4 For example, as shown in the figure.

[0025] Step 2.3: The scorer respectively... and Downsampling was performed separately to obtain the true high-level abstract feature map of a single channel. And single-channel fake high-level abstract feature maps Then, they were directly flattened into Authenticity rating and Authenticity rating ; Step 2.4: Construct the gradient penalty loss function using equation (1) : (1) In equation (1), This represents the scorer's expectation of real oral medical images. This indicates the scorer's expectation of a fake oral medicine image. It is a gradient penalty term. This is the penalty coefficient; by introducing a gradient penalty term into the loss function, the gradient norm of the scorer is explicitly constrained to be close to 1 on the interpolation path, thus approximately satisfying the 1-Lipschitz condition. Even when there is a large gap between the real data distribution and the generated data distribution, WGAN-GP can still provide effective gradient information, enabling the generator to learn more diverse data distributions.

[0026] Step 2.5: Train the oral medicine image generation network based on generative adversarial network using the backpropagation algorithm, and calculate the gradient penalty loss function. The algorithm updates all network parameters to be learned in both the scorer and generator, allowing the scorer and generator to optimize alternately until the game converges, thus obtaining the optimal oral medicine image generation model WGAN-GP. The model structure is as follows: Figure 2 As shown; Step 3: Construct a CNN-Transformer dual-stream collaborative deep neural image classification network. The image feature extraction module structure in this example is as follows: Figure 5 As shown, it includes: a primary feature extraction module, a CNN local feature extraction module based on the ResNet50 main structure, a Transformer global feature encoding module, and a classification module, and performs... and Merged training sample set Any k-th training sample Processing is performed, and at the same time, results are obtained. Predicted normal category probability and predicting the probability of anomaly categories ;make The true category label is denoted as ,when When, it means For an exception category, when When, it means This is the normal category.

[0027] Step 3.1: The primary feature extraction module uses the initial convolutional and pooling layers of ResNet50 to... Perform primary feature extraction to obtain the k-th primary feature map. The primary feature extraction module reuses the initial convolutional and pooling layers of ResNet50, and reduces the image size by a factor of 2 through two downsampling steps. It then uses 64 convolutional kernels to expand the number of channels to 64 to obtain the primary feature map. .

[0028] Step 3.2: CNN Local Feature Extraction Module Based on ResNet50 Main Structure Local feature extraction is performed on the spatial distribution characteristics of teeth, gums, oral mucosa, and jawbone to obtain the k-th local feature map. The CNN local feature extraction module consists of four layers. Layer 1 maintains the same image size. The first residual block of layers 2, 3, and 4 is downsampled using a 3×3 convolution kernel, reducing the image size by a factor of 2. The first residual block of layer 1 uses a 1×1 convolution to increase the number of channels by a factor of 4, and the first residual blocks of layers 2, 3, and 4 use a 1×1 convolution to increase the number of channels by a factor of 2, generating local feature maps. .

[0029] Step 3.3: The Transformer global feature encoding module includes: a Transformer input reshaping unit, a Transformer encoding unit, a Transformer output reshaping unit, and a CNN output flattening unit, and performs... After processing, the k-th one-dimensional global association vector is obtained. and the k-th one-dimensional local detail vector .

[0030] Step 3.3.1: The Transformer input reshaping unit will... Reconstructed into two-dimensional CNN local feature sequences ; Step 3.3.2: The Transformer encoding unit utilizes a multi-head attention mechanism and a feedforward neural network to... The process involves establishing feature dependencies between different regions of the teeth, gums, oral mucosa, and jawbone, and obtaining the k-th global feature sequence. Transformer encoding includes two layers of Transformer encoders. , Each layer contains a quad-attention mechanism and a feedforward neural network for capturing... Global relationships between sequence elements; Step 3.3.3, Transformer output reshaping unit pair After processing, the k-th one-dimensional global association vector is obtained. ; Step 3.3.4: Flattening the CNN output Flattening the spatial and channel dimensions yields the result... One-dimensional local detail vectors with the same dimensions .

[0031] Step 3.4: The classification module includes a feature fusion unit and two fully connected layers, and performs... and Processing is performed, and at the same time, results are obtained. Predicted normal category probability and predicting the probability of anomaly categories .

[0032] Step 3.4.1, will and After concatenation along the channel dimension, the k-th fused feature vector based on local spatial features and global relational features is obtained. , used for classification of two fully connected layers; Step 3.4.2, will The input is processed through a nonlinear transformation in the first fully connected layer, and the k-th intermediate feature vector is output. , This represents the number of neurons in the first fully connected layer. Step 3.4.3, will The input is processed in the second fully connected layer to obtain... Original classification score Then, it is input into the Softmax function for processing, and at the same time, it obtains... Predicted normal category probability and predicting the probability of anomaly categories .

[0033] Step 4: Construct the cross-entropy loss function of the deep neural image classification network using equation (2). : (2) In equation (2), express The number of training samples in the dataset.

[0034] Step 5: Train the deep neural image classification network using the Adam optimization algorithm and calculate the cross-entropy loss function. To update all the parameters to be learned in the network, after passing through the cross-entropy loss function The process continues until convergence, thus obtaining the optimal oral medicine image classification model, which is used to perform the binary classification task of oral medicine images.

[0035] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0036] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

[0037] Example: Table 1: Comparison of different algorithms on the validation set

[0038] Note: Vision Transformer Small is the Transformer class used in this invention.

[0039] The comparison results in Table 1 clearly show that the CNN-Transformer (Convolutional Neural Network-Transformer) of this invention significantly outperforms other image classification methods in several key image quality evaluation metrics, including accuracy, F1 score, number of parameters, and number of floating-point operations. This demonstrates that this invention can more effectively improve the extraction of local and global features from images, focusing on the edges (benign) or irregular regions (malignant) of oral ulcers, thereby enhancing the classification ability of oral medical images.

Claims

1. A method for classifying oral medical images based on a WGAN-GP fusion CNN-Transformer hybrid architecture, characterized in that, Includes the following steps: Step 1: Obtain the dimension as Real oral medical images were obtained and normalized preprocessed to obtain a preprocessed oral medical image sequence. ,in, Indicates the first A dental image, This represents the total number of oral medical images. Indicates the height of an oral medical image. Indicates the width of the oral medicine image. This represents the number of channels in an oral medicine image; let... The true category label is denoted as ,when When, it means For an exception category, when When, it means The true normal category; Step 2: Construct a generative adversarial network-based oral medicine image generation network, including: a heatmap guide, a generator, and a scorer, and then... The images were processed to obtain a forged oral medicine image sequence. Thus obtain Authenticity rating and Authenticity rating Used to construct the gradient penalty loss function The oral medicine image generation network was trained to obtain the optimal oral medicine image generation model WGAN-GP, which was used to output the first... The optimal forged oral medicine image that preserves the spatial distribution features of teeth, gums, oral mucosa, and jawbone. Thus, the optimal set of forged oral medical images is obtained. ,in, Indicates the first A fake dental image; Step 3: Construct a CNN-Transformer dual-stream collaborative deep neural image classification network, including: a primary feature extraction module, a CNN local feature extraction module based on the ResNet50 main structure, a Transformer global feature encoding module, and a classification module, and perform... and Merged training sample set Any k-th training sample Processing is performed, and at the same time, results are obtained. Predicted normal category probability and predicting the probability of anomaly categories ;make The true category label is denoted as ,when When, it means For an exception category, when When, it means Normal category; Step 4: Construct the cross-entropy loss function of the deep neural image classification network using equation (2). : (2) In equation (2), express The number of training samples in the dataset; Step 5: Train the deep neural image classification network using the Adam optimization algorithm and calculate the cross-entropy loss function. To update all the parameters to be learned in the network, after passing through the cross-entropy loss function The process continues until convergence, thus obtaining the optimal oral medicine image classification model, which is used to perform the binary classification task of oral medicine images.

2. The oral medical image classification method based on a WGAN-GP fusion CNN-Transformer hybrid architecture according to claim 1, characterized in that, Step 2 includes the following steps: Step 2.1, Heatmap Guide Processing is performed to obtain the first... A heat map ; Step 2.2, the generator processes the first... random noise vectors and Processing is performed to obtain the same as Fake oral images with the same probability distribution as the image data ; Step 2.3: The scorer respectively... and Downsampling was performed separately to obtain the true high-level abstract feature map of a single channel. And single-channel fake high-level abstract feature maps Then, they were directly flattened into Authenticity rating and Authenticity rating ; Step 2.4: Construct the gradient penalty loss function using equation (1) : (1) In equation (1), Indicates the scorer's... Expectations Indicates the scorer's... Expectations It is a gradient penalty term. It is the penalty coefficient; Step 2.5: Train the oral cavity image generation network based on generative adversarial network using the backpropagation algorithm, and calculate the gradient penalty loss function. Update all network parameters to be learned in the scorer and generator, and optimize the scorer and generator alternately until the game converges, thereby obtaining the optimal oral medicine image generation model WGAN-GP.

3. The oral medical image classification method based on a WGAN-GP fusion CNN-Transformer hybrid architecture according to claim 2, characterized in that, Step 2.1 includes the following steps: Step 2.1.1: Utilize pre-trained... Model pair Feature extraction is performed to obtain the first... High-level feature map ; Step 2.1.2, along the channel dimension Perform mean pooling to generate the first... Single-channel characteristic response map ; Step 2.1.3: Use bicubic interpolation method to... Perform upsampling to obtain the first A heat map .

4. The oral medical image classification method based on a WGAN-GP fusion CNN-Transformer hybrid architecture according to claim 3, characterized in that, Step 2.2 includes the following steps: Step 2.2.1, along the spatial dimension Perform global average pooling to obtain Global feature representation ; Step 2.2.2: Randomly sample the first... noise vectors , Let the dimension of the noise vector be . and After splicing along the first dimension, we get the th A joint input tensor ; Step 2.2.3, will The input generator performs stepwise feature amplification to generate the first... A fake oral disease image .

5. The oral medical image classification method based on a WGAN-GP fusion CNN-Transformer hybrid architecture according to claim 1, characterized in that, Step 3 includes the following steps: Step 3.1: The primary feature extraction module uses initial convolutional layers and pooling layers to... Perform primary feature extraction to obtain the k-th primary feature map. ; Step 3.2: CNN Local Feature Extraction Module Based on ResNet50 Main Structure Local feature extraction is performed on the spatial distribution characteristics of teeth, gums, oral mucosa, and jawbone to obtain the k-th local feature map. ; Step 3.3: The Transformer global feature encoding module includes: a Transformer input reshaping unit, a Transformer encoding unit, a Transformer output reshaping unit, and a CNN output flattening unit, and performs... After processing, the k-th one-dimensional global association vector is obtained. and the k-th one-dimensional local detail vector ; Step 3.4: The classification module includes a feature fusion unit and two fully connected layers, and performs... and Processing is performed to obtain Predicted normal category probability and predicting the probability of anomaly categories .

6. The oral medical image classification method based on a WGAN-GP fusion CNN-Transformer hybrid architecture according to claim 5, characterized in that, Step 3.3 includes the following steps: Step 3.3.1: The Transformer input reshaping unit will... Reconstructed into two-dimensional CNN local feature sequences ; Step 3.3.2: The Transformer encoding unit utilizes a multi-head attention mechanism and a feedforward neural network to... The process involves establishing feature dependencies between different regions of the teeth, gums, oral mucosa, and jawbone, and obtaining the k-th global feature sequence. ; Step 3.3.3, Transformer output reshaping unit pair After processing, the k-th one-dimensional global association vector is obtained. ; Step 3.3.4: Flattening the CNN output Flattening the spatial and channel dimensions yields the result... One-dimensional local detail vectors with the same dimensions .

7. The oral medical image classification method based on a WGAN-GP fusion CNN-Transformer hybrid architecture according to claim 6, characterized in that, Step 3.4 includes the following steps: Step 3.4.1, will and After concatenation along the channel dimension, the k-th fused feature vector based on local spatial features and global relational features is obtained. ; Step 3.4.2, will The input is processed through a nonlinear transformation in the first fully connected layer, and the k-th intermediate feature vector is output. , This represents the number of neurons in the first fully connected layer. Step 3.4.3, will The input is processed in the second fully connected layer to obtain... Original classification score Then, it is input into the Softmax function for processing, and at the same time, it obtains... Predicted normal category probability and predicting the probability of anomaly categories .

8. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the oral medical image classification method according to any one of claims 1-7, the processor being configured to execute the program stored in the memory.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when run by a processor, performs the steps of the oral medical image classification method according to any one of claims 1-7.