Image recognition system for oral mucosa lesion

Through the improvement of the multi-scale weighted Otsu edge segmentation and lesion area classification module, the accuracy and robustness of traditional systems in lesion area segmentation and classification are solved, efficient and accurate oral mucosal lesion image recognition is achieved, and the system's adaptability and recognition ability are improved.

CN120259286AActive Publication Date: 2025-07-04CENT SOUTH UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510731129.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Traditional oral mucosal image recognition systems have problems with low accuracy and poor robustness in the precise segmentation and classification of lesion areas. Especially under complex backgrounds and varied lesions, it is difficult to meet the requirements of high accuracy and efficiency, and the sensitivity to local characteristics is insufficient, resulting in inaccurate segmentation and classification results.

Method used

Multi-scale weighted Otsu edge segmentation and local contrast weighted optimization technology are used to combine image pyramid and Sobel edge detection for precise segmentation of lesion areas; the lesion area classification module improves feature extraction and processing capabilities through feature combination, feature interaction, HEMish activation function and variable expansion causal convolution technology.

Benefits of technology

It significantly improves the segmentation accuracy and classification accuracy of oral mucosal lesions images, enhances the robustness and adaptability of the system, can better handle various lesion types, especially the capture of details and the sensitivity of local features, and improves the reliability and practicality of the identification system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259286A_ABST
    Figure CN120259286A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing technology, provides an image recognition system for oral mucosa lesions, and aims to improve the precision and efficiency of oral image processing. The system comprises a data acquisition module, a data preprocessing module, an image segmentation module, a lesion area classification module and a report display and visualization module. Through an image pyramid, Sobel edge detection, Otsu threshold segmentation and a local contrast weighted optimization technology, an image segmentation module realizes high-precision lesion region segmentation; the lesion region classification module is combined with technologies such as feature combination, feature interaction, an HEMish activation function and variable expansion causal convolution to accurately classify oral lesion regions; through a multi-level and multi-dimensional technical means, the recognition precision, robustness and adaptability of the lesion area are remarkably improved, the problems of low precision, instable recognition and the like in a traditional system are solved, and the technical progress in the field of oral mucosa lesion image recognition is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and in particular to an image recognition system for oral mucosal lesions. Background Art

[0002] With the continuous progress of technology, image recognition and deep learning technologies are increasingly widely used in the field of medical image analysis; however, traditional oral mucosal image recognition systems still face many challenges; firstly, existing systems are still difficult to meet the requirements of high accuracy and high efficiency in the precise segmentation and classification of lesion areas, especially under complex backgrounds and variable lesion manifestations; secondly, traditional systems have low sensitivity to local features, especially in details such as edges and textures, which are often ignored, thus affecting the accurate recognition of lesion areas; finally, traditional methods rely on relatively simple image segmentation and classification technologies, ignoring the capture of subtle differences between lesion areas, resulting in inaccurate segmentation and classification results; therefore, there is an urgent need for a more efficient and accurate oral mucosal lesion image recognition system to provide more precise segmentation and classification results of lesion areas and have stronger robustness and adaptability. Summary of the Invention

[0003] The present invention provides an image recognition system for oral mucosal lesions, which adopts a new type of image segmentation and lesion area classification technology to improve the accuracy and robustness of lesion area recognition; specifically, the system combines image pyramid, Sobel edge detection, Otsu threshold segmentation and local contrast weighted optimization technology in the image segmentation module. Through multi-scale processing and edge detection, the lesion area in the oral image is accurately segmented, and the clarity and accuracy of the segmentation result are ensured through local contrast weighted optimization; the lesion area classification module of the system further improves the feature extraction and processing ability through feature combination, feature interaction, HEMish activation function and variable dilation causal convolution technology, so as to achieve precise classification of the lesion area; overall, the present invention ensures the efficient and accurate recognition of oral mucosal lesion images through multi-level and multi-dimensional technical means, and has strong robustness and adaptability.

[0004] The present invention provides an image recognition system for oral mucosal lesions, which includes a data acquisition module, a data preprocessing module, an image segmentation module, a lesion area classification module, and a report display and visualization module; The data acquisition module uses an endoscope, a high-definition camera and a digital oral camera to collect images in the oral cavity as the original oral image data; the original oral image data includes mucosal surface images and lesion area images; The data preprocessing module performs image denoising, image enhancement, image normalization, image cropping and image scaling on the original oral image data to generate preprocessed oral image data; An image segmentation module constructs a multi-scale weighted Otsu edge segmentation model by combining image pyramid, Sobel edge detection, Otsu threshold segmentation, and local contrast weighted optimization techniques. The preprocessed oral image data is processed by the multi-scale weighted Otsu edge segmentation model to generate a multi-scale edge segmentation map. A lesion area classification module establishes a convolutional neural network model. The convolutional neural network model is improved by introducing feature combination, feature interaction, HEMish activation function, and variable dilation causal convolution to construct a multi-scale distorted-variable convolutional neural network model. The multi-scale edge segmentation map is processed by the multi-scale distorted-variable convolutional neural network model to generate lesion area classification information. A report display and visualization module generates an analysis report of oral lesions based on the lesion area classification information and the multi-scale edge segmentation map. The analysis report of oral lesions includes classification results, detailed information of the lesion area, and image annotations. The analysis report of oral lesions is displayed on a visualization interface to view the probability distribution of the lesion area, classification results, and annotations on the image, assisting doctors in further observation, data archiving, and historical record comparison. The multi-scale weighted Otsu edge segmentation model includes a multi-scale image pyramid unit, a Sobel operator edge detection unit, a multi-scale Otsu threshold unit, and a generation unit. The multi-scale distorted-variable convolutional neural network model includes a fully connected layer 1, a Dropout layer, a fully connected layer 2, a residual connection, and an output layer.

[0005] Furthermore, the multi-scale image pyramid unit constructs an image pyramid, decomposes the preprocessed oral image data into 5 levels with different resolutions, and each level contains two scales to obtain multi-scale image data.

[0006] Furthermore, the Sobel operator edge detection unit applies the Sobel operator to perform edge detection on the multi-scale image data, calculates the horizontal gradient and the vertical gradient, and calculates the gradient magnitude according to the horizontal gradient and the vertical gradient to highlight the edge area through the gradient magnitude.

[0007] Furthermore, the multi-scale Otsu threshold unit calculates the gray histogram and the between-class variance of the multi-scale image data, then introduces local contrast weighted optimization to determine the optimal threshold. According to the optimal threshold, the gradient magnitude and the multi-scale image data are binarized to realize the segmentation of the foreground and background in the multi-scale image data, obtaining a multi-scale binary image. The formula used is as follows: ; where represents the index of the local area, represents the contrast of the th local area. Represents the pixel value of a local area , Represents the maximum value of the pixel values of a local area , Represents the minimum value of the pixel values of a local area , Represents the average value of all pixel values in a local area ; ; Among them, Represents the candidate threshold Represents the between-class variance Respectively represent the weights of the foreground and background in the multi-scale image data And Respectively represent the means of the foreground and background in the multi-scale image data

[0008] Furthermore, the generation unit performs edge thinning, noise point removal, and gap filling on the multi-scale binary image to generate a multi-scale edge segmentation map

[0009] Furthermore, the lesion area classification module, in the process of generating lesion area classification information, specifically includes the following steps Step D1: Feature combination: Perform product combination and pairwise product summation on the multi-scale edge segmentation map to generate combined feature data Step D2: Feature interaction: Perform non-linear transformation on the combined feature data through the ReLU activation function to capture the high-order interaction relationships of the combined feature data, generate rich high-order feature representations; enhance the model's ability to express complex relationships between features Step D3: Feature transformation: Combine the smoothness of Mish, the Gaussian function, the hyperbolic tangent function, and adjustable hyperparameters to construct the HEMish activation function; perform linear transformation on the rich high-order feature representations, and then perform non-linear transformation through the HEMish activation function to generate deep feature representations; further enhance the non-linear expression ability of the features. The formula used is as follows ; Among them, Represents the linear transformation data Represents the HEMish activation function Represents the hyperbolic tangent function Represents an adjustable parameter that controls the influence intensity of the Gaussian function Represents an adjustable parameter that controls the attenuation rate Represents the Gaussian attenuation function Represents the non-linear transformation for smoothing ​represents the amplitude factor of the sine function, represents the frequency factor of the sine function; represents a periodic function, represents an exponential function, endowing the ability of non - linear expression with periodic changes; Step D4: Feature mapping 1: Process the deep - layer feature representation through the fully - connected layer 1 to generate the fully - connected layer 1 feature data; Learn the deep - layer feature representation through the fully - connected layer 1 and output the dilation factor control signal; Step D5: Causal convolution: Further capture the long - term dependency relationship between the fully - connected layer 1 feature data through the variable - dilation causal convolution to generate the variable - causal feature data; Step D6: Batch normalization processing: Perform batch normalization processing on the variable - causal feature data to reduce the internal covariance change during training and generate the batch - normalized feature data; Step D7: Dropout processing: Randomly discard a part of neurons of the batch - normalized feature data through the Dropout layer to generate the Dropout feature data; Step D8: Feature mapping 2: Further refine the feature representation of the Dropout feature data through the fully - connected layer 2 to generate the fully - connected layer 2 feature data; Step D9: Generate classification results: Input the fully - connected layer 2 feature data into the output layer to generate the classification information of the lesion area; Step D10: Residual connection: Add a residual connection between the fully - connected layer 1 and the fully - connected layer 2 to solve the problem of gradient vanishing.

[0010] Further, step D5 specifically includes the following steps: Step D51: Introduce a variable - dilation factor to improve the dilation causal convolution, construct the variable - dilation causal convolution, initialize the variable - dilation factor, and set the range of the variable - dilation factor; Step D52: According to the range of the variable - dilation factor and the dilation factor control signal, use the dilation causal convolution to change the receptive field of the convolution kernel, capture the long - term dependency relationship between the fully - connected layer 1 feature data, and generate the variable - causal feature data.

[0011] The beneficial effects achieved by the present invention adopting the above - mentioned scheme are as follows: The present invention effectively achieves the precise segmentation of oral mucosal lesion images by adopting multi-scale weighted Otsu edge segmentation and local contrast weighted optimization techniques. This image segmentation technology can overcome the problem of inaccurate segmentation in traditional systems under complex backgrounds and various lesion types. Through multi-scale image pyramids and Sobel edge detection techniques, it ensures the clear presentation of the lesion area at different resolutions and improves the segmentation accuracy. The introduction of this technical means significantly solves the performance deficiency problem of the oral mucosal image recognition system when facing high-noise or low-contrast images, thus ensuring that the data after image preprocessing is more suitable for subsequent lesion classification and improving the recognition accuracy of the entire system. In addition, the lesion area classification module of this system improves the accuracy of lesion area classification by introducing feature combination, feature interaction, and HEMish activation function techniques. This module enhances the system's ability to capture complex features and can effectively handle the recognition of subtle differences in the lesion area. Through the improvement of these technologies, the system can better process various lesion types, especially the capture of details and the sensitivity to local features have been significantly improved. Compared with traditional systems, the present invention greatly enhances the adaptability of the system to different oral images, enabling it to accurately classify and identify various oral mucosal lesions, and further improving the reliability and practicality of the oral mucosal image recognition system. In summary, the present invention significantly improves the overall performance of the oral mucosal lesion image recognition system by combining advanced image segmentation and classification technologies. The innovative improvement of the multi-scale weighted Otsu edge segmentation technology and the lesion area classification module not only effectively improve the segmentation accuracy of the lesion area but also enhance the reliability of the classification results, providing a more accurate and efficient technical means for the automatic analysis of oral medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a schematic diagram of the modules of an image recognition system for oral mucosal lesions proposed by the present invention. Figure 2 It is the model architecture diagram of the multi-scale twisted-variable convolutional neural network model proposed in Embodiment VII. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0014] Embodiment 1, according to Figure 1, the present invention provides an image recognition system for oral mucosal lesions, which includes a data acquisition module, a data preprocessing module, an image segmentation module, a lesion area classification module, and a report display and visualization module; The data acquisition module uses an endoscope, a high-definition camera, and a digital oral camera to collect images in the oral cavity as the original oral image data; the original oral image data includes mucosal surface images and lesion area images; The data preprocessing module performs image denoising, image enhancement, image normalization, image cropping, and image scaling on the original oral image data to generate preprocessed oral image data; The image segmentation module constructs a multi-scale weighted Otsu edge segmentation model by combining image pyramid, Sobel edge detection, Otsu threshold segmentation, and local contrast weighted optimization techniques, and processes the preprocessed oral image data through the multi-scale weighted Otsu edge segmentation model to generate a multi-scale edge segmentation map; The lesion area classification module establishes a convolutional neural network model, improves the convolutional neural network model by introducing feature combination, feature interaction, HEMish activation function, and variable dilation causal convolution, and constructs a multi-scale distortion-variable convolutional neural network model; processes the multi-scale edge segmentation map through the multi-scale distortion-variable convolutional neural network model to generate lesion area classification information; The report display and visualization module generates an analysis report of oral lesions based on the lesion area classification information and the multi-scale edge segmentation map; the analysis report of oral lesions includes classification results, detailed information of the lesion area, and image annotations; displays the analysis report of oral lesions in a visualization interface to view the probability distribution of the lesion area, classification results, and annotations on the image, assisting doctors in further observation, data archiving, and historical record comparison; The multi-scale weighted Otsu edge segmentation model includes a multi-scale image pyramid unit, a Sobel operator edge detection unit, a multi-scale Otsu threshold unit, and a generation unit; The multi-scale distortion-variable convolutional neural network model includes a fully connected layer 1, a Dropout layer, a fully connected layer 2, a residual connection, and an output layer.

[0015] Embodiment 2, based on Embodiment 1, in this embodiment, the multi-scale image pyramid unit constructs an image pyramid, decomposes the preprocessed oral image data into 5 levels of different resolutions, and each level contains two scales to obtain multi-scale image data.

[0016] Embodiment 3. This embodiment is based on Embodiment 2. In this embodiment, the Sobel operator edge detection unit applies the Sobel operator to perform edge detection on the multi-scale image data, calculates the horizontal gradient and the vertical gradient, calculates the gradient magnitude based on the horizontal gradient and the vertical gradient, and highlights the edge region through the gradient magnitude.

[0017] Embodiment 4. This embodiment is based on Embodiment 3. In this embodiment, the multi-scale Otsu threshold unit calculates the gray histogram and the between-class variance of the multi-scale image data, then introduces local contrast weighting optimization to determine the optimal threshold, and binarizes the gradient magnitude and the multi-scale image data according to the optimal threshold to achieve the segmentation of the foreground and the background in the multi-scale image data, obtaining a multi-scale binary image. The formula used is as follows: ; where, represents the index of the local region, represents the contrast of the th local region, represents the pixel value of the local region , represents the maximum value of the pixel values of the local region , represents the minimum value of the pixel values of the local region , represents the average value of all the pixel values in the local region ; ; where, represents the candidate threshold, represents the between-class variance, respectively represent the weights of the foreground and the background in the multi-scale image data, and respectively represent the means of the foreground and the background in the multi-scale image data.

[0018] Embodiment 5. This embodiment is based on Embodiment 3. In this embodiment, the multi-scale Otsu threshold unit calculates the gray histogram and the between-class variance of the multi-scale image data, determines the optimal threshold, and binarizes the gradient magnitude and the multi-scale image data according to the optimal threshold to achieve the segmentation of the foreground and the background in the multi-scale image data, obtaining a multi-scale binary image.

[0019] Embodiment 6. This embodiment is based on Embodiment 4. In this embodiment, the generation unit refines the edges of the multi-scale binary image, removes noise points and fills gaps to generate a multi-scale edge segmentation map.

[0020] Embodiment 7. According to Figure 2, this embodiment is based on Embodiment Six. In this embodiment, the process of the lesion area classification module generating lesion area classification information specifically includes the following steps: Step D1: Feature combination: Perform product combination and pairwise product summation on the multi-scale edge segmentation map to generate combined feature data; Step D2: Feature interaction: Perform non-linear transformation on the combined feature data through the ReLU activation function to capture the high-order interaction relationships of the combined feature data, generating rich high-order feature representations; enhancing the model's ability to express complex relationships between features; Step D3: Feature transformation: Combine the smoothness of Mish, Gaussian function, hyperbolic tangent function, and adjustable hyperparameters to construct the HEMish activation function; perform linear transformation on the rich high-order feature representations, and then perform non-linear transformation through the HEMish activation function to generate deep feature representations; further enhancing the non-linear expression ability of features. The formula used is as follows: ; Among them, represents the linearly transformed data, represents the HEMish activation function, represents the hyperbolic tangent function; represents an adjustable parameter that controls the influence intensity of the Gaussian function; represents an adjustable parameter that controls the attenuation rate; represents the Gaussian attenuation function; represents the smooth non-linear transformation; represents the amplitude factor of the sine function, represents the frequency factor of the sine function; represents the periodic function, represents the exponential function, endowing the non-linear expression ability of periodic changes; Step D4: Feature mapping 1: Process the deep feature representations through the fully connected layer 1 to generate fully connected layer 1 feature data; learn the deep feature representations through the fully connected layer 1 and output the expansion factor control signal; Step D5: Causal convolution: Further capture the long-term dependence relationships between the fully connected layer 1 feature data through variable dilation causal convolution to generate variable causal feature data; Step D6: Batch normalization processing: Perform batch normalization processing on the variable causal feature data to reduce the internal covariance change during training and generate batch normalization feature data; Step D7: Dropout processing: Randomly discard a part of the neurons in the batch normalization feature data through the Dropout layer to generate Dropout feature data; Step D8: Feature Mapping 2: Further refine the feature representation of the Dropout feature data through the fully connected layer 2 to generate the fully connected layer 2 feature data; Step D9: Generate Classification Results: Input the fully connected layer 2 feature data into the output layer to generate the classification information of the lesion area; Step D10: Residual Connection: Add a residual connection between the fully connected layer 1 and the fully connected layer 2 to solve the problem of gradient disappearance.

[0021] Embodiment 8. This embodiment is based on Embodiment 6. In this embodiment, the process of the lesion area classification module generating the lesion area classification information specifically includes the following steps: Step D1: Feature Combination: Perform product combination and pairwise product summation on the multi-scale edge segmentation maps to generate combined feature data; Step D2: Feature Interaction: Perform a non-linear transformation on the combined feature data through the ReLU activation function to capture the high-order interaction relationships of the combined feature data and generate rich high-order feature representations; enhance the model's ability to express complex relationships between features; Step D3: Feature Transformation: Perform a linear transformation on the rich high-order feature representations and then perform a non-linear transformation through the ReLU activation function to generate deep feature representations; further improve the non-linear expression ability of the features; Step D4: Feature Mapping 1: Process the deep feature representations through the fully connected layer 1 to generate the fully connected layer 1 feature data; learn the deep feature representations through the fully connected layer 1 and output the dilation factor control signal; Step D5: Causal Convolution: Further capture the long-term dependence relationships between the fully connected layer 1 feature data through dilated causal convolution to generate variable causal feature data; Step D6: Batch Normalization Processing: Perform batch normalization processing on the variable causal feature data to reduce the internal covariance change during training and generate batch-normalized feature data; Step D7: Dropout Processing: Randomly discard a part of the neurons of the batch-normalized feature data through the Dropout layer to generate Dropout feature data; Step D8: Feature Mapping 2: Further refine the feature representation of the Dropout feature data through the fully connected layer 2 to generate the fully connected layer 2 feature data; Step D9: Generate Classification Results: Input the fully connected layer 2 feature data into the output layer to generate the classification information of the lesion area; Step D10: Residual Connection: Add a residual connection between the fully connected layer 1 and the fully connected layer 2 to solve the problem of gradient disappearance.

[0022] Embodiment 9. This embodiment is based on Embodiment 7. In this embodiment, Step D5 specifically includes the following steps: Step D51: Introduce a variable dilation factor to improve dilated causal convolution, construct variable dilation causal convolution, initialize the variable dilation factor, and set the range of the variable dilation factor; Step D52: According to the range of the variable dilation factor and the dilation factor control signal, use dilated causal convolution to change the receptive field of the convolution kernel, capture the long-term dependencies between the feature data of the fully connected layer 1, and generate variable causal feature data.

[0023] The above describes the present invention and its embodiments. Such description is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In summary, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural forms and embodiments to this technical solution without creative efforts, they should all fall within the protection scope of the present invention.

Claims

1. An image recognition system for oral mucosal lesions, comprising a data acquisition module and a data preprocessing module; the data acquisition module acquires original oral image data; the data preprocessing module preprocesses the original oral image data to generate preprocessed oral image data; characterized in that: The system further includes an image segmentation module and a lesion area classification module; The image segmentation module constructs a multi-scale weighted Otsu edge segmentation model by combining image pyramid, Sobel edge detection, Otsu threshold segmentation, and local contrast weighted optimization techniques. The pre-processed oral image data is processed by the multi-scale weighted Otsu edge segmentation model to generate a multi-scale edge segmentation map; The lesion area classification module establishes a convolutional neural network model. By introducing feature combination, feature interaction, HEMish activation function, and variable dilation causal convolution, the convolutional neural network model is improved to construct a multi-scale distortion-variable convolutional neural network model. The multi-scale edge segmentation map is processed by the multi-scale distortion-variable convolutional neural network model to generate lesion area classification information.

2. The image recognition system for oral mucosal lesions according to claim 1, characterized in that: The multi-scale weighted Otsu edge segmentation model includes a multi-scale image pyramid unit, a Sobel operator edge detection unit, a multi-scale Otsu threshold unit, and a generation unit.

3. An image recognition system for oral mucosal lesions according to claim 1, characterized in that: The multi-scale distortion-variable convolutional neural network model includes a fully connected layer 1, a Dropout layer, a fully connected layer 2, a residual connection, and an output layer.

4. An image recognition system for oral mucosal lesions according to claim 2, characterized in that: The multi-scale image pyramid unit constructs an image pyramid, decomposes the pre-processed oral image data into 5 levels of different resolutions, and each level contains two scales to obtain multi-scale image data.

5. An image recognition system for oral mucosal lesions according to claim 4, characterized in that: The Sobel operator edge detection unit applies the Sobel operator to perform edge detection on the multi-scale image data, calculates the horizontal gradient and the vertical gradient, calculates the gradient magnitude according to the horizontal gradient and the vertical gradient, and highlights the edge area through the gradient magnitude.

6. The image recognition system for oral mucosal lesions according to claim 5, wherein: The multi-scale Otsu threshold unit calculates the gray histogram and the between-class variance of the multi-scale image data, then introduces local contrast weighted optimization to determine the optimal threshold. According to the optimal threshold, the gradient magnitude and the multi-scale image data are binarized to realize the segmentation of the foreground and background in the multi-scale image data, and a multi-scale binary image is obtained.

7. The image recognition system for oral mucosal lesions according to claim 6, characterized in that: The generation unit refines the edges of the multi-scale binary image, removes noise points, and fills gaps to generate a multi-scale edge segmentation map.

8. The image recognition system for oral mucosal lesions according to claim 3, wherein: The process of the lesion area classification module generating lesion area classification information specifically includes the following steps: Step D1: Feature combination: The multi-scale edge segmentation map is subjected to product combination and pairwise product summation to generate combined feature data; Step D2: Feature interaction: The combined feature data is non-linearly transformed through the ReLU activation function to capture the high-order interaction relationship of the combined feature data and generate a rich high-order feature representation; Step D3: Feature transformation: Combining the smoothness of Mish, the Gaussian function, the hyperbolic tangent function, and adjustable hyperparameters, a HEMish activation function is constructed; The rich high-order feature representation is linearly transformed and then non-linearly transformed through the HEMish activation function to generate a deep feature representation; Step D4: Feature Mapping 1: Process the deep feature representation through the fully connected layer 1 to generate the fully connected layer 1 feature dataset; learn the deep feature representation through the fully connected layer 1 and output the dilation factor control signal; Step D5: Causal Convolution: Capture the long-term dependencies in the fully connected layer 1 feature dataset through the variable dilation causal convolution to generate the variable causal feature data; Step D6: Batch Normalization Processing: Perform batch normalization processing on the variable causal feature data to reduce the internal covariance shift during training and generate the batch normalized feature data; Step D7: Dropout Processing: Pass the batch normalized feature data through the Dropout layer to randomly discard a part of the neurons and generate the Dropout feature data; Step D8: Feature Mapping 2: Refine the feature representation of the Dropout feature data through the fully connected layer 2 to generate the fully connected layer 2 feature data; Step D9: Generate Classification Results: Input the fully connected layer 2 feature data into the output layer to generate the classification information of the lesion area; Step D10: Residual Connection: Add a residual connection between the fully connected layer 1 and the fully connected layer 2.

9. An image recognition system for oral mucosal lesions according to claim 8, wherein: The specific steps of Step D5 are as follows: Step D51: Introduce a variable dilation factor to improve the dilated causal convolution, construct the variable dilation causal convolution, initialize the variable dilation factor, and set the variable dilation factor range; Step D52: According to the variable dilation factor range and the dilation factor control signal, use the dilated causal convolution to change the receptive field of the convolution kernel, capture the long-term dependencies between the fully connected layer 1 feature data, and generate the variable causal feature data.

Citation Information

Patent Citations

  • SD-OCT image retinopathy detection system based on category discrimination and location

    CN109493954A

  • Ultrasonic image breast tumor classification method based on feature fusion and attention mechanism

    CN117746119A

  • Deep convolutional neural network with self-transfer learning

    US20190122360A1

  • Method for detecting and classifying lesion area in clinical image

    WO2022037642A1