A method for detecting oral diseases based on artificial intelligence large model

Through the oral disease detection method based on artificial intelligence large-scale models, combined with feature extraction and YOLOv8 network, the problem that the existing technology cannot effectively identify oral disease types and grading is solved, and higher detection and classification accuracy and assisted decision-making effect of lesion description is achieved.

CN119273632BActive Publication Date: 2025-05-16SHINVA MEDICAL INSTR CO LTD +2

Patent Information

Application Number
CN202411292542.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-05-16
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

The existing intelligent diagnostic methods cannot effectively identify the types and grading of oral diseases, especially because the types of oral diseases are complex and the lesions of the same types and different degrees are difficult to distinguish.

Method used

Using an oral disease detection method based on artificial intelligence large model, image and text features are extracted through feature extraction block BERT and multi-scale feature extraction block MFEM, combined with YOLOv8 network and dynamic tag allocation strategy, training and verification are carried out to improve the accuracy of disease detection and classification.

Benefits of technology

It effectively improves the detection and classification accuracy of various diseases in the oral cavity, and provides corresponding lesion area descriptions and suggestions to assist clinicians in making decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119273632B_ABST
    Figure CN119273632B_ABST
Patent Text Reader

Abstract

The present invention discloses an oral disease detection method based on an artificial intelligence large model, and relates to the field of oral medicine technology. The method comprises obtaining a tooth lesion image and a corresponding description of the internal state of the oral cavity, using a feature extraction block BERT and a multi-scale feature extraction block MFEM to respectively extract the image feature information of the tooth lesion image and the text feature information of the internal state of the oral cavity, identifying and excluding false negative samples therein, introducing a total loss function and a dynamic label allocation strategy in a YOLOv8 network, obtaining an improved YOLOv8 network, dividing a data set into a training set and a test set, sequentially inputting the improved YOLOv8 network for verification, obtaining a verified YOLOv8 network, inputting the tooth lesion image and the corresponding text information into the verified YOLOv8 network, and obtaining lesion area type detection and text description. The present invention can effectively improve the detection and classification of multiple diseases in the oral cavity, and provide corresponding lesion area descriptions and suggestions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of oral medicine technology, and in particular to an oral disease detection method based on an artificial intelligence large model. Background Art

[0002] As living standards improve, people have higher requirements for oral health and gradually realize the importance of oral health. In fact, oral diseases not only cause local infections, tissue defects, etc., but also damage overall health by affecting systemic diseases, nutritional intake, etc., and hinder physiological functions and facial beauty. If not treated in time, tooth loss will affect chewing, pronunciation, and appearance, seriously reducing the quality of life; in addition, malignant diseases (such as oral cancer, etc.) may further develop if not diagnosed and treated in time, causing irreparable damage and seriously endangering overall health; infectious diseases may spread throughout the body if not controlled quickly, interfering with the treatment of digestive, endocrine, cardiovascular and other diseases. Therefore, early detection and diagnosis of oral diseases is of great significance for maintaining oral health.

[0003] However, there is still a large gap in oral medical resources; there are many types of oral diseases (specific classifications include maxillofacial surgical diseases, endodontic diseases, periodontal diseases, oral mucosal diseases, malocclusion, tooth defects, etc.), many of which have relatively hidden symptoms, and diagnosis requires doctors to have a high level of professionalism; in addition, since the severity of oral diseases (grade I, II, III, etc.) is difficult to distinguish, patients themselves often cannot discover the severity of oral diseases in the first place, and are prone to repeated referrals or failure to seek medical treatment in time; in addition, since the professional level of dentists in different regions varies, young and grassroots doctors are prone to miss or misdiagnose some diseases due to lack of experience, which leads to failure to diagnose and effectively treat them in time, and delays the best diagnosis and treatment time for the disease. In the era of artificial intelligence, the medical industry is expected to use new tools for intelligent diagnosis or auxiliary treatment to provide strong support for doctors to make more accurate decisions. At present, existing deep learning models such as convolutional neural networks, generative adversarial networks, recursive neural networks, autoencoders and deep belief networks can realize the extraction of image features, automatic classification and recognition of medical images, and qualitative and quantitative analysis of lesions. However, in the detection of oral diseases, due to the complexity of oral diseases (such as ulcers, caries, periodontitis, etc.), and the difficulty in distinguishing oral diseases of the same type and different degrees (such as recurrent oral ulcers and oral squamous cell carcinoma lesions have similar colors and morphology, and need to use the area around the lesion tissue for auxiliary diagnosis), the existing intelligent diagnosis methods cannot effectively identify the type and grade of oral diseases.

[0004] Therefore, it is an urgent problem for those skilled in the art to propose an oral disease detection method based on a large artificial intelligence model to solve the problems existing in the prior art. Summary of the invention

[0005] In view of this, the present invention provides an oral disease detection method based on an artificial intelligence large model, which can effectively improve the detection and classification of various diseases in the oral cavity and provide corresponding lesion area descriptions and suggestions.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] An oral disease detection method based on an artificial intelligence large model comprises the following steps:

[0008] Data acquisition steps: obtaining images of dental lesions and corresponding descriptions of the internal state of the oral cavity;

[0009] Feature extraction step: The feature extraction block BERT and the multi-scale feature extraction block MFEM are used to extract the image feature information of the dental lesion image and the text feature information of the internal state of the oral cavity, respectively, to obtain a label data set;

[0010] Error elimination step: identify and exclude false negative samples in the labeled data to obtain the preprocessed labeled data set;

[0011] Network improvement steps: Introduce the total loss function and dynamic label allocation strategy into the YOLOv8 network to obtain an improved YOLOv8 network;

[0012] Training and verification steps: The preprocessed labeled data set is divided into a training set and a test set in a ratio of 8:2. The training set is input into the improved YOLOv8 network for training to obtain a trained YOLOv8 network. The test set is input into the trained YOLOv8 network for verification to obtain a verified YOLOv8 network.

[0013] Disease monitoring steps: Input the dental lesion image and the corresponding text information into the verified YOLOv8 network to obtain the lesion area type detection and text description.

[0014] In the above method, optionally, in the feature extraction step, the feature extraction block BERT uses a self-attention mechanism to calculate the influence of each word in the sentence on other words, and uses a masked language model MLM to pre-train the feature extraction block BERT.

[0015] In the above method, optionally, in the feature extraction step, the multi-scale feature extraction block MFEM includes: setting the number of input image channels to four channels, and inputting the four-channel output image into a 1x1 convolution layer, a 3x3 convolution layer, a 5x5 convolution layer, and a maximum pooling layer to obtain four corresponding tensors, concatenating the four tensors to form a new matrix tensor, fusing them through CotNet, generating a new feature information matrix and obtaining the corresponding feature vector X, and then inputting the feature vector X into CotNet to obtain the output result.

[0016] In the above method, optionally, CotNet includes: converting the feature vector X into Queries: Q = XW Q ,Keys:K=XW K ,Values:V=XW V , perform k×k groups of convolutions on all adjacent keys in the k×k spatial grid to obtain the local context information K between adjacent keys 1 , the local context information of the image is concatenated with Q to obtain the attention matrix D, and the attention matrix D is multiplied by V to obtain the total context information K 2 , and the local context information K 1 Add to get the output result.

[0017] In the above method, optionally, in the error elimination step, the image feature f generated by the cosine similarity calculation model is i is and text feature f T The similarity between them is expressed as follows:

[0018]

[0019] If Sim(f i ,f T ) is higher than a certain threshold θ, and the sample is marked as a negative sample in the label dataset, it is marked as a false negative sample.

[0020] In the above method, optionally, in the network improvement step, the total loss function expression is:

[0021] L total =θ cls L cls +θ loc L loc +θ scale L scale ,

[0022] Among them, θ cls ,θ loc ,θ scale is the weight hyperparameter, L cls is the classification loss, L loc is the positioning loss, Lscale is the scale loss.

[0023] It can be seen from the above technical solutions that, compared with the prior art, the present invention provides an oral disease detection method based on an artificial intelligence large model, which has the following beneficial effects: 1) The present invention effectively improves the diagnostic performance of the system by integrating text features and the assistance of prior knowledge, identifies and classifies lesion types, and provides text descriptions of corresponding lesions to assist clinicians in making decisions; 2) The present invention proposes a multi-scale feature extraction module MFEM, which effectively enhances the information of oral disease lesion areas and further improves the algorithm's diagnostic performance; 3) The present invention effectively improves the detection and classification of various diseases in the oral cavity, and provides corresponding lesion area descriptions and suggestions. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0025] Figure 1 This is a flow chart of an oral disease detection method based on an artificial intelligence large model disclosed in the present invention;

[0026] Figure 2 This is a schematic diagram of an oral disease detection method based on an artificial intelligence large model disclosed in the present invention;

[0027] Figure 3 It is the principle diagram of the multi-scale feature extraction module MFEM disclosed in the present invention;

[0028] Figure 4 This is the CotNet principle diagram disclosed by the present invention. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "comprise one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0031] Reference Figure 1 and Figure 2 As shown, the present invention discloses an oral disease detection method based on an artificial intelligence large model, comprising the following steps:

[0032] Data acquisition steps: obtaining images of dental lesions and corresponding descriptions of the internal state of the oral cavity;

[0033] Feature extraction step: The feature extraction block BERT and the multi-scale feature extraction block MFEM are used to extract the image feature information of the dental lesion image and the text feature information of the internal state of the oral cavity, respectively, to obtain a label data set;

[0034] Error elimination step: identify and exclude false negative samples in the labeled data to obtain the preprocessed labeled data set;

[0035] Network improvement steps: Introduce the total loss function and dynamic label allocation strategy into the YOLOv8 network to obtain an improved YOLOv8 network;

[0036] Training and verification steps: The preprocessed labeled data set is divided into a training set and a test set in a ratio of 8:2. The training set is input into the improved YOLOv8 network for training to obtain a trained YOLOv8 network. The test set is input into the trained YOLOv8 network for verification to obtain a verified YOLOv8 network.

[0037] Disease monitoring steps: Input the dental lesion image and the corresponding text information into the verified YOLOv8 network to obtain the lesion area type detection and text description.

[0038] Furthermore, in the data acquisition step, the dental lesion images are oral scan data of patients, and the number is greater than 2,000.

[0039] Furthermore, in the feature extraction step, the feature extraction block BERT uses a self-attention mechanism to calculate the influence of each word in the sentence on other words, and the masked language model MLM is used to pre-train the feature extraction block BERT.

[0040] Specifically, BERT is a pre-trained language model based on the Transformer architecture proposed by Google. It performs well in natural language processing (NLP) tasks and can capture the contextual information of words, thereby achieving higher comprehension capabilities. The core idea of ​​BERT is to understand the context of sentences in a bidirectional manner (i.e., from left to right and from right to left at the same time), while traditional language models such as GPT only use unidirectional context. This bidirectional context understanding makes BERT more accurate in capturing the relationship and context between words.

[0041] Furthermore, the self-attention mechanism includes a representation of a given input sequence X = [x1, x2, ..., x n ], the representation of each word is calculated by the following formula:

[0042]

[0043] Among them, Q (Query), K (Key), and V (Value) are representations obtained through linear transformation. k is the dimension of the Key. In BERT, the self-attention mechanism helps the model capture the dependencies between all words in a sentence.

[0044] Furthermore, during the pre-training process, some words in the input sequence will be randomly masked. Assume that the masked words are w i , the goal of the model is to predict the original vocabulary of these words. The goal of the MLM task is to minimize the following loss function:

[0045] L MLM =-∑ i∈M log P(w i |X),

[0046] Where M is the set of masked words and X is the input sequence after masking.

[0047] Further, see Figure 3 As shown in the figure, in the feature extraction step, the multi-scale feature extraction block MFEM includes: setting the number of input image channels to four channels, and inputting the four-channel output image into the 1x1 convolution layer, the 3x3 convolution layer, the 5x5 convolution layer and the maximum pooling layer to obtain four corresponding tensors, concatenating the four tensors to form a new matrix tensor, fusing them through CotNet, generating a new feature information matrix and obtaining the corresponding feature vector X, and then inputting the feature vector X into CotNet to obtain the output result.

[0048] Specifically, in most CNN architectures, researchers either stack convolutional layers of the same size as filters, or add a pooling layer, or use a dual-channel parallel method to extract and fuse features. However, these methods only extract shallow features or only extract deep features of one scale when extracting features. The feature information extracted in this way is not comprehensive and cannot achieve the best recognition effect. Therefore, this patent proposes a multi-scale feature extraction module Multi-scale Feature Extraction Module (MFEM), which extracts features of different scales in one layer of operation, so as to obtain the deep features of the image to improve the recognition accuracy. The proposed feature extraction module uses four scales of convolution kernels to extract features of different scales.

[0049] Furthermore, the number of input image channels is increased to four channels to increase the expressive power of spectral image information. Different convolution sizes provide different receptive fields, which can be used for feature extraction at different levels. The pooling operation itself has the function of extracting features, and because there are no parameters, it will not produce overfitting. Each branch of the 1x1 convolution layer, 3x3 convolution layer, 5x5 convolution layer and maximum pooling layer contains some convolution layers and batch normalization layers, as well as ReLU activation functions. Convolution layers with different convolution kernel sizes extract different weight information to generate new effective features. The generated matrix with a larger size can be first reduced in dimension, and the visual information can be aggregated at different sizes to facilitate the extraction of features from different scales.

[0050] Further, see Figure 4 As shown, CotNet includes: Converting the feature vector X into Queries: Q = XW Q ,Keys:K=XW K ,Values:V=XW V , perform k×k groups of convolutions on all adjacent keys in the k×k spatial grid to obtain the local context information K between adjacent keys 1 , the local context information of the image is concatenated with Q to obtain the attention matrix D, and the attention matrix D is multiplied by V to obtain the total context information K 2 , and the local context information K 1 Add to get the output result.

[0051] Furthermore, W is the weight matrix, K 1 ∈R H×W×C , which can also be used as a static context representation of the input X. The concatenation process obtains the attention matrix D through two consecutive 1×1 convolutions with and without ReLU activation functions, respectively. The expression is as follows:

[0052] D=[K1 ,Q]W θ W δ ,

[0053] Among them, D represents the attention matrix, K 1 represents the local context information of the image, Q represents Queries, and W θ and W δ Represents two convolutions.

[0054] K 2 It can also be used as a dynamic context for input, the expression is:

[0055]

[0056] Where V represents value, and the expression is:

[0057] Y=K 1 +K 2 .

[0058] Different convolutional layers are combined in parallel, and the result matrices processed by different convolutional layers are spliced ​​together in the depth dimension to form a deeper matrix. This method can effectively process the input image and extract richer and more diverse features. It can efficiently expand the depth and width of the network, improve the accuracy of the deep learning network, and prevent overfitting.

[0059] Furthermore, in the error elimination step, the image feature f generated by the cosine similarity calculation model is i is and text feature f T The similarity between them is expressed as follows:

[0060]

[0061] If Sim(f i ,f T ) is higher than a certain threshold θ, and the sample is marked as a negative sample in the label dataset, it is marked as a false negative sample.

[0062] Specifically, false negative samples refer to data that are misclassified as negative samples (i.e., the model thinks they are irrelevant or mismatched), but in fact they should be positive samples (i.e., relevant or matched). The core idea of ​​False Negative Exclusion (FNE) is to detect and eliminate these false negative samples through certain strategies to ensure that the model only learns on true negative samples.

[0063] Furthermore, if the similarity Sim(f i ,f T ) meets the following conditions, it is identified as a false negative sample:

[0064] IF Sim(f i ,f T )>θ,then it is a False Negative,

[0065] Here, θ is a pre-set hyperparameter determined experimentally to balance the precision and recall of false negative detection.

[0066] Furthermore, in the network improvement step, the total loss function expression is:

[0067] L total =θ cls L cls +θ loc L loc +θ scale L scale ,

[0068] Among them, θ cls ,θ loc ,θ scale is the weight hyperparameter, L cls is the classification loss, L loc is the positioning loss, L scale is the scale loss.

[0069] Furthermore, the total loss function includes classification loss, localization loss and scale loss.

[0070] Specifically, classification loss: measures the gap between the predicted category and the true category, usually calculated using cross entropy loss or focal loss, as follows:

[0071]

[0072] Among them, y i is the true category, y z is the predicted category.

[0073] Positioning loss: measures the gap between the predicted box and the real box, usually calculated using IoU or GIoU, and the expression is as follows:

[0074] L loc =1-IOU(A,B),

[0075] Among them, A is the predicted box and B is the real box.

[0076] Scale loss: used to optimize the scale and aspect ratio of the predicted box to better fit the target.

[0077] Furthermore, in the network improvement step, YOLOv8 introduced a dynamic label assignment strategy, which allows the model to more flexibly select positive and negative samples during training. This strategy can better adapt to target detection tasks in complex scenarios by adjusting the selection rules of positive and negative samples.

[0078] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0079] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An oral disease detection method based on an artificial intelligence large model, characterized in that: The following steps are involved: Data acquisition steps: obtaining images of dental lesions and corresponding descriptions of the internal state of the oral cavity; Feature extraction steps: The feature extraction block BERT and the multi-scale feature extraction block MFEM are used to extract the image feature information of the dental lesion image and the text feature information of the internal state of the oral cavity, respectively, to obtain a labeled data set; Error elimination step: identify and exclude false negative samples in the labeled data to obtain the preprocessed labeled data set; Network improvement steps: Introduce the total loss function and dynamic label allocation strategy into the YOLOv8 network to obtain an improved YOLOv8 network; Training and verification steps: The preprocessed labeled data set is divided into a training set and a test set in a ratio of 8:

2. The training set is input into the improved YOLOv8 network for training to obtain a trained YOLOv8 network. The test set is input into the trained YOLOv8 network for verification to obtain a verified YOLOv8 network. Disease monitoring steps: Input the dental lesion image and the corresponding text information into the verified YOLOv8 network to obtain the lesion area type detection and text description; In the feature extraction step, the multi-scale feature extraction block MFEM includes: setting the number of input image channels to four channels, inputting the four-channel output image into the 1x1 convolution layer, the 3x3 convolution layer, the 5x5 convolution layer and the maximum pooling layer to obtain four corresponding tensors, concatenating the four tensors to form a new matrix tensor, fusing them through CotNet, generating a new feature information matrix and obtaining the corresponding feature vector X, and then inputting the feature vector X into CotNet to obtain the output result; In the error elimination step, the image feature f generated by the cosine similarity calculation model is i and text features f T The similarity between them is expressed as follows: If Sim(f i ,f T ) is higher than a certain threshold θ, and the sample is marked as a negative sample in the label dataset, it is marked as a false negative sample.

2. The method for detecting oral diseases based on an artificial intelligence large model according to claim 1, characterized in that: In the feature extraction step, the feature extraction block BERT uses the self-attention mechanism to calculate the influence of each word in the sentence on other words, and the masked language model MLM is used to pre-train the feature extraction block BERT.

3. The method for detecting oral diseases based on an artificial intelligence large model according to claim 1, characterized in that: CotNet includes: Convert feature vector X to Queries: Q = XW Q ,Keys:K=XW K ,Values:V=XW V , perform k×k groups of convolutions on all adjacent keys in the k×k spatial grid to obtain the local context information K between adjacent keys 1 , the local context information of the image is concatenated with Q to obtain the attention matrix D, and the attention matrix D is multiplied by V to obtain the total context information K 2 , and the local context information K 1 Add to get the output result.

4. The method for detecting oral diseases based on an artificial intelligence large model according to claim 1, characterized in that: In the network improvement step, the total loss function expression is: L total =θ cls L cls +θ loc L loc +θ scale L scale , Among them, θ cls ,θ loc ,θ scale is the weight hyperparameter, L cls is the classification loss, L loc is the positioning loss, L scale is the scale loss.

Citation Information

Patent Citations

  • Ultrasonic diagnosis intelligent interaction system based on liver attribute analysis

    CN117333462A

  • Deep learning dental implant classification method based on text prompt training

    CN118570523A

Cited By

  • Lightweight oral image AI diagnosis system

    CN120932865A