Oral Disease Auxiliary Detection Method, Device, Equipment and Medium
The method integrates and processes diverse dental data using advanced models to construct a knowledge graph, enhancing oral disease detection accuracy and adaptability.
Patent Information
- Application Number
- CN202510629255.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The prior art is difficult to effectively integrate heterogeneous data in the oral medical field, identify and extract entities and their relationships in oral diseases, and it is impossible to quickly update new results.
The BioBERT+BiLSTM+CRF model is used for entity recognition and the BioBERT+Text-CNN model for relationship extraction, construct oral disease knowledge graph, and use the RAG pre-trained large language model for question-and-answer to achieve fusion and dynamic update of multimodal data.
It realizes efficient integration and accurate identification of oral disease data, provides a scalable knowledge graph, supports highly interactive knowledge query and retrieval, provides decision support for clinicians and researchers, and ensures the timeliness and accuracy of Q&A results.
Smart Images

Figure CN120144730B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an oral disease auxiliary detection method, device, equipment and medium. Background Art
[0002] In recent years, people have paid more and more attention to oral health, and how to assist medical staff in quickly diagnosing oral diseases has become a key issue of concern. However, the existing technologies mainly have the following problems:
[0003] First of all, the data sources in the field of oral medicine are diverse, including clinical data, literature data, experimental data, etc., and these data often exist in different formats and structures. How to effectively integrate these heterogeneous data and use them for oral disease auxiliary detection is the primary problem faced.
[0004] Secondly, there are many professional terms in the field of oral medicine, and there are a large number of synonyms and near-synonyms. How to effectively distinguish and identify these words is also the key to the research.
[0005] Finally, with the continuous in-depth research of oral medicine, new research results emerge continuously. How to effectively accommodate the new results is also a problem that needs to be solved in the research.
[0006] In view of this, it is necessary to provide an oral disease auxiliary detection method that is efficient, accurate and easy to update. Summary of the Invention
[0007] In view of the above content, it is necessary to provide an oral disease auxiliary detection method, device, equipment and medium, aiming to solve the problem of being unable to effectively assist in detecting oral diseases.
[0008] An oral disease auxiliary detection method, the oral disease auxiliary detection method includes:
[0009] Collect oral disease data according to the configured data collection strategy;
[0010] Perform multi-modal fusion processing on the oral disease data to obtain target features;
[0011] Use the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and use the BioBERT+Text-CNN model to perform relationship extraction on the target features to obtain target relationships;
[0012] Construct an ontology, and construct triple data based on the target entities and the target relationships;
[0013] Integrate the ontology and the triple data and import them into the configured database to obtain a target knowledge graph;
[0014] Perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG to obtain an oral disease Q&A model;
[0015] In response to an oral disease auxiliary detection instruction, obtain the user's question;
[0016] Use the oral disease Q&A model to respond to the user's question to obtain a response result;
[0017] Send the response result to the trigger of the oral disease auxiliary detection instruction.
[0018] According to a preferred embodiment of the present invention, the collecting of oral disease data according to the configured data collection strategy includes:
[0019] Retrieve literature resources related to oral medicine in a biomedical literature database, and collect clinical guidelines and clinical pathway texts related to oral medicine as unstructured data;
[0020] Extract medical record nodes, test results, and examination images related to oral medicine in outpatient electronic medical records as semi-structured data;
[0021] Retrieve drug data related to oral medicine in a drug database as structured data.
[0022] According to a preferred embodiment of the present invention, the multi-modal fusion processing of the oral disease data to obtain target features includes:
[0023] Identify the image data and text data in the oral disease data;
[0024] Extract the first feature of the image data from the region of interest of the candidate region network pooling layer;
[0025] Extract a 4D vector for image data region border localization as the second feature;
[0026] Convert the image attribute category of the image data into one-hot encoding to obtain the third feature;
[0027] Combine the first feature, the second feature, and the third feature to obtain the fourth feature;
[0028] Input the fourth feature into a fully connected layer for conversion to obtain target image features aligned with the text data;
[0029] Use the BioBERT model to extract word vectors of the text data to obtain target text features;
[0030] Learn the intra-modal correlation of the target image features using the self-attention mechanism, and use a one-dimensional convolutional neural network to encode the context information of the target text features to learn the intra-modal correlation of the target text features;
[0031] Use the cross-attention mechanism to capture the inter-modal correlation between the target image features and the target text features;
[0032] After completing the learning, obtain the structural features of the oral disease data, fuse the target image features and the target text features using multi-modal factorized bilinear pooling, and fuse the target image features, the target text features and the corresponding structural features through a gating network to obtain the target features.
[0033] According to a preferred embodiment of the present invention, the entity recognition of the target features using the BioBERT+BiLSTM+CRF model to obtain the target entity includes:
[0034] Initialize the first BioBERT model; wherein, the first BioBERT model has the same hyperparameters as the BERT model, and the dimension after word vectorization is 1024 dimensions;
[0035] Split each sentence of length K in the target features into word vectors, segment vectors and position vectors with dimensions of (1, K, 1024);
[0036] Add the word vectors, segment vectors and position vectors element-wise to obtain a target vector;
[0037] Input the target vector into the first BioBERT model for training to obtain a text vector representation;
[0038] Input the text vector representation into the BiLSTM layer to obtain text semantic information;
[0039] Input the text semantic information into the CRF layer to obtain the predicted label of each word in the target features;
[0040] Obtain a pre-constructed dictionary; wherein, the dictionary includes standardized medical terms and entity information;
[0041] Based on the dictionary, perform entity alignment on the predicted label of each word to obtain the target entity.
[0042] According to a preferred embodiment of the present invention, the entity alignment of the predicted label of each word based on the dictionary to obtain the target entity includes:
[0043] Match the predicted label of each word with the dictionary to map the predicted label of each word to the normalized dictionary;
[0044] Configure a unique identification code for the entity of each word in the dictionary after mapping, so that the same entity has a unified identification code in different texts or data sources;
[0045] Determine each entity after configuring the identification code as the target entity.
[0046] According to a preferred embodiment of the present invention, the obtaining the target relationship by using the BioBERT+Text-CNN model to perform relation extraction on the target feature includes:
[0047] Initialize the second BioBERT model; wherein, the second BioBERT model uses the masked language model mechanism to predict masked words with unmasked words to learn the lexical semantics and context information in the text, and predicts whether the next sentence is randomly replaced through the next sentence prediction mechanism to capture the discourse structure and semantic coherence of the text;
[0048] Input the target feature into the second BioBERT model to perform text processing on the target feature by using a multi-layer Transformer structure to obtain the semantic vector representation of each word;
[0049] Input the semantic vector representation of each word into the Text-CNN model to perform a convolution operation on the semantic vector representation of each word through convolution kernels of different sizes to obtain a convolution result;
[0050] Use the max pooling layer of the Text-CNN model to perform dimensionality reduction processing on the convolution result to extract the maximum value output by each convolution kernel;
[0051] Input the maximum value output by each convolution kernel into the classification layer of the Text-CNN model to obtain the target relationship.
[0052] According to a preferred embodiment of the present invention, after using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target feature to obtain the target entity, and using the BioBERT+Text-CNN model to perform relation extraction on the target feature to obtain the target relationship, the method further includes:
[0053] Construct a validation set;
[0054] Based on the validation set, perform accuracy evaluation on the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model to obtain an evaluation result;
[0055] Adjust the hyperparameters and model structure of the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model according to the evaluation results.
[0056] An oral disease auxiliary detection device, the oral disease auxiliary detection device includes:
[0057] A collection unit for collecting oral disease data according to a configured data collection strategy;
[0058] A processing unit for performing multimodal fusion processing on the oral disease data to obtain target features;
[0059] An extraction unit for using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and using the BioBERT+Text-CNN model to perform relationship extraction on the target features to obtain target relationships;
[0060] A construction unit for constructing an ontology and constructing triple data based on the target entities and the target relationships;
[0061] An integration unit for integrating the ontology and the triple data and importing them into a configured database to obtain a target knowledge graph;
[0062] An access unit for performing vector indexing on the knowledge graph and accessing a large language model pre-trained based on RAG to obtain an oral disease Q&A model;
[0063] An acquisition unit for obtaining a user's question in response to an oral disease auxiliary detection instruction;
[0064] A response unit for using the oral disease Q&A model to respond to the user's question to obtain a response result;
[0065] A sending unit for sending the response result to the trigger of the oral disease auxiliary detection instruction.
[0066] A computer device, the computer device includes:
[0067] A memory storing at least one instruction; and
[0068] A processor for executing the instructions stored in the memory to implement the oral disease auxiliary detection method.
[0069] A computer-readable storage medium storing at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the oral disease auxiliary detection method.
[0070] As can be seen from the above technical solutions, the present invention can comprehensively collect oral disease data according to the configured data collection strategy, and perform multi-modal fusion processing on the oral disease data to achieve the effective integration of heterogeneous data in the field of oral medicine; use the BioBERT+BiLSTM+CRF model for entity recognition, and use the BioBERT+Text-CNN model for relationship extraction to accurately identify and extract oral medical entities and their relationships, so as to assist in constructing a comprehensive and detailed oral disease knowledge graph; perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG, which can perform highly interactive knowledge query and retrieval, provide decision-making support for oral clinicians and researchers, and because the graph is extensible and can be dynamically updated, it can also maintain the timeliness of knowledge, thus making the Q&A results more accurate and reliable. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 is a flowchart of a preferred embodiment of the oral disease assisted detection method of the present invention;
[0072] Figure 2 is a functional module diagram of a preferred embodiment of the oral disease assisted detection device of the present invention;
[0073] Figure 3 is a schematic structural diagram of a computer device of a preferred embodiment for implementing the oral disease assisted detection method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0075] As Figure 1 shown, it is a flowchart of a preferred embodiment of the oral disease assisted detection method of the present invention. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted.
[0076] The oral disease assisted detection method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0077] The computer device can be any electronic product that can interact with users, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0078] The computer device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0079] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0080] Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0081] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0082] The network where the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0083] S10, collect oral disease data according to the configured data collection strategy.
[0084] In this embodiment, collecting oral disease data according to the configured data collection strategy includes:
[0085] Retrieve literature resources related to oral medicine in the biomedical literature database, and collect clinical guidelines and clinical pathway texts related to oral medicine as unstructured data;
[0086] Extract the medical record nodes, test results, and examination images related to oral healthcare from outpatient electronic medical records as semi-structured data;
[0087] Retrieve drug data related to oral healthcare from a drug database as structured data.
[0088] For example: Biomedical literature databases such as PubMed can be used to retrieve literature resources related to oral healthcare, obtain literature datasets including drug side effects, gene expression, and those related to oral diseases, and collect clinical guidelines and clinical pathway texts related to the oral cavity as the unstructured data.
[0089] Another example: Medical record nodes such as treatment plans, past medical history, current medical history, and chief complaints, as well as test results and examination images, can be extracted from outpatient electronic medical records as the semi-structured data, and the semi-structured data can be converted into a structured data structure using natural language processing.
[0090] Another example: Data such as drug interactions, metabolic enzymes, and drug targets can be obtained from drug databases such as DrugBank, drug side effect data can be obtained using the SIDER (Side Effect Resource) database, drug target data can be obtained through BindingDB (Bindind Database), TTD (Therapeutic Target Database), and DrugBank, and drug-gene relationship data can be obtained from DGIdb (Drug–Gene Interaction Database) and LINCSL1000 (The Library of Integrated Network-Based Cellular Signatures L1000), and used as the structured data.
[0091] Through the above embodiments, oral disease data can be comprehensively collected according to a certain data collection strategy, providing sufficient data support for subsequent processing.
[0092] S11. Perform multimodal fusion processing on the oral disease data to obtain target features.
[0093] In this embodiment, the performing multimodal fusion processing on the oral disease data to obtain target features includes:
[0094] Identify the image data and text data in the oral disease data;
[0095] Extract the first feature of the image data from the Region of Interest (ROI) of the Region Proposal Network (RPN) pooling layer;
[0096] Extract a 4D vector for the region bounding box localization of the image data as the second feature;
[0097] Convert the image attribute category of the image data into one-hot encoding to obtain the third feature;
[0098] Combine the first feature, the second feature, and the third feature to obtain the fourth feature;
[0099] Input the fourth feature into a fully connected layer for transformation to obtain the target image feature aligned with the text data;
[0100] Use the BioBERT (Bidirectional Encoder Representations from Transformers) model to extract the word vectors of the text data to obtain the target text feature;
[0101] Utilize the self-attention mechanism to learn the intra-modal correlation of the target image feature, and use a one-dimensional Convolutional Neural Networks (CNN) to encode the context information of the target text feature to learn the intra-modal correlation of the target text feature;
[0102] Use the cross-attention mechanism to capture the inter-modal correlation between the target image feature and the target text feature;
[0103] After completing the learning, obtain the structural feature of the oral disease data, use multi-modal factorized bilinear pooling to fuse the target image feature and the target text feature, and fuse the target image feature, the target text feature, and the corresponding structural feature through a gating network to obtain the target feature.
[0104] Among them, the first feature can describe the visual features of each region of interest in the image.
[0105] Among them, the second feature is used to represent the position and size of each detected object, usually a 4D vector containing the coordinates of the upper left corner and the lower right corner.
[0106] Among them, the third feature is used to represent the category to which the object belongs.
[0107] Among them, the fourth feature is input into a fully connected layer for conversion, aiming to enable the comparison and fusion of image features and text features in the same feature space.
[0108] Among them, the self-attention mechanism is used to learn the intra-modal correlation of the target image features. The self-attention mechanism enables the model to focus on the associations between different parts of the image when processing the image. For example, when processing medical images, through the self-attention mechanism, the model can focus on the relationship between the lesion area and the surrounding tissues. In an image, each pixel point or feature block can be regarded as a "word", and the self-attention mechanism calculates the association weights between them, so as to better capture the semantic information within the image.
[0109] Among them, a one-dimensional convolutional neural network is used to encode the context information of the target text features to learn the intra-modal correlation of the target text features. The one-dimensional convolutional neural network is suitable for processing sequential data such as text. It slides different-sized convolutional kernels over the text sequence to extract context information of different lengths. For example, when analyzing the medical record text of oral diseases, the one-dimensional convolutional neural network can capture the key phrases in the symptom description and the relationships between them.
[0110] Among them, the cross-attention mechanism is used to capture the inter-modal correlation between the target image features and the target text features. The cross-attention mechanism is used to connect two different modalities of images and texts. In the oral field, it can establish the connection between objects such as teeth and gums in the image and the words describing these parts in the text. For example, when the text mentions "red and swollen gums", the cross-attention mechanism enables the model to focus on the corresponding gum area in the image, thereby capturing the cross-modal correlation.
[0111] Among them, the structural features may include, but are not limited to: geometric structures such as the shape, size, and position of image data, hierarchical features existing in the image at different resolutions, the grammatical structure and discourse structure of text data, the association relationships between data, etc.
[0112] Among them, multimodal factorized bilinear pooling (MFB) is used to fuse the target image features and the target text features. MFB is used to fuse text and visual features. It generates a new feature representation through bilinear interaction of the two features, and can capture the complex interaction relationships between the two modalities. For example, in the diagnosis of oral diseases, oral image features and medical record text features are fused to obtain more comprehensive information.
[0113] Among them, the gating network fuses the target image feature, the target text feature and the corresponding structural feature. On the basis of MFB fusion, the gating network further fuses visual features, text features and structural features. The gating network determines the importance of different features by learning weights, so as to obtain the final feature representation. For example, when constructing an oral disease knowledge graph, this fusion method can integrate various information to improve the performance of the model.
[0114] In the above embodiment, the cross-modal related learning algorithm is optimized to extract the multi-modal features of each entity.
[0115] S12. Use the BioBERT+BiLSTM+CRF model to perform entity recognition on the target feature to obtain target entities, and use the BioBERT+Text-CNN model to perform relation extraction on the target feature to obtain target relations.
[0116] In this embodiment, the use of the BioBERT+BiLSTM+CRF model to perform entity recognition on the target feature to obtain target entities includes:
[0117] Initialize the first BioBERT model; among them, the first BioBERT model has the same hyperparameters as the BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture), and the dimension after word vectorization is 1024 dimensions;
[0118] Split each sentence of length K in the target feature into word vectors, segment vectors and position vectors with dimensions of (1, K, 1024);
[0119] Add the word vectors, segment vectors and position vectors element by element to obtain a target vector;
[0120] Input the target vector into the first BioBERT model for training to obtain a text vector representation;
[0121] Input the text vector representation into a BiLSTM (Bidirectional Long Short-Term Memory) layer to obtain text semantic information;
[0122] Input the text semantic information into a CRF (Conditional Random Fields) layer to obtain the predicted label of each word in the target feature;
[0123] Obtain a pre-built dictionary; wherein, the dictionary includes standardized medical terms and entity information;
[0124] Perform entity alignment on the predicted labels of each word based on the dictionary to obtain the target entity.
[0125] Wherein, the first BioBERT model can be the BioBERTv1.1 version, and the first BioBERT model is a version that has been pre-trained for 1M steps on PubMed.
[0126] Next, a specific example will be used to illustrate the method of element-wise addition of the word vector, segment vector, and position vector:
[0127] Suppose the word vector generated for the word "oral cavity" is (only for illustration here, actually 1024 values), with a dimension of 1×1024. Since the sentence has K = 8 words, the overall dimension of the word vector is (1, 8, 1024).
[0128] The segment vector is used to distinguish different text segments or sentences. In the case of single-sentence input, generally the entire sentence belongs to the same segment, so the segment vector corresponding to each word is the same. Suppose the generated segment vector is with a dimension of 1×1024. Similarly, for the entire sentence, the dimension of the segment vector is also (1, 8, 1024), and the segment vector values at each position are the same.
[0129] The position vector represents the position information of the word in the sentence. For the word at position i a corresponding 1024-dimensional position vector will be generated. For example, "oral cavity" is the 2nd word in the sentence (counting from 1), and suppose its position vector is with a dimension of 1×1024. The dimension of the position vector for the entire sentence is (1, 8, 1024).
[0130] For each word in the sentence, add its corresponding word vector, segment vector, and position vector element-wise, then the target vector v2 = w2 + s + p2 is obtained. For example: v2[0] = w2[0] + s[0] + p2[0] = 0.1 + 0.05 + 0.01 = 0.16, and so on, to obtain a new 1024-dimensional vector v2, which is the new vector representation of the word "oral cavity" after fusing the word vector, segment vector, and position vector information.
[0131] For the other words in the sentence, perform the element-wise addition operation on these three vectors in the same way. Finally, the new vector representation after fusing the entire sentence is obtained, and the dimension is still (1, 8, 1024), where each position (corresponding to each word) has a fused 1024-dimensional vector.
[0132] Among them, inputting the text vector representation into the BiLSTM layer can process the text from both the forward and backward directions simultaneously, capture context information, and learn more comprehensive semantic features. Through the processing of the BiLSTM layer, the semantic information in the text is further extracted and integrated, providing more discriminative features for subsequent entity recognition.
[0133] Among them, inputting the text semantic information into the CRF layer can consider the dependency relationships between tags, and use the context information between words in the sentence and the constraint relationships between tags to predict the tags of each word, thereby achieving accurate named entity recognition and automatically identifying oral medical-related entities, including drugs, genes, and diseases, etc.
[0134] Among them, the dictionary may include the UMLS (Unified Medical Language System) vocabulary, the MeSH (Medical Subject Headings) vocabulary, and EntrezGene (the National Center for Biotechnology Information Gene Database), etc.
[0135] Among them, the entity alignment of the predicted tags of each word based on the dictionary to obtain the target entity includes:
[0136] Match the predicted tag of each word with the dictionary to map the predicted tag of each word to the normalized dictionary;
[0137] Configure a unique identification code for the entity of each word in the dictionary, so that the same entity has a unified identification code in different texts or data sources;
[0138] Determine each entity after configuring the identification code as the target entity.
[0139] Through the above embodiments, entities can be accurately and efficiently recognized.
[0140] In this embodiment, the obtaining of the target relationship by extracting the relationship of the target features using the BioBERT+Text-CNN model includes:
[0141] Initialize the second BioBERT model; among them, the second BioBERT model uses the masked language model mechanism to predict masked words using unmasked words to learn the lexical semantics and context information in the text, and predicts whether the next sentence is randomly replaced through the next sentence prediction mechanism to capture the discourse structure and semantic coherence of the text;
[0142] Input the target feature into the second BioBERT model to perform text processing on the target feature using a multi-layer Transformer structure, and obtain the semantic vector representation of each word;
[0143] Input the semantic vector representation of each word into the Text-CNN model to perform convolution operations on the semantic vector representation of each word through convolutional kernels of different sizes, and obtain the convolution result;
[0144] Use the max pooling layer of the Text-CNN model to perform dimensionality reduction processing on the convolution result to extract the maximum value output by each convolutional kernel;
[0145] Input the maximum value output by each convolutional kernel into the classification layer of the Text-CNN model to obtain the target relationship.
[0146] Among them, more comprehensive semantic vector expression information in the biomedical field can be obtained through the second BioBERT model.
[0147] Among them, after the target feature is input into the second BioBERT model, the semantic vector representation of each word obtained contains rich context information and biomedical semantic features, providing a basis for subsequent relationship classification.
[0148] Among them, the Text-CNN model performs convolution operations on the input semantic vectors through convolutional kernels of different sizes to capture important local semantic information, which may be closely related to the relationship between entities. And after being processed by the convolutional layer, the max pooling layer is used to perform dimensionality reduction on the convolution result to extract the maximum value output by each convolutional kernel, further highlighting the key semantic features.
[0149] Among them, the classification layer performs classification prediction of entity relationships according to features, that is, determines the specific relationship between entities.
[0150] In this embodiment, after using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target feature to obtain the target entity, and using the BioBERT+Text-CNN model to perform relationship extraction on the target feature to obtain the target relationship, the method further includes:
[0151] Construct a validation set;
[0152] Based on the validation set, perform accuracy evaluation on the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model to obtain the evaluation result;
[0153] Adjust the hyperparameters and model structures of the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model according to the evaluation results.
[0154] Among them, the performance of the model can be measured by calculating indicators such as the accuracy, recall rate, and F1 value of the corresponding model.
[0155] Through the above embodiments, the model can be adjusted in a timely manner according to the accuracy of the model, thereby ensuring the reliability of the identified entities and the extracted relationships.
[0156] S13, construct an ontology and construct triple data based on the target entity and the target relationship.
[0157] In this embodiment, the ontology can be constructed using ontology editing and knowledge acquisition software.
[0158] Among them, the triples can include forms such as (gene, association, disease).
[0159] S14, integrate the ontology and the triple data and import them into the configuration database to obtain the target knowledge graph.
[0160] This embodiment effectively integrates heterogeneous data such as text descriptions and oral images for knowledge graph embedding, and integrates multi-modal data such as research literature, electronic medical records, and clinical guidelines in the field of oral medicine. It uses deep learning algorithms to mine various entities and complex relationships between entities, and at the same time incorporates existing knowledge in public databases, thereby constructing a comprehensive and detailed oral disease knowledge graph.
[0161] In this embodiment, when there are new research results or medical data, the knowledge graph can be extended and updated to maintain the reliability of the knowledge graph.
[0162] S15, perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG (Retrieval-augmented Generation) to obtain an oral disease Q&A model.
[0163] This embodiment constructs a Q&A system based on the knowledge graph using the large language model RAG technology, provides an effective knowledge management platform, can support highly interactive knowledge query and retrieval, automatically design treatment plans and disposal plans, and provide decision-making support for oral clinical medical staff and scientific research personnel.
[0164] In this embodiment, since the knowledge graph has scalability, by updating the knowledge graph, it can indirectly assist the large language model to respond more accurately.
[0165] S16. In response to the oral disease assisted detection instruction, obtain the user's question.
[0166] In this embodiment, the oral disease assisted detection instruction can be triggered through a specified interface.
[0167] In this embodiment, the user's question can be input in text form or in voice form.
[0168] S17. Use the oral disease Q&A model to respond to the user's question and obtain a response result.
[0169] In this embodiment, the oral disease Q&A model can use the knowledge graph as a knowledge base to respond to the user's question. Due to the comprehensiveness of the knowledge graph, the answer can be made more accurate.
[0170] S18. Send the response result to the trigger of the oral disease assisted detection instruction.
[0171] Among them, the trigger can be medical staff, scientific research personnel, etc.
[0172] This embodiment efficiently integrates multi-modal heterogeneous data in the field of oral medicine, accurately identifies and extracts oral medical entities and their relationships, and constructs an extensible and dynamically updated knowledge graph, which can facilitate data update and maintenance, so that the knowledge of the Q&A system remains up-to-date, improves the accuracy of the response of the Q&A system, and thus provides effective support for relevant personnel to assist in the detection of oral diseases.
[0173] It can be seen from the above technical solutions that the present invention can comprehensively collect oral disease data according to the configured data collection strategy, and perform multi-modal fusion processing on the oral disease data to achieve the effective integration of heterogeneous data in the field of oral medicine; use the BioBERT+BiLSTM+CRF model for entity recognition and the BioBERT+Text-CNN model for relationship extraction to accurately identify and extract oral medical entities and their relationships, so as to assist in constructing a comprehensive and detailed oral disease knowledge graph; perform vector indexing on the knowledge graph and connect to a large language model pre-trained based on RAG, which can perform highly interactive knowledge query and retrieval, provide decision-making support for oral clinicians and scientific research personnel, and because the graph is extensible and can be dynamically updated, it can also maintain the timeliness of knowledge, so that the Q&A results are more accurate and reliable.
[0174] Such as Figure 2As shown, it is a functional block diagram of a preferred embodiment of the oral disease auxiliary detection device of the present invention. The oral disease auxiliary detection device 11 includes a collection unit 110, a processing unit 111, an extraction unit 112, a construction unit 113, an integration unit 114, an access unit 115, an acquisition unit 116, a response unit 117, and a sending unit 118. The module / unit referred to in the present invention means a series of computer program segments that can be executed by a processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0175] Among them, the collection unit 110 is used to collect oral disease data according to the configured data collection strategy;
[0176] The processing unit 111 is used to perform multi-modal fusion processing on the oral disease data to obtain target features;
[0177] The extraction unit 112 is used to perform entity recognition on the target features using the BioBERT+BiLSTM+CRF model to obtain target entities, and perform relationship extraction on the target features using the BioBERT+Text-CNN model to obtain target relationships;
[0178] The construction unit 113 is used to construct an ontology and construct triple data based on the target entities and the target relationships;
[0179] The integration unit 114 is used to integrate the ontology and the triple data and import them into the configuration database to obtain a target knowledge graph;
[0180] The access unit 115 is used to perform vector indexing on the knowledge graph and access a large language model pre-trained based on RAG to obtain an oral disease Q&A model;
[0181] The acquisition unit 116 is used to obtain a user's question in response to an oral disease auxiliary detection instruction;
[0182] The response unit 117 is used to respond to the user's question using the oral disease Q&A model to obtain a response result;
[0183] The sending unit 118 is used to send the response result to the trigger of the oral disease auxiliary detection instruction.
[0184] As can be seen from the above technical solutions, the present invention can comprehensively collect oral disease data according to the configured data collection strategy, and perform multi-modal fusion processing on the oral disease data to achieve the effective integration of heterogeneous data in the field of oral medicine; use the BioBERT+BiLSTM+CRF model for entity recognition, and use the BioBERT+Text-CNN model for relationship extraction to accurately identify and extract oral medical entities and their relationships, so as to assist in constructing a comprehensive and detailed oral disease knowledge graph; perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG, which can perform highly interactive knowledge query and retrieval, provide decision-making support for oral clinicians and researchers, and because the graph is extensible and can be dynamically updated, it can also maintain the timeliness of knowledge, so that the question-and-answer results are more accurate and reliable.
[0185] As Figure 3 shown, it is a schematic structural diagram of a computer device of a preferred embodiment for implementing the oral disease auxiliary detection method of the present invention.
[0186] The computer device 1 may include a memory 12, a processor 13 and a bus (the arrow in the figure is the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as an oral disease auxiliary detection program.
[0187] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 can be either a bus structure or a star structure. The computer device 1 may also include more or fewer other hardware or software than shown, or different component arrangements. For example, the computer device 1 may also include input and output devices, network access devices, etc.
[0188] It should be noted that the computer device 1 is only an example, and other existing or future possible electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are included herein by reference.
[0189] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as the mobile hard disk of the computer device 1. In some other embodiments, the memory 12 can also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the computer device 1. The memory 12 can be used not only to store application software installed in the computer device 1 and various types of data, such as the code of the oral disease auxiliary detection program, etc., but also to temporarily store data that has been output or will be output.
[0190] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged together, including the combination of one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the computer device 1, connecting various components of the entire computer device 1 through various interfaces and circuits. By running or executing programs or modules stored in the memory 12 (such as executing the oral disease auxiliary detection program, etc.), and by calling the data stored in the memory 12, it can execute various functions of the computer device 1 and process data.
[0191] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application program to implement the steps in the above-mentioned embodiments of various oral disease auxiliary detection methods, such as Figure 1 the steps shown.
[0192] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an acquisition unit 110, a processing unit 111, an extraction unit 112, a construction unit 113, an integration unit 114, an access unit 115, an acquisition unit 116, a response unit 117, and a sending unit 118.
[0193] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the oral disease auxiliary detection method described in various embodiments of the present invention.
[0194] If the modules / units integrated in the computer device 1 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present invention, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned method embodiments may be implemented.
[0195] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, etc.
[0196] Furthermore, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0197] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.
[0198] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, in Figure 3 it is only represented by a single straight line, but it does not mean that there is only one bus or one type of bus. The bus is arranged to realize the connection and communication between the memory 12 and at least one processor 13, etc.
[0199] Although not shown, the computer device 1 may further include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, so as to realize functions such as charge management, discharge management, and power consumption management through the power management device. The power supply may also include any components such as one or more DC or AC power supplies, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The computer device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0200] Furthermore, the computer device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.
[0201] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device 1 and to display a visual user interface.
[0202] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0203] Those skilled in the art can understand that Figure 3 the shown structure does not limit the computer device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0204] In combination with Figure 1 , the memory 12 in the computer device 1 stores multiple instructions to implement an oral disease assisted detection method, and the processor 13 can execute the multiple instructions to implement:
[0205] Collect oral disease data according to the configured data collection strategy;
[0206] Perform multimodal fusion processing on the oral disease data to obtain target features;
[0207] Use the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and use the BioBERT+Text-CNN model to perform relation extraction on the target features to obtain target relations;
[0208] Construct an ontology, and construct triple data based on the target entities and the target relations;
[0209] Integrate the ontology and the triple data and import them into the configuration database to obtain a target knowledge graph;
[0210] Perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG to obtain an oral disease Q&A model;
[0211] In response to an oral disease assisted detection instruction, obtain a user's question;
[0212] Use the oral disease Q&A model to respond to the user's question to obtain a response result;
[0213] Send the response result to the trigger of the oral disease auxiliary detection instruction.
[0214] Specifically, for the specific implementation method of the above instructions by the processor 13, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0215] It should be noted that all the data involved in this case are legally obtained. The non-company software tools or components appearing in the embodiments of this application are only for illustrative introduction and do not represent actual use.
[0216] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0217] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0218] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0219] In addition, in each embodiment of the present invention, the various functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0220] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0221] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0222] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the present invention can also be implemented by one unit or device through software or hardware. The terms such as "first" and "second" are used to denote names and do not denote any particular order.
[0223] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An auxiliary detection method for oral diseases, characterized in that, The oral disease auxiliary detection method includes: Collecting oral disease data according to a configured data collection strategy; Performing multi-modal fusion processing on the oral disease data to obtain target features, including: identifying image data and text data in the oral disease data; extracting first features of the image data from regions of interest in the candidate region network pooling layer; extracting a 4-dimensional vector for image data region border localization as second features; converting the image attribute category of the image data into one-hot encoding to obtain third features; combining the first features, the second features, and the third features to obtain fourth features; inputting the fourth features into a fully connected layer for transformation to obtain target image features aligned with the text data; using a BioBERT model to extract word vectors of the text data to obtain target text features; using a self-attention mechanism to learn the intra-modal correlation of the target image features, and using a one-dimensional convolutional neural network to encode the context information of the target text features to learn the intra-modal correlation of the target text features; using a cross-attention mechanism to capture the inter-modal correlation between the target image features and the target text features; after completing the learning, obtaining the structural features of the oral disease data, using multi-modal factorization bilinear pooling to fuse the target image features and the target text features, and using a gated network to fuse the target image features, the target text features, and the corresponding structural features to obtain the target features; Using a BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and using a BioBERT+Text-CNN model to perform relation extraction on the target features to obtain target relations; Constructing an ontology, and constructing triple data based on the target entities and the target relations; Integrating the ontology and the triple data and importing them into a configuration database to obtain a target knowledge graph; Performing vector indexing on the knowledge graph and connecting to a large language model pre-trained based on RAG to obtain an oral disease Q&A model; Responding to an oral disease auxiliary detection instruction to obtain a user's question; Using the oral disease Q&A model to respond to the user's question to obtain a response result; Sending the response result to the trigger of the oral disease auxiliary detection instruction.
2. The oral disease assisted detection method according to claim 1, wherein, The collecting oral disease data according to a configured data collection strategy includes: Retrieving literature resources related to oral medicine in a biomedical literature database, and collecting clinical guidelines and clinical pathway texts related to oral medicine as unstructured data; Extracting medical record record nodes, test results, and examination images related to oral medicine in outpatient electronic medical records as semi-structured data; Retrieving drug data related to oral medicine in a drug database as structured data.
3. The oral disease auxiliary detection method according to claim 1, wherein The using a BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities includes: Initialize the first BioBERT model; wherein, the first BioBERT model has the same hyperparameters as the BERT model, and the dimension after word vectorization is 1024 dimensions; Split each sentence of length K in the target feature into word vectors, segment vectors, and position vectors with dimensions of (1, K, 1024); Add the word vectors, segment vectors, and position vectors element-wise to obtain a target vector; Input the target vector into the first BioBERT model for training to obtain a text vector representation; Input the text vector representation into the BiLSTM layer to obtain text semantic information; Input the text semantic information into the CRF layer to obtain the predicted label for each word in the target feature; Obtain a pre-constructed dictionary; wherein, the dictionary includes standardized medical terms and entity information; Perform entity alignment on the predicted label of each word based on the dictionary to obtain the target entity.
4. The oral disease auxiliary detection method according to claim 3, wherein The performing entity alignment on the predicted label of each word based on the dictionary to obtain the target entity includes: Match the predicted label of each word with the dictionary to map the predicted label of each word to the normalized dictionary; Configure a unique identification code for the entity of each word in the dictionary after mapping, so that the same entity has a unified identification code in different texts or data sources; Determine each entity after configuring the identification code as the target entity.
5. The oral disease auxiliary detection method according to claim 1, characterized in that The performing relation extraction on the target feature by using the BioBERT+Text-CNN model to obtain the target relation includes: Initialize the second BioBERT model; wherein, the second BioBERT model uses the masked language model mechanism to predict masked words with unmasked words to learn the lexical semantics and context information in the text, and predicts whether the next sentence is randomly replaced through the next sentence prediction mechanism to capture the discourse structure and semantic coherence of the text; Input the target feature into the second BioBERT model to perform text processing on the target feature by using a multi-layer Transformer structure to obtain the semantic vector representation of each word; Input the semantic vector representation of each word into the Text-CNN model to perform a convolution operation on the semantic vector representation of each word through convolutional kernels of different sizes to obtain a convolution result; Use the max pooling layer of the Text-CNN model to perform dimensionality reduction processing on the convolution result to extract the maximum value output by each convolutional kernel; Input the maximum value output by each convolutional kernel into the classification layer of the Text-CNN model to obtain the target relation.
6. The oral disease auxiliary detection method according to claim 1, wherein, After performing entity recognition on the target feature by using the BioBERT+BiLSTM+CRF model to obtain the target entity, and performing relation extraction on the target feature by using the BioBERT+Text-CNN model to obtain the target relation, the method further includes: Construct a validation set; Perform accuracy evaluation on the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model based on the validation set to obtain an evaluation result; Adjust the hyperparameters and model structures of the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model according to the evaluation results.
7. An oral disease auxiliary detection device, characterized in that, The oral disease auxiliary detection device includes: A collection unit for collecting oral disease data according to a configured data collection strategy; A processing unit for performing multimodal fusion processing on the oral disease data to obtain target features, including: identifying image data and text data in the oral disease data; extracting first features of the image data from regions of interest in the candidate region network pooling layer; extracting a 4-dimensional vector for positioning the image data region border as second features; converting the image attribute category of the image data into one-hot encoding to obtain third features; combining the first features, the second features, and the third features to obtain fourth features; inputting the fourth features into a fully connected layer for conversion to obtain target image features aligned with the text data; using a BioBERT model to extract word vectors of the text data to obtain target text features; using a self-attention mechanism to learn the intra-modal correlation of the target image features, and using a one-dimensional convolutional neural network to encode the context information of the target text features to learn the intra-modal correlation of the target text features; using a cross-attention mechanism to capture the inter-modal correlation between the target image features and the target text features; after completing the learning, obtaining the structural features of the oral disease data, using multimodal factorized bilinear pooling to fuse the target image features and the target text features, and fusing the target image features, the target text features, and the corresponding structural features through a gating network to obtain the target features; An extraction unit for performing entity recognition on the target features using a BioBERT+BiLSTM+CRF model to obtain target entities, and performing relation extraction on the target features using a BioBERT+Text-CNN model to obtain target relations; A construction unit for constructing an ontology and constructing triple data based on the target entities and the target relations; An integration unit for integrating the ontology and the triple data and importing them into a configuration database to obtain a target knowledge graph; An access unit for performing vector indexing on the knowledge graph and accessing a large language model pre-trained based on RAG to obtain an oral disease Q&A model; An acquisition unit for obtaining a user's question in response to an oral disease auxiliary detection instruction; A response unit for using the oral disease Q&A model to respond to the user's question to obtain a response result; A sending unit for sending the response result to the trigger of the oral disease auxiliary detection instruction.
8. A computer device, characterized in that, The computer device includes: A memory storing at least one instruction; and A processor for executing the instructions stored in the memory to implement the oral disease auxiliary detection method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: At least one instruction is stored in the computer-readable storage medium, and the at least one instruction is executed by a processor in a computer device to implement the oral disease assisted detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Medical field entity and relation extraction method based on training model
CN116383392A
Large language model knowledge question-answering method and system fused with multi-modal knowledge graph
CN118627628A
Paper data availability classification method and device, equipment and storage medium
CN119226516A
Cited By
Oral diagnosis and treatment knowledge graph dynamic construction system based on deep learning
CN121278115A