Oral disease auxiliary detection method, device, equipment and medium

Through the combination of multimodal fusion processing and deep learning models, the entities and relationships in oral disease data are integrated and identified, and a dynamic updated knowledge graph is built, which solves the problem of low data integration and recognition efficiency in the existing technology, and achieves efficient and accurate oral disease assisted detection and question-and-answer support.

CN120144730AActive Publication Date: 2025-06-13ZHEJIANG UNIV

Patent Information

Application Number
CN202510629255.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively integrate and utilize diverse oral disease data, identify complex professional terms and relationships, and be timely compatible with new research results, resulting in insufficient efficiency and accuracy of assisted oral disease detection.

Method used

Multimodal fusion processing technology is used to fuse the image and text data in oral disease data, entity recognition is used using BioBERT+BiLSTM+CRF model, relationship extraction is used using BioBERT+Text-CNN model to build a knowledge graph, and a question-and-answer model is constructed through a large language model pre-trained by RAG.

Benefits of technology

It realizes effective integration and identification of oral disease data, improves the efficiency and accuracy of oral disease auxiliary detection, can dynamically update the knowledge graph, maintain the timeliness of knowledge, and provide highly accurate question-and-answer support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144730A_ABST
    Figure CN120144730A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an oral disease auxiliary detection method, device, equipment and medium, oral disease data is comprehensively collected according to a configuration data collection strategy, and multi-modal fusion is performed on the oral disease data to realize effective integration of heterogeneous data in the field of oral medical treatment; according to the method, a BioBERT + BiLSTM + CRF model is used for entity recognition, a BioBERT + Text-CNN model is used for relation extraction, oral medical entities and relations thereof are accurately recognized and extracted, and therefore construction of a comprehensive and fine oral disease knowledge graph is assisted; according to the technical scheme, the knowledge graph is subjected to vector indexing, the big language model based on RAG pre-training is accessed, high-interactivity knowledge query and retrieval can be performed, decision support is provided for oral clinicians and researchers, the graph can be expanded and dynamically updated, and the timeliness of knowledge can be kept, so that the question and answer result is more accurate and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and medium for assisting in the detection of oral diseases. Background Art

[0002] In recent years, people have paid more and more attention to oral health, and how to assist medical staff in quickly diagnosing oral diseases has become a key issue of concern. However, the existing technologies mainly have the following problems:

[0003] First of all, the data sources in the field of oral medicine are diverse, including clinical data, literature data, experimental data, etc., and these data often exist in different formats and structures. How to effectively integrate these heterogeneous data and use them for assisting in the detection of oral diseases is the primary problem faced.

[0004] Secondly, there are many professional terms in the field of oral medicine, and there are a large number of synonyms and near-synonyms. How to effectively distinguish and identify these words is also the key to the research.

[0005] Finally, with the continuous in-depth research of oral medicine, new research results emerge continuously. How to effectively be compatible with the new results is also a problem that needs to be solved in the research.

[0006] In view of this, it is necessary to provide a method for assisting in the detection of oral diseases that is efficient, accurate and easy to update. Summary of the Invention

[0007] In view of the above content, it is necessary to provide a method, device, equipment and medium for assisting in the detection of oral diseases, aiming to solve the problem of being unable to effectively assist in the detection of oral diseases.

[0008] A method for assisting in the detection of oral diseases, the method for assisting in the detection of oral diseases includes: Collect oral disease data according to the configured data collection strategy; Perform multi-modal fusion processing on the oral disease data to obtain target features; Use the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and use the BioBERT+Text-CNN model to perform relationship extraction on the target features to obtain target relationships; Construct an ontology, and construct triple data based on the target entities and the target relationships; Integrate the ontology and the triple data and import them into the configured database to obtain a target knowledge graph; Perform vector indexing on the knowledge graph and connect to a large language model pre-trained based on RAG to obtain an oral disease Q&A model; In response to an oral disease assistance detection instruction, obtain a user's question; Use the oral disease Q&A model to respond to the user's question and obtain a response result; Send the response result to the trigger of the oral disease auxiliary detection instruction.

[0009] According to a preferred embodiment of the present invention, the collection of oral disease data according to the configured data collection strategy includes: Retrieve literature resources related to oral medicine in a biomedical literature database, and collect clinical guidelines and clinical pathway texts related to oral medicine as unstructured data; Extract medical record nodes, test results, and examination images related to oral medicine from outpatient electronic medical records as semi-structured data; Retrieve drug data related to oral medicine in a drug database as structured data.

[0010] According to a preferred embodiment of the present invention, the multi-modal fusion processing of the oral disease data to obtain target features includes: Identify image data and text data in the oral disease data; Extract the first feature of the image data from the region of interest of the candidate region network pooling layer; Extract a 4D vector for the region border localization of the image data as the second feature; Convert the image attribute category of the image data into a one-hot encoding to obtain the third feature; Combine the first feature, the second feature, and the third feature to obtain the fourth feature; Input the fourth feature into a fully connected layer for conversion to obtain a target image feature aligned with the text data; Use the BioBERT model to extract word vectors of the text data to obtain target text features; Use the self-attention mechanism to learn the intra-modal correlation of the target image feature, and use a one-dimensional convolutional neural network to encode the context information of the target text feature to learn the intra-modal correlation of the target text feature; Use the cross-attention mechanism to capture the inter-modal correlation between the target image feature and the target text feature; After completing the learning, obtain the structural features of the oral disease data, use multi-modal factorized bilinear pooling to fuse the target image feature and the target text feature, and fuse the target image feature, the target text feature, and the corresponding structural features through a gating network to obtain the target feature.

[0011] According to a preferred embodiment of the present invention, the entity recognition of the target feature by using the BioBERT+BiLSTM+CRF model to obtain the target entity includes: Initialize the first BioBERT model; wherein, the first BioBERT model has the same hyperparameters as the BERT model, and the dimension after word vectorization is 1024 dimensions; Split each sentence of length K in the target feature into word vectors, segment vectors, and position vectors with dimensions of (1, K, 1024); Add the word vectors, segment vectors, and position vectors element by element to obtain a target vector; Input the target vector into the first BioBERT model for training to obtain a text vector representation; Input the text vector representation into the BiLSTM layer to obtain text semantic information; Input the text semantic information into the CRF layer to obtain the predicted label of each word in the target feature; Obtain a pre-constructed dictionary; wherein, the dictionary includes standardized medical terms and entity information; Based on the dictionary, perform entity alignment on the predicted label of each word to obtain the target entity.

[0012] According to a preferred embodiment of the present invention, the entity alignment of the predicted label of each word based on the dictionary to obtain the target entity includes: Match the predicted label of each word with the dictionary to map the predicted label of each word to the normalized dictionary; Configure a unique identification code for the entity of each word in the dictionary after mapping, so that the same entity has a unified identification code in different texts or data sources; Determine each entity after configuring the identification code as the target entity.

[0013] According to a preferred embodiment of the present invention, the relationship extraction of the target feature by using the BioBERT+Text-CNN model to obtain the target relationship includes: Initialize the second BioBERT model; wherein, the second BioBERT model uses the masked language model mechanism to predict the masked word by using the unmasked word to learn the lexical semantics and context information in the text, and predicts whether the next sentence is randomly replaced through the next sentence prediction mechanism to capture the discourse structure and semantic coherence of the text; Input the target feature into the second BioBERT model to perform text processing on the target feature by using a multi-layer Transformer structure to obtain the semantic vector representation of each word; Input the semantic vector representation of each word into the Text-CNN model to perform convolution operations on the semantic vector representation of each word through convolutional kernels of different sizes to obtain convolution results; Use the max pooling layer of the Text-CNN model to perform dimensionality reduction on the convolution results to extract the maximum value output by each convolutional kernel; Input the maximum value output by each convolutional kernel into the classification layer of the Text-CNN model to obtain the target relationship.

[0014] According to a preferred embodiment of the present invention, after using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and using the BioBERT+Text-CNN model to perform relationship extraction on the target features to obtain target relationships, the method further includes: Construct a validation set; Based on the validation set, evaluate the accuracy of the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model to obtain an evaluation result; According to the evaluation result, adjust the hyperparameters and model structure of the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model.

[0015] An oral disease auxiliary detection device, the oral disease auxiliary detection device includes: A collection unit for collecting oral disease data according to a configured data collection strategy; A processing unit for performing multimodal fusion processing on the oral disease data to obtain target features; An extraction unit for using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and using the BioBERT+Text-CNN model to perform relationship extraction on the target features to obtain target relationships; A construction unit for constructing an ontology and constructing triple data based on the target entities and the target relationships; An integration unit for integrating the ontology and the triple data and importing them into a configured database to obtain a target knowledge graph; An access unit for performing vector indexing on the knowledge graph and accessing a large language model pre-trained based on RAG to obtain an oral disease Q&A model; An acquisition unit for obtaining a user's question in response to an oral disease auxiliary detection instruction; A response unit for using the oral disease Q&A model to respond to the user's question to obtain a response result; A sending unit, configured to send the response result to the trigger of the oral disease auxiliary detection instruction.

[0016] A computer device, comprising: A memory, storing at least one instruction; and A processor, configured to execute the instruction stored in the memory to implement the oral disease auxiliary detection method.

[0017] A computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in a computer device to implement the oral disease auxiliary detection method.

[0018] As can be seen from the above technical solutions, the present invention can comprehensively collect oral disease data according to the configured data collection strategy, and perform multi-modal fusion processing on the oral disease data to achieve effective integration of heterogeneous data in the field of oral medicine; use the BioBERT+BiLSTM+CRF model for entity recognition, and use the BioBERT+Text-CNN model for relationship extraction to accurately identify and extract oral medical entities and their relationships, so as to assist in constructing a comprehensive and detailed oral disease knowledge graph; perform vector indexing on the knowledge graph and connect to a large language model pre-trained based on RAG, which can perform highly interactive knowledge query and retrieval, provide decision support for oral clinicians and researchers, and because the graph is extensible and can be dynamically updated, it can also maintain the timeliness of knowledge, so that the Q&A results are more accurate and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flowchart of a preferred embodiment of the oral disease auxiliary detection method of the present invention; Figure 2 is a functional module diagram of a preferred embodiment of the oral disease auxiliary detection device of the present invention; Figure 3 is a structural schematic diagram of a computer device of a preferred embodiment for implementing the oral disease auxiliary detection method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] As Figure 1 shown, it is a flowchart of a preferred embodiment of the oral disease auxiliary detection method of the present invention. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.

[0022] The oral disease auxiliary detection method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0023] The computer device can be any electronic product that can interact with users. For example, personal computers, tablets, smartphones, personal digital assistants (PDAs), game consoles, Internet Protocol Television (IPTV), smart wearable devices, etc.

[0024] The computer device may also include network devices and / or user devices. Among them, the network devices include, but are not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0025] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0026] Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0027] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0028] The network where the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.

[0029] S10. Collect oral disease data according to the configured data collection strategy.

[0030] In this embodiment, the collecting oral disease data according to the configured data collection strategy includes: Retrieve literature resources related to oral healthcare in biomedical literature databases, and collect clinical guidelines and clinical pathway texts related to oral healthcare as unstructured data; Extract medical record nodes, test results, and examination images related to oral healthcare from outpatient electronic medical records as semi-structured data; Retrieve drug data related to oral healthcare in drug databases as structured data.

[0031] For example: Biomedical literature databases such as PubMed can be used to retrieve literature resources related to oral healthcare, obtain literature datasets including drug side effects, gene expression, and those related to oral diseases, and collect clinical guidelines and clinical pathway texts related to the oral cavity as the unstructured data.

[0032] For another example: Medical record nodes such as treatment plans, past medical history, current medical history, and chief complaints, as well as test results and examination images can be extracted from outpatient electronic medical records as the semi-structured data, and the semi-structured data can be converted into a structured data structure using natural language processing.

[0033] For yet another example: Data such as drug interactions, metabolic enzymes, and drug targets can be obtained from drug databases such as DrugBank, drug side effect data can be obtained using the SIDER (Side Effect Resource) database, drug target data can be obtained through the BindingDB (Bindind Database), TTD (Therapeutic Target Database), and DrugBank, and drug-gene relationship data can be obtained from the DGIdb (Drug–Gene Interaction Database) and LINCSL1000 (The Library of Integrated Network-Based Cellular Signatures L1000), and used as the structured data.

[0034] Through the above embodiments, oral disease data can be comprehensively collected according to a certain data collection strategy, providing sufficient data support for subsequent processing.

[0035] S11. Perform multi-modal fusion processing on the oral disease data to obtain target features.

[0036] In this embodiment, the performing multi-modal fusion processing on the oral disease data to obtain target features includes: Identify the image data and text data in the oral disease data; Extract the first feature of the image data from the region of interest (ROI) of the candidate region network (Region Proposal Network, RPN) pooling layer; Extract the 4D vector for image data region border localization as the second feature; Convert the image attribute category of the image data into one-hot encoding to obtain the third feature; Combine the first feature, the second feature, and the third feature to obtain the fourth feature; Input the fourth feature into a fully connected layer for conversion to obtain a target image feature aligned with the text data; Use the BioBERT (Bidirectional Encoder Representations from Transformers, biomedical language representation model) model to extract the word vectors of the text data to obtain the target text feature; Use the self-attention mechanism to learn the intra-modal correlation of the target image feature, and use a one-dimensional convolutional neural network (Convolutional Neural Networks, CNN) to encode the context information of the target text feature to learn the intra-modal correlation of the target text feature; Use the cross-attention mechanism to capture the inter-modal correlation between the target image feature and the target text feature; After completing the learning, obtain the structural feature of the oral disease data, use multi-modal factorized bilinear pooling to fuse the target image feature and the target text feature, and fuse the target image feature, the target text feature, and the corresponding structural feature through a gating network to obtain the target feature.

[0037] Among them, the first feature can describe the visual features of each region of interest in the image.

[0038] Among them, the second feature is used to represent the position and size of each detected object, usually a 4D vector containing the coordinates of the upper left corner and the lower right corner.

[0039] Among them, the third feature is used to represent the category to which the object belongs.

[0040] Among them, the fourth feature is input into a fully connected layer for transformation, aiming to enable image features and text features to be compared and fused in the same feature space.

[0041] Among them, the self-attention mechanism is used to learn the intra-modal correlation of the target image features. The self-attention mechanism enables the model to focus on the associations between different parts of the image when processing the image. For example, when processing medical images, through the self-attention mechanism, the model can focus on the relationship between the lesion area and the surrounding tissues. In an image, each pixel point or feature block can be regarded as a "word", and the self-attention mechanism calculates the association weights between them, so as to better capture the semantic information in the image.

[0042] Among them, a one-dimensional convolutional neural network is used to encode the context information of the target text features to learn the intra-modal correlation of the target text features. The one-dimensional convolutional neural network is suitable for processing sequential data such as text. It slides different-sized convolutional kernels over the text sequence to extract context information of different lengths. For example, when analyzing the medical record text of oral diseases, the one-dimensional convolutional neural network can capture the key phrases in the symptom description and the relationships between them.

[0043] Among them, the cross-attention mechanism is used to capture the inter-modal correlation between the target image features and the target text features. The cross-attention mechanism is used to connect two different modalities of images and texts. In the field of oral cavity, it can establish the connection between objects such as teeth and gums in the image and the words describing these parts in the text. For example, when the text mentions "red and swollen gums", the cross-attention mechanism enables the model to focus on the corresponding gum area in the image, so as to capture the cross-modal correlation.

[0044] Among them, the structural features may include, but are not limited to: geometric structures such as the shape, size, and position of image data, hierarchical features existing in the image at different resolutions, the grammatical structure and discourse structure of text data, the association relationships between data, etc.

[0045] Among them, the Multimodal factorized bilinear pooling (MFB) is used to fuse the target image features and the target text features. MFB is used to fuse text and visual features. It generates a new feature representation through bilinear interaction of the two features, and can capture the complex interaction relationship between the two modalities. For example, in oral disease diagnosis, oral image features and medical record text features are fused to obtain more comprehensive information.

[0046] Among them, the target image features, the target text features and the corresponding structural features are fused through a gating network. On the basis of MFB fusion, the gating network further fuses visual features, text features and structural features. The gating network determines the importance of different features by learning weights, so as to obtain the final feature representation. For example, when constructing an oral disease knowledge graph, this fusion method can integrate various information to improve the performance of the model.

[0047] In the above embodiment, multi-modal features of each entity can be extracted by optimizing the cross-modal related learning algorithm.

[0048] S12, using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and using the BioBERT+Text-CNN model to perform relation extraction on the target features to obtain target relations.

[0049] In this embodiment, the using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities includes: Initializing the first BioBERT model; among them, the first BioBERT model has the same hyperparameters as the BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture), and the dimension after word vectorization is 1024 dimensions; Splitting each sentence of length K in the target features into word vectors, segment vectors and position vectors with dimensions of (1, K, 1024); Adding the word vectors, segment vectors and position vectors element by element to obtain a target vector; Inputting the target vector into the first BioBERT model for training to obtain a text vector representation; Input the text vector representation into a BiLSTM (Bidirectional Long Short-Term Memory) layer to obtain text semantic information; Input the text semantic information into a CRF (Conditional Random Fields) layer to obtain the predicted label for each word in the target feature; Obtain a pre-constructed dictionary; wherein, the dictionary includes standardized medical terms and entity information; Perform entity alignment on the predicted label of each word based on the dictionary to obtain the target entity.

[0050] Among them, the first BioBERT model can be the BioBERTv1.1 version, and the first BioBERT model is a version that has been pre-trained on PubMed for 1M steps.

[0051] Next, a specific example will be used to illustrate the method of adding the word vector, segment vector, and position vector element by element: Suppose the word vector generated by the word "oral cavity" is (This is only for illustration, actually 1024 values), with a dimension of 1×1024. Since the sentence has K = 8 words, the overall dimension of the word vector is (1, 8, 1024).

[0052] The segment vector is used to distinguish different text segments or sentences. In the case of single-sentence input, generally the entire sentence belongs to the same segment, so the segment vector corresponding to each word is the same. Suppose the generated segment vector is with a dimension of 1×1024. Similarly, for the entire sentence, the dimension of the segment vector is also (1, 8, 1024), and the segment vector values at each position are the same.

[0053] The position vector represents the position information of the word in the sentence. For the word at position i , a corresponding 1024-dimensional position vector will be generated. For example, "oral cavity" is the 2nd word in the sentence (counting from 1), and suppose its position vector is with a dimension of 1×1024. The dimension of the position vector for the entire sentence is (1, 8, 1024).

[0054] For each word in the sentence, add the corresponding word vector, segment vector, and position vector element-wise to obtain the target vector v2 = w2 + s + p2. For example: v2[0] = w2[0] + s[0] + p2[0] = 0.1 + 0.05 + 0.01 = 0.16, and so on, to obtain a new 1024-dimensional vector v2, which is the new vector representation of the word "oral cavity" after integrating the word vector, segment vector, and position vector information.

[0055] For other words in the sentence, perform the element-wise addition operation on these three vectors in the same way. Finally, obtain the new vector representation after integrating the entire sentence, and the dimension is still (1, 8, 1024), where each position (corresponding to each word) has a fused 1024-dimensional vector.

[0056] Among them, inputting the text vector representation into the BiLSTM layer can process the text both forward and backward simultaneously, capture context information, and learn more comprehensive semantic features. Through the processing of the BiLSTM layer, further extract and integrate the semantic information in the text, providing more discriminative features for subsequent entity recognition.

[0057] Among them, inputting the text semantic information into the CRF layer can consider the dependency relationships between labels, and use the context information between words in the sentence and the constraint relationships between labels to predict the labels of each word, thereby achieving accurate named entity recognition and automatically identifying oral medical-related entities, including drugs, genes, and diseases, etc.

[0058] Among them, the dictionary can include the UMLS (Unified Medical Language System) vocabulary, MeSH (Medical Subject Headings) vocabulary, and EntrezGene (National Center for Biotechnology Information Gene Database), etc.

[0059] Among them, the entity alignment of the predicted labels of each word based on the dictionary to obtain the target entity includes: Match the predicted label of each word with the dictionary to map the predicted label of each word to the normalized dictionary; Configure a unique identification code for the entity of each word in the dictionary, so that the same entity has a unified identification code in different texts or data sources; Determine each entity after configuring the identification code as the target entity.

[0060] Through the above embodiments, entities can be accurately and efficiently recognized.

[0061] In this embodiment, the obtaining of the target relationship by performing relation extraction on the target feature using the BioBERT+Text-CNN model includes: Initialize the second BioBERT model; wherein, the second BioBERT model uses the masked language model mechanism to predict masked words using unmasked words to learn lexical semantics and context information in the text, and predicts whether the next sentence is randomly replaced through the next sentence prediction mechanism to capture the discourse structure and semantic coherence of the text; Input the target feature into the second BioBERT model to perform text processing on the target feature using a multi-layer Transformer structure to obtain the semantic vector representation of each word; Input the semantic vector representation of each word into the Text-CNN model to perform convolution operations on the semantic vector representation of each word using convolution kernels of different sizes to obtain a convolution result; Use the max pooling layer of the Text-CNN model to perform dimensionality reduction processing on the convolution result to extract the maximum value output by each convolution kernel; Input the maximum value output by each convolution kernel into the classification layer of the Text-CNN model to obtain the target relationship.

[0062] Among them, more comprehensive semantic vector expression information in the biomedical field can be obtained through the second BioBERT model.

[0063] Among them, after inputting the target feature into the second BioBERT model, the semantic vector representation of each word obtained contains rich context information and biomedical semantic features, providing a basis for subsequent relation classification.

[0064] Among them, the Text-CNN model performs convolution operations on the input semantic vectors using convolution kernels of different sizes to capture important local semantic information, which may be closely related to the relationship between entities. And after being processed by the convolution layer, the max pooling layer is used to perform dimensionality reduction on the convolution result to extract the maximum value output by each convolution kernel, further highlighting the key semantic features.

[0065] Among them, the classification layer performs classification prediction of entity relationships according to features, that is, determines the specific relationship between entities.

[0066] In this embodiment, after obtaining the target entity by performing entity recognition on the target feature using the BioBERT+BiLSTM+CRF model and obtaining the target relationship by performing relation extraction on the target feature using the BioBERT+Text-CNN model, the method further includes: Construct a validation set; Accurately evaluate the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model based on the validation set to obtain an evaluation result; Adjust the hyperparameters and model structure of the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model according to the evaluation result.

[0067] Among them, the performance of the model can be measured by calculating indicators such as the accuracy rate, recall rate, and F1 value of the corresponding model.

[0068] Through the above embodiments, the model can be adjusted in a timely manner according to the accuracy rate of the model, thereby ensuring the reliability of the identified entities and the extracted relationships.

[0069] S13. Construct an ontology and construct triple data based on the target entity and the target relationship.

[0070] In this embodiment, the ontology can be constructed using ontology editing and knowledge acquisition software.

[0071] Among them, the triple can include forms such as (gene, association, disease).

[0072] S14. Integrate the ontology and the triple data and import them into the configuration database to obtain a target knowledge graph.

[0073] This embodiment effectively integrates heterogeneous data such as text descriptions and oral images for knowledge graph embedding, and integrates multi-modal data such as research literature, electronic medical records, and clinical guidelines in the field of oral medicine. It uses deep learning algorithms to mine various types of entities and complex relationships between entities, and at the same time incorporates existing knowledge in public databases, thereby constructing a comprehensive and detailed oral disease knowledge graph.

[0074] In this embodiment, when there are new research results or medical data, the knowledge graph can be expanded and updated to maintain the reliability of the knowledge graph.

[0075] S15. Perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG (Retrieval-augmented Generation) to obtain an oral disease Q&A model.

[0076] This embodiment constructs a Q&A system based on the knowledge graph using the RAG technology of the large language model, provides an effective knowledge management platform, can support highly interactive knowledge query and retrieval, automatically design treatment plans and disposal plans, and provide decision-making support for oral clinical medical staff and scientific research personnel.

[0077] In this embodiment, since the knowledge graph is scalable, updating the knowledge graph can indirectly assist the large language model to respond more accurately.

[0078] S16. In response to an oral disease auxiliary detection instruction, obtain the user's question.

[0079] In this embodiment, the oral disease auxiliary detection instruction can be triggered through a specified interface.

[0080] In this embodiment, the user's question can be input in text form or in voice form.

[0081] S17. Use the oral disease Q&A model to respond to the user's question to obtain a response result.

[0082] In this embodiment, the oral disease Q&A model can use the knowledge graph as a knowledge base to respond to the user's question. Due to the comprehensiveness of the knowledge graph, the answer can be more accurate.

[0083] S18. Send the response result to the trigger of the oral disease auxiliary detection instruction.

[0084] Among them, the trigger can be medical staff, scientific research personnel, etc.

[0085] This embodiment efficiently integrates multi-modal heterogeneous data in the field of oral medicine, accurately identifies and extracts oral medical entities and their relationships, and constructs an expandable and dynamically updated knowledge graph, which can facilitate data update and maintenance, so that the knowledge of the Q&A system remains timely, improves the accuracy of the Q&A system's response, and thus provides effective support for relevant personnel to assist in the detection of oral diseases.

[0086] It can be seen from the above technical solutions that the present invention can comprehensively collect oral disease data according to the configured data collection strategy, and perform multi-modal fusion processing on the oral disease data to achieve the effective integration of heterogeneous data in the field of oral medicine; use the BioBERT+BiLSTM+CRF model for entity recognition and the BioBERT+Text-CNN model for relationship extraction to accurately identify and extract oral medical entities and their relationships, so as to assist in constructing a comprehensive and detailed oral disease knowledge graph; perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG, which can perform highly interactive knowledge query and retrieval, provide decision support for oral clinicians and scientific research personnel, and because the graph is expandable and can be dynamically updated, it can also maintain the timeliness of knowledge, so that the Q&A results are more accurate and reliable.

[0087] Such as Figure 2As shown, it is a functional module diagram of a preferred embodiment of the oral disease auxiliary detection device of the present invention. The oral disease auxiliary detection device 11 includes a collection unit 110, a processing unit 111, an extraction unit 112, a construction unit 113, an integration unit 114, an access unit 115, an acquisition unit 116, a response unit 117, and a sending unit 118. The module / unit referred to in the present invention means a series of computer program segments that can be executed by a processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0088] Among them, the collection unit 110 is used to collect oral disease data according to the configured data collection strategy; The processing unit 111 is used to perform multi-modal fusion processing on the oral disease data to obtain target features; The extraction unit 112 is used to perform entity recognition on the target features using the BioBERT+BiLSTM+CRF model to obtain target entities, and perform relationship extraction on the target features using the BioBERT+Text-CNN model to obtain target relationships; The construction unit 113 is used to construct an ontology and construct triple data based on the target entities and the target relationships; The integration unit 114 is used to integrate the ontology and the triple data and import them into the configuration database to obtain a target knowledge graph; The access unit 115 is used to perform vector indexing on the knowledge graph and access a large language model pre-trained based on RAG to obtain an oral disease Q&A model; The acquisition unit 116 is used to obtain a user's question in response to an oral disease auxiliary detection instruction; The response unit 117 is used to respond to the user's question using the oral disease Q&A model to obtain a response result; The sending unit 118 is used to send the response result to the trigger of the oral disease auxiliary detection instruction.

[0089] As can be seen from the above technical solutions, the present invention can comprehensively collect oral disease data according to the configured data collection strategy, and perform multi-modal fusion processing on the oral disease data to achieve the effective integration of heterogeneous data in the field of oral medicine; use the BioBERT+BiLSTM+CRF model for entity recognition and the BioBERT+Text-CNN model for relation extraction to accurately identify and extract oral medical entities and their relationships, thereby assisting in constructing a comprehensive and detailed oral disease knowledge graph; perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG, enabling highly interactive knowledge query and retrieval, providing decision support for oral clinicians and researchers, and since the graph is extensible and can be dynamically updated, it can also maintain the timeliness of knowledge, thus making the question-and-answer results more accurate and reliable.

[0090] As Figure 3 shown, it is a schematic structural diagram of a computer device of a preferred embodiment for implementing the oral disease assisted detection method of the present invention.

[0091] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure is the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as an oral disease assisted detection program.

[0092] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 can be either a bus structure or a star structure. The computer device 1 may also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the computer device 1 may also include input / output devices, network access devices, etc.

[0093] It should be noted that the computer device 1 is only an example, and other existing or future possible electronic products that can be adapted to the present invention should also be included in the protection scope of the present invention and are included herein by reference.

[0094] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as the mobile hard disk of the computer device 1. In other embodiments, the memory 12 can also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the computer device 1. The memory 12 can not only be used to store application software installed on the computer device 1 and various types of data, such as the code of the oral disease auxiliary detection program, etc., but also be used to temporarily store the data that has been output or will be output.

[0095] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including the combination of one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the computer device 1, connecting various components of the entire computer device 1 through various interfaces and lines. By running or executing the programs or modules stored in the memory 12 (such as executing the oral disease auxiliary detection program, etc.), and calling the data stored in the memory 12, it can execute various functions of the computer device 1 and process data.

[0096] The processor 13 executes the operating system of the computer device 1 and various installed application programs. The processor 13 executes the application program to implement the steps in the above-mentioned embodiments of various oral disease auxiliary detection methods, such as Figure 1 the steps shown.

[0097] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an acquisition unit 110, a processing unit 111, an extraction unit 112, a construction unit 113, an integration unit 114, an access unit 115, a retrieval unit 116, a response unit 117, and a transmission unit 118.

[0098] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above-mentioned software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the oral disease assisted detection method described in various embodiments of the present invention.

[0099] If the modules / units integrated in the computer device 1 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented.

[0100] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, etc.

[0101] Further, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0102] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. A blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0103] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, in Figure 3 it is only represented by a single straight line, but it does not mean that there is only one bus or one type of bus. The bus is arranged to implement the connection and communication between the memory 12 and at least one processor 13, etc.

[0104] Although not shown, the computer device 1 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 13 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The computer device 1 may further include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0105] Furthermore, the computer device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the computer device 1 and other computer devices.

[0106] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the computer device 1 and to display a visual user interface.

[0107] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0108] Those skilled in the art can understand that Figure 3 the structure shown does not constitute a limitation on the computer device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0109] In combination with Figure 1 , the memory 12 in the computer device 1 stores multiple instructions to implement an oral disease assisted detection method, and the processor 13 can execute the multiple instructions to achieve: Collect oral disease data according to the configured data collection strategy; Perform multi-modal fusion processing on the oral disease data to obtain target features; Use the BioBERT+BiLSTM+CRF model to perform entity recognition on the target features to obtain target entities, and use the BioBERT+Text-CNN model to perform relationship extraction on the target features to obtain target relationships; Construct an ontology, and construct triple data based on the target entities and the target relationships; Integrate the ontology and the triple data and import them into the configuration database to obtain a target knowledge graph; Perform vector indexing on the knowledge graph and connect it to a large language model pre-trained based on RAG to obtain an oral disease Q&A model; In response to an oral disease assisted detection instruction, obtain a user's question; Use the oral disease Q&A model to respond to the user's question to obtain a response result; Send the response result to the trigger of the oral disease assisted detection instruction.

[0110] Specifically, the specific implementation method of the above instructions by the processor 13 can refer to Figure 1Descriptions of relevant steps in corresponding embodiments are not elaborated herein.

[0111] It should be noted that all data involved in this case are legally obtained. The non-company software tools or components appearing in the embodiments of this application are only introduced by way of example and do not represent actual use.

[0112] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0113] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0114] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional module.

[0116] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0117] Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

[0118] In addition, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. A plurality of elements or devices recited in the present invention may also be implemented by one element or device through software or hardware. The terms such as "first" and "second" are used to denote names and do not denote any particular order.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An oral disease auxiliary detection method, characterized in that: The oral disease auxiliary detection method comprises: Collect oral disease data according to the configured data collection strategy; Performing multimodal fusion processing on the oral disease data to obtain target features; Using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target feature to obtain the target entity, and using the BioBERT+Text-CNN model to perform relationship extraction on the target feature to obtain the target relationship; Constructing an ontology, and constructing triple data based on the target entity and the target relationship; Integrate the ontology and the triple data and import them into a configuration database to obtain a target knowledge graph; Performing vector indexing on the knowledge graph and accessing a large language model pre-trained based on RAG to obtain an oral disease question-answering model; In response to the oral disease auxiliary detection instruction, obtaining a user question; Using the oral disease question-answering model to respond to the user's question and obtain a response result; The response result is sent to the trigger of the oral disease auxiliary detection instruction.

2. The oral disease auxiliary detection method according to claim 1, characterized in that: The collecting of oral disease data according to the configured data collection strategy includes: Search for literature resources related to oral medicine in the biomedical literature database, and collect clinical guidelines and clinical pathway texts related to oral medicine as unstructured data; Extract medical record nodes, test results, and examination images related to oral care from outpatient electronic medical records as semi-structured data; Drug data related to oral healthcare is retrieved from a drug database as structured data.

3. The oral disease auxiliary detection method according to claim 1, characterized in that: The multimodal fusion processing is performed on the oral disease data to obtain target features including: Identifying image data and text data in the oral disease data; Extracting a first feature of the image data from a region of interest of a candidate region network pooling layer; Extracting a 4-dimensional vector of the border positioning of the image data area as a second feature; Convert the image attribute category of the image data into a one-hot encoding to obtain a third feature; Combining the first feature, the second feature and the third feature to obtain a fourth feature; Inputting the fourth feature into a fully connected layer for conversion to obtain a target image feature aligned with the text data; The BioBERT model is used to extract word vectors of the text data to obtain target text features; Using a self-attention mechanism to learn the intra-modal correlation of the target image features, and using a one-dimensional convolutional neural network to encode the context information of the target text features to learn the intra-modal correlation of the target text features; Using a cross-attention mechanism to capture the inter-modal correlation between the target image features and the target text features; After completing the learning, the structural features of the oral disease data are obtained, the target image features and the target text features are fused using multimodal factor decomposition bilinear pooling, and the target image features, the target text features and the corresponding structural features are fused through a gating network to obtain the target features.

4. The oral disease auxiliary detection method according to claim 1, characterized in that: The method of using the BioBERT+BiLSTM+CRF model to perform entity recognition on the target feature to obtain the target entity includes: Initialize a first BioBERT model; wherein the first BioBERT model has the same hyperparameters as the BERT model, and the dimension of the word vectorization is 1024 dimensions; Split each sentence of length K in the target feature into a word vector, a segment vector, and a position vector of dimension (1, K, 1024); Add the word vector, segment vector and position vector element by element to obtain a target vector; Inputting the target vector into the first BioBERT model for training to obtain a text vector representation; Input the text vector representation into the BiLSTM layer to obtain text semantic information; Input the text semantic information into the CRF layer to obtain the predicted label of each word in the target feature; Obtaining a pre-built dictionary; wherein the dictionary includes standardized medical terms and entity information; The predicted label of each word is aligned based on the dictionary to obtain the target entity.

5. The oral disease auxiliary detection method according to claim 4, characterized in that: The predicted label of each word is aligned based on the dictionary to obtain the target entity, which includes: Matching the predicted label of each word with the dictionary to map the predicted label of each word to the normalized dictionary; Assign a unique identification code to each entity of the mapped word in the dictionary, so that the same entity has a unified identification code in different texts or data sources; Each entity configured with the identification code is determined as the target entity.

6. The oral disease auxiliary detection method according to claim 1, characterized in that: The method of using the BioBERT+Text-CNN model to extract the relationship of the target features to obtain the target relationship includes: Initializing a second BioBERT model; wherein the second BioBERT model uses a masked language model mechanism to predict masked words using unmasked words to learn lexical semantics and contextual information in the text, and predicts whether the next sentence is randomly replaced through a next sentence prediction mechanism to capture the text's paragraph structure and semantic coherence; Inputting the target feature into the second BioBERT model to perform text processing on the target feature using a multi-layer Transformer structure to obtain a semantic vector representation of each word; The semantic vector representation of each word is input into the Text-CNN model to perform convolution operations on the semantic vector representation of each word through convolution kernels of different sizes to obtain the convolution results; Performing dimensionality reduction processing on the convolution result by using the maximum pooling layer of the Text-CNN model to extract the maximum value of each convolution kernel output; The maximum value of each convolution kernel output is input into the classification layer of the Text-CNN model to obtain the target relationship.

7. The oral disease auxiliary detection method according to claim 1, characterized in that: After the target entity is identified by using the BioBERT+BiLSTM+CRF model, and the target relationship is extracted by using the BioBERT+Text-CNN model, the method further includes: Construct a validation set; Based on the validation set, the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model are evaluated for accuracy to obtain evaluation results; The hyperparameters and model structure of the BioBERT+BiLSTM+CRF model and the BioBERT+Text-CNN model are adjusted according to the evaluation results.

8. An oral disease auxiliary detection device, characterized in that: The oral disease auxiliary detection device comprises: A collection unit, for collecting oral disease data according to a configured data collection strategy; A processing unit, used for performing multimodal fusion processing on the oral disease data to obtain target features; An extraction unit, used to perform entity recognition on the target feature using a BioBERT+BiLSTM+CRF model to obtain a target entity, and to perform relationship extraction on the target feature using a BioBERT+Text-CNN model to obtain a target relationship; A construction unit, used for constructing an ontology and constructing triple data based on the target entity and the target relationship; An integration unit, used for integrating the ontology and the triple data and importing them into a configuration database to obtain a target knowledge graph; An access unit, used to perform vector indexing on the knowledge graph and access a large language model pre-trained based on RAG to obtain an oral disease question-answering model; An acquisition unit, used to respond to the oral disease auxiliary detection instruction and acquire the user's question; A response unit, used to respond to the user's question using the oral disease question-answering model to obtain a response result; A sending unit is used to send the response result to the trigger of the oral disease auxiliary detection instruction.

9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the oral disease auxiliary detection method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the oral disease auxiliary detection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video text matching model training method and device and video text matching method and device

    CN115204301A

  • Target segmentation method based on non-local feature aggregation neural network

    CN115223080A

  • Medical field entity and relation extraction method based on training model

    CN116383392A

  • Large language model knowledge question-answering method and system fused with multi-modal knowledge graph

    CN118627628A

  • Clinical examination result auditing method and system based on artificial intelligence and big data

    CN118629571A

Cited By

  • Oral disease image report generation method and device, equipment and medium

    CN120452657A

  • Temporomandibular joint disease diagnosis method and system based on large language model

    CN121354862A