Oral disease diagnosis and treatment question and answer method, device, equipment, medium and product
By constructing an oral healthcare knowledge graph and matching it with query subgraphs, the problem of insufficient accuracy and interpretability of oral disease diagnosis and treatment questions and answers in existing technologies is solved, resulting in more accurate and interpretable diagnosis and treatment answers.
Patent Information
- Application Number
- CN202510989927.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-18
AI Technical Summary
When using general large-scale language models or existing retrieval-enhanced generation (RAG) techniques for oral disease diagnosis and answering, the accuracy of the answers is low and the interpretability is lacking.
Based on the oral disease questions input by the user, a query subgraph is determined and matched with a pre-built oral medical knowledge graph. The complete subgraph is determined by the matching results and input into the diagnosis and treatment language model to obtain accurate diagnosis and treatment answers.
It improves the accuracy and interpretability of oral disease diagnosis and treatment Q&A, ensuring the professionalism and logical coherence of the answers.
Smart Images

Figure CN120973893A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of diagnosis and treatment question and answer, and in particular to a diagnosis and treatment question and answer method, device, equipment, medium and product for oral diseases. BACKGROUND
[0002] With the increasing public attention to oral health, the demand for oral medical services is rapidly growing. However, there is a relative shortage of high-level oral medical professionals, which is difficult to meet the huge diagnosis and treatment demand. This situation prompts people to try to use artificial intelligence technology to improve the efficiency and quality of oral disease diagnosis and treatment.
[0003] In the prior art, the following two methods are usually used for oral disease diagnosis and treatment question and answer: (1) a general large language model is used, but the general large language model is insufficient in understanding the professional terms and complex relationships in the oral disease field, resulting in low accuracy of answers; (2) the existing retrieval-augmented generation (RAG) technology is used, but the existing RAG technology mostly only uses text vector retrieval to obtain relevant content, which is difficult to handle the complex relationships of multiple text entities in the oral disease diagnosis and treatment problem, resulting in answers that may not be accurate or lack of explainability. SUMMARY
[0004] The present application provides a diagnosis and treatment question and answer method, device, equipment, medium and product for oral diseases, to solve the defects that the existing technology only uses a general large language model or uses the existing RAG technology for oral disease diagnosis and treatment question and answer, resulting in inaccurate answers or lack of explainability. The technical solution of the present application determines a query subgraph based on the oral disease problem input by the user, matches the query subgraph and a pre-constructed oral medical knowledge graph, determines a complete subgraph corresponding to the query subgraph according to the matching result, the complete subgraph can accurately represent the professional terms and complex relationships in the oral disease field, and then inputs a large language model to obtain an oral disease diagnosis and treatment answer, which can improve the accuracy and explainability of the diagnosis and treatment answer.
[0005] The present application provides a diagnosis and treatment question and answer method for oral diseases, comprising the following steps.
[0006] An oral disease problem input by a user is obtained, and a query subgraph is determined based on the oral disease problem; the query subgraph represents a structured oral disease problem; The query subgraph and an oral medical knowledge graph are matched, and a complete subgraph corresponding to the query subgraph is determined according to the matching result; the oral medical knowledge graph is constructed based on multi-source heterogeneous diagnosis and treatment text data in the oral disease field; inputting the complete subgraph and the oral disease question into a diagnosis and treatment large language model to obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model; the diagnosis and treatment large language model is obtained by training a general large language model based on oral disease diagnosis and treatment sample data.
[0007] According to the oral disease diagnosis and treatment question and answer method provided by the application, the query subgraph is determined based on the oral disease question, and the query subgraph includes: performing semantic analysis and recognition processing on the oral disease question to obtain a plurality of oral disease text entities corresponding to the oral disease question, and an entity type and a relationship structure corresponding to each oral disease text entity, respectively; determining the query subgraph based on each oral disease text entity, the entity type corresponding to each oral disease text entity, and the relationship structure.
[0008] According to the oral disease diagnosis and treatment question and answer method provided by the application, the oral medical knowledge graph is constructed in the following manner: obtaining the multi-source heterogeneous diagnosis and treatment text data in the oral disease field, and performing standardization processing on the multi-source heterogeneous diagnosis and treatment text data to obtain standard diagnosis and treatment text data; determining a plurality of standard diagnosis and treatment nodes in the oral medical knowledge graph based on the standard diagnosis and treatment text data; inputting each standard diagnosis and treatment node into a pre-trained semantic embedding model to obtain a diagnosis and treatment semantic vector corresponding to each standard diagnosis and treatment node output by the semantic embedding model, respectively; constructing the oral medical knowledge graph based on all the standard diagnosis and treatment nodes and all the diagnosis and treatment semantic vectors.
[0009] According to the oral disease diagnosis and treatment question and answer method provided by the application, the standard diagnosis and treatment text data includes textbook type text data and example type text data; determining a plurality of standard diagnosis and treatment nodes in the oral medical knowledge graph based on the standard diagnosis and treatment text data, including: splitting the textbook type text data into a plurality of first text blocks, and inputting all the first text blocks into a first general large language model to obtain a plurality of first text entities output by the first general large language model, and constructing a tree classification graph based on all the first text entities; splitting the example type text data into a plurality of second text blocks, and inputting all the second text blocks into the first general large language model to obtain a plurality of second text entities output by the first general large language model; performing alignment and fusion processing on all the first text entities and all the second text entities to obtain each standard diagnosis and treatment node after alignment and fusion.
[0010] According to the oral disease diagnosis and treatment question answering method provided by the application, the query subgraph and the oral medical knowledge graph are matched, and a complete subgraph corresponding to the query subgraph is determined according to a matching result, comprising: For each initial unknown node in all query nodes in the query subgraph, the initial unknown node is input into the semantic embedding model to obtain a query semantic vector corresponding to the initial unknown node output by the semantic embedding model; Determine the first similarity matching degree of the query semantic vector corresponding to the initial unknown node and all the diagnosis and treatment semantic vectors in the oral medical knowledge graph, respectively; The diagnosis and treatment semantic vector corresponding to the largest first similarity matching degree greater than the preset matching threshold is determined as the first semantic vector corresponding to the initial unknown node; and the complete node corresponding to the initial unknown node is determined based on the standard diagnosis and treatment node corresponding to the first semantic vector; For each initial known node in all query nodes in the query subgraph, the initial known node is matched based on the keyword corresponding to the initial known node and the oral medical knowledge graph, and the complete node corresponding to the initial known node is determined based on the standard diagnosis and treatment node matched in the oral medical knowledge graph; Based on all the complete nodes, the complete subgraph corresponding to the query subgraph is determined.
[0011] According to the oral disease diagnosis and treatment question answering method provided by the application, the method further comprises: For each initial known node, in the case that the keyword corresponding to the initial known node does not match the standard diagnosis and treatment node in the oral medical knowledge graph, the initial known node is input into the semantic embedding model to obtain a query semantic vector corresponding to the initial known node output by the semantic embedding model; the second similarity matching degree of the query semantic vector corresponding to the initial known node and all the diagnosis and treatment semantic vectors in the oral medical knowledge graph is determined; the diagnosis and treatment semantic vector corresponding to the largest second similarity matching degree greater than the preset matching threshold is determined as the second semantic vector corresponding to the initial known node; and the complete node corresponding to the initial known node is determined based on the standard diagnosis and treatment node corresponding to the second semantic vector.
[0012] According to the oral disease diagnosis and treatment question answering method provided by the application, the method further comprises: For each initial unknown node, in the case that all the first similarity matching degrees corresponding to the initial unknown node are less than the preset matching threshold, the initial unknown node is determined as a target unknown node; determine the initial known node as a target unknown node in a case that the keyword corresponding to the initial known node is not matched to the standard diagnosis and treatment node in the oral medical knowledge graph, and all the second similarity matching degrees corresponding to the initial known node are less than the preset matching threshold; For each of the target unknown nodes, determine the adjacent query node corresponding to the target unknown node based on the relationship structure corresponding to the target unknown node; determine the complete node corresponding to the target unknown node based on the standard diagnosis and treatment node corresponding to the adjacent query node.
[0013] According to the oral disease diagnosis and treatment question and answer method provided by the application, the complete subgraph and the oral disease question are input into the diagnosis and treatment large language model, and the oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model is obtained, which comprises: determine the initial diagnosis and treatment prompt word based on the node definition, node relationship structure and node attribute corresponding to the complete subgraph; determine the target diagnosis and treatment prompt word based on the initial diagnosis and treatment prompt word and the oral disease question; input the target diagnosis and treatment prompt word into the diagnosis and treatment large language model, and obtain the oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model.
[0014] According to the oral disease diagnosis and treatment question and answer method provided by the application, the method further comprises: input the preset oral disease diagnosis and treatment question bank into the diagnosis and treatment large language model, and obtain the diagnosis and treatment prediction answer corresponding to the preset oral disease diagnosis and treatment question bank output by the diagnosis and treatment large language model; input the preset evaluation prompt word, the preset oral disease diagnosis and treatment question bank, the diagnosis and treatment prediction answer, and the diagnosis and treatment standard answer corresponding to the preset oral disease diagnosis and treatment question bank into the second general large language model, and obtain the diagnosis and treatment evaluation result of the diagnosis and treatment large language model output by the second general large language model.
[0015] The application also provides an oral disease diagnosis and treatment question and answer device, comprising the following modules: a question module for acquiring an oral disease question input by a user and determining a query subgraph based on the oral disease question; the query subgraph represents a structured oral disease question; a matching module for matching the query subgraph and an oral medical knowledge graph, and determining a complete subgraph corresponding to the query subgraph according to the matching result; the oral medical knowledge graph is constructed based on multi-source heterogeneous diagnosis and treatment text data in the field of oral diseases; The answer module is used to input the complete subgraph and the oral disease question into the diagnostic language model to obtain the oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnostic language model; the diagnostic language model is obtained by training a general language model based on oral disease diagnosis and treatment sample data.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the oral disease diagnosis and treatment question-and-answer method as described above.
[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the oral disease diagnosis and treatment question-and-answer method as described above.
[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the oral disease diagnosis and treatment question-and-answer method as described above.
[0019] This invention provides a method, apparatus, device, medium, and product for answering oral disease diagnosis and treatment questions. It acquires oral disease questions input by a user and determines a query subgraph based on these questions. The query subgraph represents a structured oral disease question. The query subgraph is matched with an oral medical knowledge graph, and the complete subgraph corresponding to the query subgraph is determined based on the matching result. The oral medical knowledge graph is constructed based on multi-source heterogeneous diagnostic and treatment text data in the oral disease field. The complete subgraph and the oral disease question are input into a large-scale diagnostic and treatment language model to obtain the oral disease diagnosis and treatment answer corresponding to the question, output by the large-scale diagnostic and treatment language model. The large-scale diagnostic and treatment language model is obtained by training a general large-scale language model based on oral disease diagnosis and treatment sample data. This invention's technical solution determines a query subgraph based on user-input oral disease questions, matches the query subgraph with a pre-constructed oral medical knowledge graph, and determines the complete subgraph corresponding to the query subgraph based on the matching result. This complete subgraph can accurately represent professional terms and complex relationships in the oral disease field, and then inputs it into a large-scale language model to obtain oral disease diagnosis and treatment answers, thereby improving the accuracy and interpretability of the diagnosis and treatment responses. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1is one of flow schematic diagrams of the oral disease diagnosis and treatment question and answer method provided by the present application.
[0022] Figure 2 is one of flow schematic diagrams of the oral disease diagnosis and treatment question and answer method provided by the present application.
[0023] Figure 3 is one of schematic diagrams of the preset oral disease diagnosis and treatment question bank provided by the present application.
[0024] Figure 4 is one of schematic diagrams of the preset oral disease diagnosis and treatment question bank provided by the present application.
[0025] Figure 5 is one of schematic diagrams of the preset oral disease diagnosis and treatment question bank provided by the present application.
[0026] Figure 6 is a scoring standard schematic diagram provided by the present application.
[0027] Figure 7 is a structural schematic diagram of the oral disease diagnosis and treatment question and answer device provided by the present application.
[0028] Figure 8 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0029] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.
[0030] In view of the above problems in the prior art, the present application provides an oral disease diagnosis and treatment question and answer method, Figure 1 is one of flow schematic diagrams of the oral disease diagnosis and treatment question and answer method provided by the present application, as Figure 1 shown, the method comprises the following steps 110 to 130.
[0031] Step 110: obtaining an oral disease question input by a user, and determining a query subgraph based on the oral disease question; the query subgraph represents the structured oral disease question.
[0032] Specifically, a user-inputted oral disease question can be obtained, which can be, for example, “What can cause gums to always bleed?”. After obtaining the oral disease question, a query subgraph can be determined based on the oral disease question, the query subgraph being organized in the form of an instance graph, and the query subgraph being used for a matching query in the oral medical knowledge graph to obtain a diagnosis and treatment answer. In addition, the query subgraph represents a structured oral disease question.
[0033] In one embodiment, the query subgraph is determined based on the oral disease question, comprising: performing semantic analysis and recognition on the oral disease question to obtain a plurality of oral disease text entities corresponding to the oral disease question, and an entity type and a relationship structure corresponding to each of the oral disease text entities; determining the query subgraph based on each of the oral disease text entities, the entity type and the relationship structure corresponding to each of the oral disease text entities.
[0034] Specifically, the oral disease question can be subjected to semantic analysis and recognition to obtain a plurality of oral disease text entities corresponding to the oral disease question. In the process of extracting text entities, potential oral disease text entities can also be extracted, which do not correspond to keywords but correspond to entity types and relationship structures. In the subsequent process, the potential oral disease text entities can be determined as initial unknown nodes, which can be represented by “?”. The entity type and the relationship structure corresponding to each of the oral disease text entities can also be determined. The entity type is the type of the keyword, which can include, for example, “symptom”, “part”, and “treatment method”, etc. One oral disease text entity can correspond to one entity type or multiple entity types. The relationship structure represents the relationship between each of the oral disease text entities, which can include, for example, a containing relationship or a triggering relationship. For example, the relationship between the text entity “periodontitis” and the text entity “pulp tissue inflammation” is a triggering relationship, i.e., periodontitis triggers pulp tissue inflammation.
[0035] Further, the query subgraph can also be determined based on each of the oral disease text entities, the entity type and the relationship structure corresponding to each of the oral disease text entities. It is easy to understand that each of the oral disease text entities forms a query node in the query subgraph, the entity type can be a description of the query node, and the relationship structure forms an edge between the query nodes in the query subgraph.
[0036] In the above embodiment, the semantic analysis and identification processing is performed on the oral disease problem to obtain the text entity corresponding to the oral disease problem, and the type and relationship structure of the text entity, thereby constructing a query subgraph. The query subgraph can clearly represent each text entity and its relationship, and can help the large language model understand the professional terms and complex relationships in the oral disease field.
[0037] Step 120: matching the query subgraph and the oral medical knowledge graph, and determining the complete subgraph corresponding to the query subgraph according to the matching result; the oral medical knowledge graph is constructed based on multi-source heterogeneous diagnosis and treatment text data in the oral disease field.
[0038] Specifically, the oral medical knowledge graph can be constructed in advance based on multi-source heterogeneous diagnosis and treatment text data in the oral disease field. The multi-source heterogeneous diagnosis and treatment text data may, for example, include multi-source data such as oral medical textbooks, diagnosis and treatment clinical guidelines, and case documents.
[0039] After obtaining the query subgraph, the query subgraph and the oral medical knowledge graph can be matched, and the complete subgraph corresponding to the query subgraph can be determined according to the matching result. It is easy to understand that the complete subgraph can more clearly represent the oral disease problem input by the user.
[0040] In one embodiment, the oral medical knowledge graph is constructed by the following method: Obtaining the multi-source heterogeneous diagnosis and treatment text data in the oral disease field, and performing standardization processing on the multi-source heterogeneous diagnosis and treatment text data to obtain standard diagnosis and treatment text data; Determining a plurality of standard diagnosis and treatment nodes in the oral medical knowledge graph based on the standard diagnosis and treatment text data; Inputting each of the standard diagnosis and treatment nodes into a pre-trained semantic embedding model to obtain a diagnosis and treatment semantic vector corresponding to each of the standard diagnosis and treatment nodes output by the semantic embedding model; Constructing the oral medical knowledge graph based on all the standard diagnosis and treatment nodes and all the diagnosis and treatment semantic vectors.
[0041] Specifically, the multi-source heterogeneous diagnosis and treatment text data in the oral disease field can be obtained. It is easy to understand that since the formats of the various data in the multi-source heterogeneous diagnosis and treatment text data may not be the same, the multi-source heterogeneous diagnosis and treatment text data can be standardized to obtain standard diagnosis and treatment text data.
[0042] Further, a plurality of standard diagnosis and treatment nodes in the dental medical knowledge graph can be determined based on the standard diagnosis and treatment text data. As can be easily understood, a standard diagnosis and treatment node is essentially a semantic block, and the plurality of standard diagnosis and treatment nodes are connected to form a knowledge graph, and the edges between nodes represent the relationship structure between nodes. Each standard diagnosis and treatment node can also be input into a pre-trained semantic embedding model. The semantic embedding model can perform semantic encoding on each standard diagnosis and treatment node to generate a diagnosis and treatment semantic vector corresponding to each standard diagnosis and treatment node. The semantic embedding model can be, for example, a Sentence-BERT (Sentence Embeddings using Siamese BERT-Networks) model. As can be easily understood, the diagnosis and treatment semantic vector and the standard diagnosis and treatment node have a one-to-one mapping relationship. Further, the dental medical knowledge graph can be constructed based on all standard diagnosis and treatment nodes and all diagnosis and treatment semantic vectors. It should be noted that all standard diagnosis and treatment nodes can constitute a dental medical knowledge tree graph and can be stored in a graph database, and the diagnosis and treatment semantic vector can be stored in a vector database.
[0043] In the above embodiment, by standardizing the multi-source heterogeneous diagnosis and treatment text data, the diagnosis and treatment text data from multiple sources is unified in format, and the correspondence between the standard diagnosis and treatment nodes and the diagnosis and treatment semantic vectors facilitates subsequent matching and retrieval of the query subgraph in the dental medical knowledge graph, thereby improving the accuracy of diagnosis and treatment question answering.
[0044] In one embodiment, the standard diagnosis and treatment text data includes textbook text data and example text data. The method further includes determining a plurality of standard diagnosis and treatment nodes in the dental medical knowledge graph based on the standard diagnosis and treatment text data. The method further includes splitting the textbook text data into a plurality of first text blocks, inputting all the first text blocks into a first general large language model to obtain a plurality of first text entities output by the first general large language model, and constructing a tree classification graph based on all the first text entities. The method further includes splitting the example text data into a plurality of second text blocks, inputting all the second text blocks into the first general large language model to obtain a plurality of second text entities output by the first general large language model. The method further includes performing alignment and fusion processing on all the first text entities and all the second text entities to obtain each aligned and fused standard diagnosis and treatment node.
[0045] Specifically, the standard diagnosis and treatment text data includes textbook type text data and example type text data. The textbook type text data may, for example, include oral medicine professional textbooks, standard diagnosis and treatment manuals, licensed physician examination guide books, and oral medicine guideline files, and the like textbook type text data having a chapter structure. The example type text data may, for example, include cases, medical literature, and the like example type text data.
[0046] The textbook type text data can be split into a plurality of first text blocks, and the splitting process can employ a hybrid chunking strategy combining regular expressions and semantic analysis. For example, all textbook type text data can be divided into a plurality of paragraph texts, ensuring the logical independence of each paragraph text, and further utilizing a language model to analyze each paragraph text, and then dividing the paragraph texts with consistent topics into the same first text block. The textbook type text data can also be split according to chapters. After obtaining a plurality of first text blocks, all first text blocks can be input into a first general large language model to obtain a plurality of first text entities output by the first general large language model, and a tree-like classification diagram can be constructed based on all first text entities. The tree-like classification diagram includes a plurality of first text entities, and the plurality of first text entities are classified by a pre-set term (which may, for example, include disease, symptom, treatment, detection, and drug, etc.).
[0047] Further, the example type text data can be split into a plurality of second text blocks, and the splitting process can employ semantic embedding vectors to divide the text blocks according to semantic consistency. This strategy ensures that the generated second text blocks have logical integrity and length adaptability, facilitating subsequent processing. Further, all second text blocks can be input into the first general large language model, and the first general large language model can extract second text entities by category, thereby obtaining a plurality of second text entities output by the first general large language model, which can include entity name, definition, symptom, indication, and contraindication, etc. field content.
[0048] As can be easily understood, there is an "ownership" relationship between the first text entities and the second text entities, so it is necessary to perform alignment and fusion processing on all first text entities and all second text entities. The alignment and fusion processing can include alignment processing and fusion processing. The alignment processing can include multi-level alignment matching in the following order: (1) exact matching: matching the name of the second text entity directly with the first text entity; (2) fuzzy matching: matching using the edit distance or character similarity between the second text entity and the first text entity; (3) semantic vector matching: respectively performing semantic encoding on the first text entity and the second text entity, and matching according to the semantic vectors obtained after semantic encoding. After alignment processing, fusion processing can be performed, which can be based on data source priority and description credibility for fusion, and a unified standard description is generated through a language model, and finally merged into a standard diagnosis and treatment node in the oral medical knowledge graph.
[0049] In the above embodiment, the construction of the tree classification diagram by the first text block lays a foundation for the construction of the oral medical knowledge graph, and the second text entity in the text block is extracted by the general large language model, ensuring the professionalism, format consistency and information integrity of the extraction result. In addition, the first text entity and the second text entity are aligned and fused, so that the first text entity and the second text entity form an attribution relationship, and the subsequently constructed graph structure is more clear and the semantics are more consistent.
[0050] In one embodiment, the matching of the query subgraph and the oral medical knowledge graph, and the determination of the complete subgraph corresponding to the query subgraph according to the matching result, comprises: For each initial unknown node in all query nodes in the query subgraph, input the initial unknown node into the semantic embedding model to obtain a query semantic vector corresponding to the initial unknown node output by the semantic embedding model; determine the first similarity matching degree of the query semantic vector corresponding to the initial unknown node and all the diagnosis and treatment semantic vectors in the oral medical knowledge graph, respectively; determine the diagnosis and treatment semantic vector corresponding to the largest first similarity matching degree greater than the preset matching threshold as the first semantic vector corresponding to the initial unknown node; determine the complete node corresponding to the initial unknown node based on the standard diagnosis and treatment node corresponding to the first semantic vector; For each initial known node in all query nodes in the query subgraph, match the initial known node corresponding keyword with the oral medical knowledge graph, and determine the complete node corresponding to the initial known node based on the standard diagnosis and treatment node in the oral medical knowledge graph matched; determine the complete subgraph corresponding to the query subgraph based on all the complete nodes.
[0051] Specifically, the query subgraph includes a plurality of query nodes, and the query nodes are divided into initial unknown nodes and initial known nodes. The initial unknown node can be a query node corresponding to a potential (or indefinite) oral disease text entity, which can be represented by “?” in the query subgraph. For each initial unknown node in all query nodes in each query subgraph, the first similarity matching degree of the query semantic vector corresponding to the initial unknown node and all the diagnosis and treatment semantic vectors in the oral medical knowledge graph can be calculated. The calculation of the first similarity matching degree can adopt various ways, which are not specifically limited in the embodiments of the present application, for example, a fuzzy matching algorithm can be used.
[0052] Further, the first similarity matching degrees corresponding to the initial unknown nodes can be compared with the preset matching threshold respectively, in the case that the first similarity matching degrees are greater than the preset matching threshold, the first similarity matching degrees greater than the preset matching threshold are compared, and the diagnosis and treatment semantic vector corresponding to the maximum first similarity matching degree is determined as the first semantic vector corresponding to the initial unknown node, and then the complete node corresponding to the initial unknown node can be determined based on the standard diagnosis and treatment node corresponding to the first semantic vector, for example, the text content corresponding to the initial unknown node can be replaced by the text content corresponding to the standard diagnosis and treatment node, so as to obtain the complete node corresponding to the initial unknown node. The preset matching threshold can be set as needed, for example, it can be 80%, and the present embodiment does not make specific limitation here.
[0053] For each initial known node in all query nodes in the query subgraph, the keyword corresponding to the initial known node can be matched with the dental medical knowledge graph, and then the complete node corresponding to the initial known node is determined based on the standard diagnosis and treatment node in the dental medical knowledge graph matched.
[0054] After obtaining the complete nodes corresponding to all query nodes, all complete nodes can be connected to obtain a complete subgraph corresponding to the query subgraph. The complete subgraph covers entities, attributes, semantic paths and expected reasoning targets involved in the user's question, and for example, the triple relationship of "disease-symptom", "disease-treatment" and "symptom-sign" can be represented in the complete subgraph.
[0055] In the above embodiment, the query subgraph is expanded to a complete subgraph capable of more completely expressing disease diagnosis and treatment information through similarity matching between the query subgraph and the dental medical knowledge graph, so that the subsequent large language model can better understand the user's question.
[0056] In one embodiment, the method further comprises: For each initial known node, in the case that the keyword corresponding to the initial known node is not matched to the standard diagnosis and treatment node in the dental medical knowledge graph, the initial known node is input to the semantic embedding model to obtain the query semantic vector of the initial known node output by the semantic embedding model; the second similarity matching degrees of the query semantic vector corresponding to the initial known node and all diagnosis and treatment semantic vectors in the dental medical knowledge graph are determined; the diagnosis and treatment semantic vector corresponding to the maximum second similarity matching degree greater than the preset matching threshold is determined as the second semantic vector corresponding to the initial known node; and the complete node corresponding to the initial known node is determined based on the standard diagnosis and treatment node corresponding to the second semantic vector.
[0057] Specifically, for each initial known node, in a case that the keyword corresponding to the initial known node is not matched to the standard diagnosis and treatment node in the oral medical knowledge graph, the initial known node can also be input to the semantic embedding model to obtain the query semantic vector of the initial known node output by the semantic embedding model.
[0058] Further, the second similarity matching degrees of the initial known node can be determined respectively with all diagnosis and treatment semantic vectors in the oral medical knowledge graph, the calculation of the second similarity matching degrees can adopt various ways, the present embodiment does not make specific limitation here, for example, a fuzzy matching algorithm can be adopted. Further, the second similarity matching degrees of the initial known node can be compared respectively with the preset matching threshold, in a case that there is a second similarity matching degree greater than the preset matching threshold, the second similarity matching degrees greater than the preset matching threshold are compared, the diagnosis and treatment semantic vector corresponding to the maximum second similarity matching degree is determined as the second semantic vector of the initial known node, and then the complete node corresponding to the initial known node can be determined based on the standard diagnosis and treatment node corresponding to the second semantic vector, for example, the text content corresponding to the initial known node can be replaced with the text content corresponding to the standard diagnosis and treatment node, so as to obtain the complete node corresponding to the initial known node.
[0059] In the above embodiment, in a case that the keyword corresponding to the initial known node is difficult to be matched to the keyword in the oral medical knowledge graph, the initial known node can be replaced, so as to ensure that the complete node of the initial known node can be determined.
[0060] In one embodiment, the method further comprises: In a case that all the first similarity matching degrees corresponding to the initial unknown node are less than the preset matching threshold, the initial unknown node is determined as a target unknown node; In a case that the keyword corresponding to the initial known node is not matched to the standard diagnosis and treatment node in the oral medical knowledge graph, and all the second similarity matching degrees corresponding to the initial known node are less than the preset matching threshold, the initial known node is determined as a target unknown node; For each target unknown node, an adjacent query node corresponding to the target unknown node is determined based on the relationship structure corresponding to the target unknown node; A complete node corresponding to the target unknown node is determined based on the standard diagnosis and treatment node corresponding to the adjacent query node.
[0061] Specifically, for each initial unknown node, if all the first similarity matching degrees corresponding to the initial unknown node are less than the preset matching threshold, the initial unknown node can be determined as a target unknown node.
[0062] For each initial known node, if the keyword corresponding to the initial known node is not matched to the standard diagnosis and treatment node in the dental medical knowledge graph (i.e., the keyword is not queried in the dental medical knowledge graph), and all the second similarity matching degrees corresponding to the initial known node are less than the preset matching threshold, the initial known node can be determined as a target unknown node. It is easy to understand that the target unknown node can also be represented by "?" in the query subgraph, and the target unknown node needs to be completed through the following steps.
[0063] For each target unknown node, the adjacent query node corresponding to the target unknown node can be determined based on the relationship structure corresponding to the target unknown node, i.e., the adjacent query node connected to the target unknown node is determined according to the edge of the target unknown node in the query subgraph. It is easy to understand that the adjacent query node can be one or more. The complete node corresponding to the target unknown node can be further determined based on the relationship structure of the standard diagnosis and treatment node corresponding to the adjacent query node.
[0064] For example, query node A is determined as a target unknown node, and query node B is further determined as an adjacent query node of query node A. Query node B corresponds to standard diagnosis and treatment node b, and standard diagnosis and treatment node b is connected to standard diagnosis and treatment node e. Since query node A is connected to query node B, query node B corresponds to standard diagnosis and treatment node b, and standard diagnosis and treatment node b is connected to standard diagnosis and treatment node e, it can be considered that query node A is more approximate to standard diagnosis and treatment node e, and thus the complete node corresponding to query node A can be determined according to standard diagnosis and treatment node e.
[0065] In the above embodiment, for each target unknown node, the complete node corresponding to the target unknown node can be determined through the relationship structure of the standard diagnosis and treatment node corresponding to the adjacent query node, thereby completing the information of the target unknown node, further making the complete subgraph more accurate, and making the complete subgraph corresponding to the query subgraph closer to the real semantic.
[0066] Step 130: inputting the complete subgraph and the dental disease problem into the diagnosis and treatment large language model to obtain the dental disease diagnosis and treatment answer corresponding to the dental disease problem output by the diagnosis and treatment large language model; the diagnosis and treatment large language model is obtained by training a general large language model based on dental disease diagnosis and treatment sample data.
[0067] Specifically, the general large language model can be fine-tuned and trained in advance based on the oral disease diagnosis and treatment sample data to obtain a diagnosis and treatment large language model. After obtaining the complete subgraph, the complete subgraph and the oral disease question can be input into the diagnosis and treatment large language model to obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model. Inputting the complete subgraph and the oral disease question into the diagnosis and treatment large language model can make the diagnosis and treatment large language model respond to the context and output more accurate answers.
[0068] In one embodiment, the inputting the complete subgraph and the oral disease question into the diagnosis and treatment large language model to obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model comprises: determining an initial diagnosis and treatment prompt word based on the node definition, the node relationship structure and the node attribute corresponding to the complete subgraph; determining a target diagnosis and treatment prompt word based on the initial diagnosis and treatment prompt word and the oral disease question; inputting the target diagnosis and treatment prompt word into the diagnosis and treatment large language model to obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model.
[0069] Specifically, an initial diagnosis and treatment prompt word can be determined based on the node definition, the node relationship structure (which may, for example, include a triple relationship) and the node attribute (which can be description information of each node in the complete subgraph) corresponding to the complete subgraph. Then, the initial diagnosis and treatment prompt word and the oral disease question can be fused to obtain a target diagnosis and treatment prompt word. Finally, the target diagnosis and treatment prompt word can be input into the diagnosis and treatment large language model to obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model.
[0070] In the above embodiment, the target diagnosis and treatment prompt word is determined based on the complete subgraph and the oral disease question, and then input into the large language model, which ensures that the oral disease diagnosis and treatment answer output by the diagnosis and treatment large language model has high explainability and traceability.
[0071] For example, Figure 2 is a flowchart of the oral disease diagnosis and treatment question and answer method provided by the present application, as Figure 2As shown, the user input oral disease question is: "In the treatment of a periodontitis-induced pulp tissue inflammation, is the disinfectant drug used also used in a preventive measure for children's deep dental sulcus? If so, describe the specific use of the drug in both cases and its mechanism of action", based on the oral disease question, the query subgraph determined is "periodontitis"→ "? oral disease"→ "? drug"← "? prevention", where "?" represents an unknown node, and then matched with the oral medical knowledge graph, the complete subgraph obtained is "periodontitis"→ "retrograde pulpitis"→ "disinfectant drug"← "root canal treatment", after inputting the complete subgraph and the oral disease question into the diagnosis and treatment large language model, the obtained oral disease diagnosis and treatment answer can be in the form of pure text, or in the form of a tree diagram as shown. Figure 2 As shown in the tree diagram.
[0072] The oral disease diagnosis and treatment question and answer method provided by the application acquires a user input oral disease question, and determines a query subgraph based on the oral disease question; the query subgraph represents a structured oral disease question; the query subgraph and the oral medical knowledge graph are matched, and a complete subgraph corresponding to the query subgraph is determined according to the matching result; the oral medical knowledge graph is constructed based on multi-source heterogeneous diagnosis and treatment text data in the oral disease field; the complete subgraph and the oral disease question are input into a diagnosis and treatment large language model, and an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model is obtained; the diagnosis and treatment large language model is obtained by training a general large language model based on oral disease diagnosis and treatment sample data. The technical scheme of the application determines a query subgraph based on a user input oral disease question, matches the query subgraph with a pre-constructed oral medical knowledge graph, determines a complete subgraph corresponding to the query subgraph according to the matching result, the complete subgraph can accurately represent professional terms and complex relationships in the oral disease field, and then inputting the large language model obtains an oral disease diagnosis and treatment answer, which can improve the accuracy and explainability of the diagnosis and treatment answer.
[0073] In one embodiment, the method further comprises: inputting a preset oral disease diagnosis and treatment question bank into the diagnosis and treatment large language model to obtain a diagnosis and treatment prediction answer corresponding to the preset oral disease diagnosis and treatment question bank output by the diagnosis and treatment large language model; inputting a preset evaluation prompt, the preset oral disease diagnosis and treatment question bank, the diagnosis and treatment prediction answer, and a diagnosis and treatment standard answer corresponding to the preset oral disease diagnosis and treatment question bank into a second general large language model to obtain a diagnosis and treatment evaluation result of the diagnosis and treatment large language model output by the second general large language model.
[0074] Specifically, the preset oral disease diagnosis and treatment question bank can be input into the diagnosis and treatment large language model to obtain a diagnosis and treatment prediction answer corresponding to the preset oral disease diagnosis and treatment question bank output by the diagnosis and treatment large language model. The preset oral disease diagnosis and treatment question bank can include subjective questions and objective questions. Figure 3 is one of the schematic diagrams of the preset oral disease diagnosis and treatment question bank provided by the present application. The objective questions in the preset oral disease diagnosis and treatment question bank can be as shown in Figure 3 The objective questions in the preset oral disease diagnosis and treatment question bank can include multiple module categories, each module category can include multiple subfields, and each subfield can include multiple specific contents / disciplines. The subjective questions can include basic diagnosis and treatment technical ability assessment questions and clinical reasoning ability assessment questions. Figure 4 is the second schematic diagram of the preset oral disease diagnosis and treatment question bank provided by the present application. The basic diagnosis and treatment technical ability assessment questions can be as shown in Figure 4 Figure 5 is the third schematic diagram of the preset oral disease diagnosis and treatment question bank provided by the present application. The clinical reasoning ability assessment questions can be as shown in Figure 5
[0075] Further, the preset evaluation prompt words, the preset oral disease diagnosis and treatment question bank, the diagnosis and treatment prediction answer, and the diagnosis and treatment standard answer corresponding to the preset oral disease diagnosis and treatment question bank can be input into the second general large language model to obtain a diagnosis and treatment evaluation result of the diagnosis and treatment large language model output by the second general large language model. The preset evaluation prompt words can be used to prompt the second general large language model to score the diagnosis and treatment prediction answer output by the diagnosis and treatment large language model, Figure 6 is a scoring standard schematic diagram provided by the present application, as shown in Figure 6 The second general large language model can evaluate the diagnosis and treatment prediction answer from five aspects of correctness, integrity, safety, understanding, and logic to obtain the diagnosis and treatment evaluation result of the diagnosis and treatment large language model.
[0076] In the prior art, when evaluating the effect of a medical question and answer system or an auxiliary diagnosis and treatment tool, there is a lack of strict evaluation system. The traditional evaluation method often only focuses on the accuracy of the answer and fails to comprehensively investigate the dimensions of clinical reasoning process, diagnosis and treatment skill application, and expression standardization, which is not conducive to discovering the deficiencies of the diagnosis and treatment large language model in real clinical application. The use of the above embodiments makes the performance evaluation of the diagnosis and treatment large language model have a basis and be fully quantified, so that the short boards of the diagnosis and treatment large language model in clinical reasoning, skill application, etc. can be found for continuous optimization.
[0077] The oral disease diagnosis and treatment question and answer device provided by the present application is described below. The oral disease diagnosis and treatment question and answer device described below can be correspondingly referred to the oral disease diagnosis and treatment question and answer method described above.
[0078] Figure 7 is a structural schematic diagram of an oral disease diagnosis and treatment question and answer device provided by the present application, as shown in Figure 7 The oral disease diagnosis and treatment question and answer device 700 comprises the following modules: A question module 710 is configured to acquire an oral disease question input by a user, and determine a query subgraph based on the oral disease question; the query subgraph represents a structured oral disease question; A matching module 720 is configured to match the query subgraph and an oral medical knowledge graph, and determine a complete subgraph corresponding to the query subgraph according to a matching result; the oral medical knowledge graph is constructed based on multi-source and heterogeneous diagnosis and treatment text data in the oral disease field; An answer module 730 is configured to input the complete subgraph and the oral disease question into a diagnosis and treatment large language model, and obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model; the diagnosis and treatment large language model is obtained by training a general large language model based on oral disease diagnosis and treatment sample data.
[0079] In one embodiment, the question module 710 is specifically configured to: perform semantic analysis and recognition processing on the oral disease question, and obtain a plurality of oral disease text entities corresponding to the oral disease question, and an entity type and a relationship structure corresponding to each oral disease text entity; determine the query subgraph based on each oral disease text entity, the entity type and the relationship structure corresponding to each oral disease text entity.
[0080] In one embodiment, the oral disease diagnosis and treatment question and answer device further comprises a construction module, which is specifically configured to: acquire the multi-source and heterogeneous diagnosis and treatment text data in the oral disease field, and perform standardization processing on the multi-source and heterogeneous diagnosis and treatment text data to obtain standard diagnosis and treatment text data; determine a plurality of standard diagnosis and treatment nodes in the oral medical knowledge graph based on the standard diagnosis and treatment text data; input each standard diagnosis and treatment node into a pre-trained semantic embedding model to obtain a diagnosis and treatment semantic vector corresponding to each standard diagnosis and treatment node output by the semantic embedding model; construct the oral medical knowledge graph based on all the standard diagnosis and treatment nodes and all the diagnosis and treatment semantic vectors.
[0081] In one embodiment, the standard diagnosis and treatment text data comprises textbook text data and example text data; the construction module is specifically further configured to: The textbook type text data is split into a plurality of first text blocks, and all the first text blocks are input into a first general large language model to obtain a plurality of first text entities output by the first general large language model, and a tree classification diagram is constructed based on all the first text entities; The instance type text data is split into a plurality of second text blocks, and all the second text blocks are input into the first general large language model to obtain a plurality of second text entities output by the first general large language model; The first text entities and the second text entities are aligned and fused to obtain each standard diagnosis and treatment node after alignment and fusion.
[0082] In one embodiment, the matching module 720 is specifically configured to: For each initial unknown node in all query nodes in the query subgraph, the initial unknown node is input into the semantic embedding model to obtain a query semantic vector corresponding to the initial unknown node output by the semantic embedding model; Determine the first similarity matching degree of the query semantic vector corresponding to the initial unknown node and all diagnosis and treatment semantic vectors in the oral medical knowledge graph, respectively; Determine the diagnosis and treatment semantic vector corresponding to the largest first similarity matching degree greater than the preset matching threshold as the first semantic vector corresponding to the initial unknown node; determine the complete node corresponding to the initial unknown node based on the standard diagnosis and treatment node corresponding to the first semantic vector; For each initial known node in all query nodes in the query subgraph, match the initial known node corresponding keyword with the oral medical knowledge graph, and determine the complete node corresponding to the initial known node based on the standard diagnosis and treatment node in the oral medical knowledge graph matched; Determine the complete subgraph corresponding to the query subgraph based on all the complete nodes.
[0083] In one embodiment, the oral disease diagnosis and treatment question answering device further includes a known completion module, which is specifically configured to: For each of the initial known nodes, in a case where the keywords corresponding to the initial known nodes are not matched to the standard diagnosis and treatment nodes in the oral medical knowledge graph, inputting the initial known nodes to the semantic embedding model to obtain query semantic vectors of the initial known nodes output by the semantic embedding model; determining second similarity matching degrees of the query semantic vectors of the initial known nodes with all the diagnosis and treatment semantic vectors in the oral medical knowledge graph respectively; determining a diagnosis and treatment semantic vector corresponding to a maximum second similarity matching degree greater than the preset matching threshold as a second semantic vector corresponding to the initial known node; and determining a complete node corresponding to the initial known node based on a standard diagnosis and treatment node corresponding to the second semantic vector.
[0084] In one embodiment, the oral disease diagnosis and treatment question and answer device further comprises an unknown completion module, which is specifically configured to: For each of the initial unknown nodes, in a case where all the first similarity matching degrees corresponding to the initial unknown nodes are less than the preset matching threshold, determining the initial unknown node as a target unknown node; For each of the initial known nodes, in a case where the keywords corresponding to the initial known nodes are not matched to the standard diagnosis and treatment nodes in the oral medical knowledge graph, and all the second similarity matching degrees corresponding to the initial known nodes are less than the preset matching threshold, determining the initial known node as a target unknown node; For each of the target unknown nodes, determining an adjacent query node corresponding to the target unknown node based on a relationship structure corresponding to the target unknown node; Determining a complete node corresponding to the target unknown node based on a standard diagnosis and treatment node corresponding to the adjacent query node.
[0085] In one embodiment, the answer module 730 is specifically configured to: Determining an initial diagnosis and treatment prompt word based on the node definition, the node relationship structure and the node attribute corresponding to the complete subgraph; Determining a target diagnosis and treatment prompt word based on the initial diagnosis and treatment prompt word and the oral disease question; Inputting the target diagnosis and treatment prompt word to the diagnosis and treatment large language model to obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model.
[0086] In one embodiment, the oral disease diagnosis and treatment question and answer device further comprises an evaluation module, which is specifically configured to: Inputting a preset oral disease diagnosis and treatment question bank to the diagnosis and treatment large language model to obtain diagnosis and treatment prediction answers corresponding to the preset oral disease diagnosis and treatment question bank output by the diagnosis and treatment large language model; The preset evaluation prompt word, the preset oral disease diagnosis and treatment question bank, the diagnosis and treatment prediction answer, and the diagnosis and treatment standard answer corresponding to the preset oral disease diagnosis and treatment question bank are input into a second general large language model, and a diagnosis and treatment evaluation result of the diagnosis and treatment large language model output by the second general large language model is obtained.
[0087] The oral disease diagnosis and treatment question and answer device provided by the application obtains an oral disease question input by a user, and determines a query subgraph based on the oral disease question; the query subgraph represents a structured oral disease question; the query subgraph and an oral medical knowledge graph are matched, and a complete subgraph corresponding to the query subgraph is determined according to a matching result; the oral medical knowledge graph is constructed based on multi-source and heterogeneous diagnosis and treatment text data in the oral disease field; the complete subgraph and the oral disease question are input into a diagnosis and treatment large language model, and an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model is obtained; the diagnosis and treatment large language model is obtained by training a general large language model based on oral disease diagnosis and treatment sample data. The technical scheme of the application determines a query subgraph based on an oral disease question input by a user, matches the query subgraph and a pre-constructed oral medical knowledge graph, and determines a complete subgraph corresponding to the query subgraph according to a matching result. The complete subgraph can accurately represent professional terms and complex relationships in the oral disease field, and then input into a large language model to obtain an oral disease diagnosis and treatment answer, which can improve the accuracy and explainability of diagnosis and treatment answers.
[0088] Figure 8 An example of an entity structure diagram of an electronic device is shown in FIG. 8A. Figure 8 As shown in FIG. 8A, the electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communications bus 840. The processor 810 can invoke a logical instruction in the memory 830 to execute an oral disease diagnosis and treatment question and answer method, which includes: obtaining an oral disease question input by a user, and determining a query subgraph based on the oral disease question; the query subgraph represents a structured oral disease question; matching the query subgraph and an oral medical knowledge graph, and determining a complete subgraph corresponding to the query subgraph according to a matching result; the oral medical knowledge graph is constructed based on multi-source and heterogeneous diagnosis and treatment text data in the oral disease field; inputting the complete subgraph and the oral disease question into a diagnosis and treatment large language model, and obtaining an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model; the diagnosis and treatment large language model is obtained by training a general large language model based on oral disease diagnosis and treatment sample data.
[0089] Further, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as standalone products, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or partially contribute to the prior art, or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0090] On the other hand, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the oral disease diagnosis and treatment question and answer method provided by the above-mentioned methods, which comprises: obtaining an oral disease question input by a user, and determining a query subgraph based on the oral disease question; the query subgraph represents a structured oral disease question; matching the query subgraph and an oral medical knowledge graph, and determining a complete subgraph corresponding to the query subgraph according to a matching result; the oral medical knowledge graph is constructed based on multi-source and heterogeneous diagnosis and treatment text data in the field of oral diseases; inputting the complete subgraph and the oral disease question into a diagnosis and treatment large language model to obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model; the diagnosis and treatment large language model is obtained by training a general large language model based on oral disease diagnosis and treatment sample data.
[0091] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement an oral disease diagnosis and treatment question and answer method provided by the above-mentioned methods, which comprises: obtaining an oral disease question input by a user, and determining a query subgraph based on the oral disease question; the query subgraph represents a structured oral disease question; matching the query subgraph and an oral medical knowledge graph, and determining a complete subgraph corresponding to the query subgraph according to a matching result; the oral medical knowledge graph is constructed based on multi-source and heterogeneous diagnosis and treatment text data in the field of oral diseases; inputting the complete subgraph and the oral disease question into the diagnosis and treatment large language model to obtain an oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnosis and treatment large language model; the diagnosis and treatment large language model is obtained by training a general large language model based on oral disease diagnosis and treatment sample data.
[0092] The device embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and the necessary general hardware platform, and of course, it can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0094] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A question-and-answer method for diagnosing oral diseases, characterized in that, include: Obtain the oral disease questions input by the user, and determine the query subgraph based on the oral disease questions; The query subgraph represents the structured oral disease problem; The query subgraph and the oral healthcare knowledge graph are matched, and the complete subgraph corresponding to the query subgraph is determined based on the matching result; the oral healthcare knowledge graph is constructed based on multi-source heterogeneous diagnostic and treatment text data in the field of oral diseases. The complete subgraph and the oral disease question are input into the diagnostic language model to obtain the oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnostic language model; the diagnostic language model is obtained by training a general language model based on oral disease diagnosis and treatment sample data.
2. The oral disease diagnosis and treatment question-and-answer method according to claim 1, characterized in that, The process of determining the query subgraph based on the oral disease problem includes: The oral disease problem is subjected to semantic parsing and recognition processing to obtain multiple oral disease text entities corresponding to the oral disease problem, as well as the entity type and relation structure corresponding to each oral disease text entity; The query subgraph is determined based on each of the oral disease text entities, the entity type corresponding to each of the oral disease text entities, and the relational structure.
3. The oral disease diagnosis and treatment question-and-answer method according to claim 2, characterized in that, The oral healthcare knowledge graph is constructed in the following ways: Acquire the multi-source heterogeneous diagnostic and treatment text data in the field of oral diseases, and perform standardization processing on the multi-source heterogeneous diagnostic and treatment text data to obtain standard diagnostic and treatment text data; Based on the standard diagnostic and treatment text data, multiple standard diagnostic and treatment nodes in the oral medical knowledge graph are determined; Each of the standard diagnostic and treatment nodes is input into a pre-trained semantic embedding model to obtain the diagnostic and treatment semantic vectors corresponding to each of the standard diagnostic and treatment nodes output by the semantic embedding model. The oral healthcare knowledge graph is constructed based on all the aforementioned standard diagnostic and treatment nodes and all the aforementioned diagnostic and treatment semantic vectors.
4. The oral disease diagnosis and treatment question-and-answer method according to claim 3, characterized in that, The standard diagnostic and treatment text data includes textbook-type text data and example-type text data; The process of determining multiple standard diagnostic and treatment nodes in the oral healthcare knowledge graph based on the standard diagnostic and treatment text data includes: The textbook text data is split into multiple first text blocks, and all first text blocks are input into a first general large language model to obtain multiple first text entities output by the first general large language model. A tree classification diagram is then constructed based on all first text entities. The instance class text data is split into multiple second text blocks, and all the second text blocks are input into the first general large language model to obtain multiple second text entities output by the first general large language model; All the first text entities and all the second text entities are aligned and fused to obtain the standard diagnostic nodes after alignment and fusion.
5. The oral disease diagnosis and treatment question-and-answer method according to claim 4, characterized in that, The process of matching the query subgraph with the oral healthcare knowledge graph, and determining the complete subgraph corresponding to the query subgraph based on the matching results, includes: For each initially unknown node in all query nodes in the query subgraph, the initially unknown node is input into the semantic embedding model to obtain the query semantic vector corresponding to the initially unknown node output by the semantic embedding model; Determine the first similarity matching degree between the query semantic vector corresponding to the initial unknown node and all the diagnosis and treatment semantic vectors in the oral medical knowledge graph; The diagnostic semantic vector corresponding to the largest first similarity matching degree among the first similarity matching degrees that are greater than the preset matching threshold is determined as the first semantic vector corresponding to the initial unknown node; the complete node corresponding to the initial unknown node is determined based on the standard diagnostic node corresponding to the first semantic vector. For each initially known node in all query nodes in the query subgraph, the keywords corresponding to the initially known node are matched with the oral medical knowledge graph, and the complete node corresponding to the initially known node is determined based on the standard treatment nodes in the matched oral medical knowledge graph. The complete subgraph corresponding to the query subgraph is determined based on all the complete nodes.
6. The oral disease diagnosis and treatment question-and-answer method according to claim 5, characterized in that, The method further includes: For each of the initial known nodes, if the keyword corresponding to the initial known node does not match a standard treatment node in the oral healthcare knowledge graph, the initial known node is input into the semantic embedding model to obtain the query semantic vector corresponding to the initial known node output by the semantic embedding model; the second similarity matching degree between the query semantic vector corresponding to the initial known node and all the treatment semantic vectors in the oral healthcare knowledge graph is determined; the treatment semantic vector corresponding to the largest second similarity matching degree among the second similarity matching degrees greater than the preset matching threshold is determined as the second semantic vector corresponding to the initial known node; the complete node corresponding to the initial known node is determined based on the standard treatment node corresponding to the second semantic vector.
7. The oral disease diagnosis and treatment question-and-answer method according to claim 6, characterized in that, The method further includes: For each of the initial unknown nodes, if all the first similarity matching degrees corresponding to the initial unknown node are less than the preset matching threshold, the initial unknown node is determined as the target unknown node; For each of the initial known nodes, if the keyword corresponding to the initial known node does not match the standard treatment node in the oral medical knowledge graph, and all the second similarity matching degrees corresponding to the initial known node are less than the preset matching threshold, the initial known node is determined as the target unknown node. For each of the aforementioned unknown target nodes, the adjacent query node corresponding to the unknown target node is determined based on the relational structure corresponding to the unknown target node; The complete node corresponding to the target unknown node is determined based on the standard diagnosis node corresponding to the adjacent query node.
8. The oral disease diagnosis and treatment question-and-answer method according to any one of claims 1 to 6, characterized in that, The process of inputting the complete subgraph and the oral disease question into the diagnostic language model to obtain the oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnostic language model includes: Initial diagnostic prompts are determined based on the node definitions, node relationship structures, and node attributes corresponding to the complete subgraph. Target treatment prompts are determined based on the initial treatment prompts and the oral disease problem; The target diagnosis and treatment prompts are input into the diagnosis and treatment language model to obtain the oral disease diagnosis and treatment answers corresponding to the oral disease questions output by the diagnosis and treatment language model.
9. The oral disease diagnosis and treatment question-and-answer method according to any one of claims 1 to 6, characterized in that, The method further includes: Input the preset oral disease diagnosis and treatment question bank into the diagnosis and treatment language model to obtain the diagnosis and treatment prediction answers corresponding to the preset oral disease diagnosis and treatment question bank output by the diagnosis and treatment language model; The preset evaluation prompts, the preset oral disease diagnosis and treatment question bank, the diagnosis and treatment prediction answers, and the corresponding standard answers of the preset oral disease diagnosis and treatment question bank are input into the second general language model to obtain the diagnosis and treatment evaluation results of the diagnosis and treatment language model output by the second general language model.
10. A diagnostic and question-and-answer device for oral diseases, characterized in that, include: The question module is used to obtain oral disease questions input by the user and determine a query subgraph based on the oral disease questions; The query subgraph represents the structured oral disease problem; The matching module is used to match the query subgraph with the oral medical knowledge graph, and determine the complete subgraph corresponding to the query subgraph based on the matching result; the oral medical knowledge graph is constructed based on multi-source heterogeneous diagnosis and treatment text data in the field of oral diseases; The answer module is used to input the complete subgraph and the oral disease question into the diagnostic language model to obtain the oral disease diagnosis and treatment answer corresponding to the oral disease question output by the diagnostic language model; the diagnostic language model is obtained by training a general language model based on oral disease diagnosis and treatment sample data.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the oral disease diagnosis and treatment question-and-answer method as described in any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the oral disease diagnosis and treatment question-and-answer method as described in any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the oral disease diagnosis and treatment question-and-answer method as described in any one of claims 1 to 9.