Medical education-oriented knowledge graph construction, retrieval and application system

By integrating medical knowledge graphs, large language models and graph neural network models, the problems of low efficiency and lack of deep intelligent interaction in traditional Chinese medicine knowledge processing in the existing technology are solved, efficient integration and deep understanding of medical knowledge are achieved, and the level of intelligence in medical education and research is improved.

CN120012902APending Publication Date: 2025-05-16SICHUAN UNIV +1
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510067873.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology has problems such as inefficiency in the processing and application of medical knowledge, incomplete knowledge graph and lack of deep intelligent interaction capabilities, which is difficult to meet the needs of complex medical knowledge application scenarios.

Method used

By deeply integrating medical knowledge graphs, large language models and graph neural network models, we build a knowledge graph construction, search and application system for medical education, and use graph neural network models to learn and understand knowledge graphs, provide structured knowledge support for large language models, and generate detailed explanation content and case analysis through the knowledge retrieval and content generation module.

Benefits of technology

It has achieved efficient integration and in-depth understanding of medical knowledge, improved the efficiency and quality of the dissemination and application of medical knowledge, and promoted the intelligent development of medical education and research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012902A_ABST
    Figure CN120012902A_ABST
Patent Text Reader

Abstract

The invention discloses a medical education-oriented knowledge graph construction, retrieval and application system, and the system comprises a data collection and integration module which is used for collecting and integrating various medical professional knowledge, and forming a medical knowledge material library; the intelligent algorithm module is used for constructing a knowledge graph according to the medical knowledge material library and then providing structured knowledge support for learning and reasoning of the knowledge graph by using a graph neural network model; the knowledge retrieval and content generation module is used for performing retrieval in the knowledge graph according to the input content and providing the retrieved knowledge to the large language model to obtain the corresponding content; and the knowledge application service module is used for providing medical education service according to learning and understanding of the graph neural network model on the knowledge graph and integration of the big language model on semantic information in the knowledge graph. The intelligent development of the medical field is better promoted, and more efficient and accurate support and service are provided for medical education, research and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application of knowledge graphs, large language models and graph neural networks in the medical field, and in particular to a knowledge graph construction, retrieval and application system for medical education. Background Art

[0002] Knowledge graph is a semantic network that reveals the relationship between entities and has a wide range of applications in many fields. In terms of medical knowledge management, existing technologies integrate medical knowledge resources by constructing medical knowledge graphs, attempting to achieve rapid retrieval and intelligent application of knowledge. For example, some medical knowledge base systems use knowledge graphs to relate medical concepts such as diseases, symptoms, and treatment methods, and users can query relevant medical knowledge by entering keywords. However, these traditional medical knowledge graph applications often have limitations. Most of their construction processes rely on manual annotation or semi-automatic methods, which have low processing efficiency for large-scale medical knowledge and make it difficult to ensure the integrity and timeliness of knowledge graphs. Moreover, at the knowledge application level, they mainly focus on simple knowledge query and retrieval, lack deep intelligent interaction capabilities, and cannot meet the needs of complex medical knowledge application scenarios, such as intelligent question and answer and interactive feedback in simulated medical teaching.

[0003] Large language models such as GPT have made significant progress in the field of natural language processing. By learning from massive amounts of text data, they can generate coherent and logical text answers, and perform well in general-purpose intelligent question-answering, text creation, and other aspects. However, when applied in the medical field, there are many problems with directly using large language models. Although the general data used in the training of large language models covers some medical knowledge, it does not cover the specific knowledge system in medical professional textbooks. This leads to inaccurate, shallow, or even wrong answers when facing medical professional questions. For example, when answering delicate questions about the physiological mechanisms of the human body in medicine or the pathogenesis of specific diseases at the medical theoretical level, it is difficult for large language models to give accurate answers that meet the requirements of teaching and learning based on professional medical textbook knowledge.

[0004] Graph neural networks have unique advantages in processing graph structured data and have been applied to some knowledge graph-related research to mine the potential features of entities and relationships in knowledge graphs. However, in the application of combining medical knowledge graphs with intelligent question and answer, existing technologies have failed to fully tap the potential of graph neural networks. In most cases, graph neural networks are only used for simple feature extraction of knowledge graph structures and are not effectively deeply integrated with large language models. This isolated application method makes it impossible to fully utilize the graph neural network's understanding of knowledge graph relationships and the powerful language generation capabilities of large language models in the intelligent question and answer process, resulting in the intelligent question and answer system's reasoning ability and answer quality when processing medical knowledge. It is difficult to meet the needs of medical education and research for accurate, comprehensive, and in-depth knowledge interaction.

[0005] Retrieval-augmented generation (RAG) technology is a method that emerged to improve the shortcomings of large language models in the application of knowledge in specific fields. It retrieves relevant information from external knowledge bases (such as knowledge graphs, etc.) before the large language model generates an answer, and then provides this information as additional input to the large language model to enhance the accuracy and pertinence of its answers. There have been some attempts based on RAG technology in the medical field. For example, some medical auxiliary question-answering systems use RAG to retrieve relevant information from medical literature databases to assist large language models in answering questions. However, the current application of RAG technology in medicine still faces challenges. On the one hand, the accuracy and efficiency of the retrieval process need to be improved. The existing retrieval algorithms may not be able to quickly and accurately locate the most critical information from the massive amount of medical knowledge, resulting in the introduction of information that may not be optimal, affecting the quality of the final answer. On the other hand, the integration of RAG technology and graph neural networks is not close enough, and the graph neural network's ability to mine knowledge graph structure information has not been fully utilized to optimize the retrieval process and information integration, resulting in shortcomings in the processing of complex medical knowledge relationships, making it difficult to meet the diverse and high-precision intelligent question-answering and knowledge application service needs in medical education and research.

[0006] In summary, the existing knowledge graph construction and application technology, the application of large language models alone, and graph neural networks all have shortcomings in medical intelligent question answering and knowledge application services. The present invention aims to overcome these shortcomings and provide a more efficient, accurate and feature-rich intelligent knowledge application solution based on medical knowledge graphs and multi-model fusion. Summary of the invention

[0007] In view of the above-mentioned deficiencies in the prior art, the present invention provides a knowledge graph construction, retrieval and application system for medical education, which solves the problem that the prior art cannot accurately and deeply answer medical professional questions and is inefficient.

[0008] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0009] A knowledge graph construction, retrieval and application system for medical education, including:

[0010] Data collection and integration module, used to collect and integrate various medical expertise to form a medical knowledge material library;

[0011] Intelligent algorithm module, used to build knowledge graph based on medical knowledge material library; using graph neural network model to learn and understand knowledge graph and provide structured knowledge support for large language model;

[0012] The knowledge retrieval and content generation module is used to search the knowledge graph based on the input content, provide the retrieved knowledge to the large language model, and obtain detailed explanations of relevant knowledge points, case analysis, and extended knowledge;

[0013] The knowledge application service module is used to provide medical education services based on the learning and understanding of the knowledge graph by the graph neural network model and the integration of semantic information in the knowledge graph by the large language model; the knowledge application services include intelligent question and answer services, automatic question setting services, and answer interpretation services.

[0014] The beneficial effects of the present invention are as follows: the present invention overcomes many deficiencies in the existing technology in the processing and application of medical knowledge by deeply integrating the medical knowledge graph, the large language model, and the graph neural network model, more efficiently realizes the integration and in-depth understanding of medical knowledge, improves the efficiency and quality of knowledge dissemination and application in the medical field, and promotes the intelligent development of medical education and research. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a structural diagram of a knowledge graph construction, retrieval and application system for medical education proposed in the present invention;

[0016] Figure 2 Flowchart of building learning mode for system operation offline;

[0017] Figure 3 Flowchart of the online search generation model for the system. DETAILED DESCRIPTION

[0018] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0019] like Figure 1 As shown, the present invention proposes a knowledge graph construction, retrieval and application system for medical education, including:

[0020] Data collection and integration module, used to collect and integrate various medical expertise to form a medical knowledge material library;

[0021] Intelligent algorithm module, used to build knowledge graph based on medical knowledge material library; using graph neural network model to learn and understand knowledge graph and provide structured knowledge support for large language model;

[0022] The knowledge retrieval and content generation module is used to search the knowledge graph based on the input content, integrate the retrieved local knowledge and global knowledge and provide them to the large language model to obtain detailed explanations of relevant knowledge points, case analysis, and extended knowledge;

[0023] The knowledge application service module is used to provide medical education services based on the learning and understanding of the knowledge graph by the graph neural network model and the integration of semantic information in the knowledge graph by the large language model; the knowledge application services include intelligent question and answer services (AI teaching assistants), automatic question setting services (AI question setting), and answer interpretation services (AI paper marking).

[0024] The present invention closely integrates the knowledge graph, graph neural network model and large language model, and uses the graph neural network model's ability to understand the knowledge graph relationship to provide structured knowledge support for the large language model. At the same time, it uses the large language model's powerful language generation ability to improve the expression of the graph neural network model's reasoning results, thereby enhancing the reasoning and answering quality of the knowledge application system for medical education in medical knowledge processing.

[0025] In one embodiment of the present invention, collecting and integrating various medical expertise specifically includes:

[0026] We collect various medical textbooks as core content, and use optical character recognition (OCR), image recognition, and text extraction to convert the collected textbooks into digital format to ensure that the textbook content can be effectively read and processed by the computer system.

[0027] Collect video and audio materials of medical teachers' lectures as auxiliary content, use speech recognition technology to convert the video and audio materials into text, and associate them with the corresponding timeline; obtain medical teachers' lecture courseware and handouts, and extract the text, charts, and formulas therein.

[0028] An index system is established to associate knowledge points in textbook content, teacher lecture videos and audios, teacher lecture courseware and handouts; the index system is indexed based on the medical discipline system and knowledge classification dimensions, such as human body systems (digestive system, respiratory system, etc.), disease types (infectious diseases, genetic diseases, etc.), and medical concepts (cell structure, physiological mechanism, etc.); after content extraction from the collected textbooks, teacher lecture videos and audios, teacher lecture courseware and handouts, the extracted content is structurally annotated according to the dimensions in the index system, and each chapter and paragraph in the textbook and each page and slide in the courseware and handouts are annotated and classified in the index system according to the main knowledge points they cover, so as to facilitate subsequent knowledge extraction and map construction, and finally form a medical knowledge material library.

[0029] This solution uses multi-source data such as digital medical textbooks to form a very detailed and comprehensive medical knowledge material library, which prepares for the construction of subsequent knowledge graphs, improves the speed and accuracy of knowledge extraction, fusion and updating, and ensures that the knowledge graph can more comprehensively and timely reflect the medical knowledge system. With the continuous development of the medical field, new knowledge and new research results continue to emerge. The knowledge graph construction solution of the present invention can easily integrate new versions of medical textbooks, the latest teaching content and other resources to achieve dynamic updating of the knowledge graph. By regularly scanning and extracting data from data sources, new or modified medical knowledge can be discovered and incorporated in a timely manner to ensure that the knowledge graph always reflects the cutting-edge status of the medical field. This is an advantage that many existing static knowledge graphs cannot achieve, and can better meet the long-term application needs of medical education and research.

[0030] The intelligent algorithm module in the present invention includes:

[0031] The knowledge extraction submodule is used to automatically extract knowledge from the medical knowledge library using a large language model, including medical entities (such as disease names, drug names, human organ names, etc.), relationships (such as the association between diseases and symptoms, the correspondence between drugs and disease treatments, etc.) and attribute information (such as descriptions of the pathogenesis of diseases, characteristics of drug action, etc.).

[0032] The knowledge fusion submodule is used to fuse the knowledge from different sources extracted from the medical knowledge material library and eliminate redundant and contradictory information. The fusion process specifically includes data alignment and mapping as well as knowledge integration and reconstruction. First, the knowledge from different sources extracted from the medical knowledge material library is matched and matched. After data alignment, the knowledge from different sources is merged and optimized. By building a unified knowledge structure, supplementing missing information and strengthening knowledge associations, a unified, complete and enhanced knowledge graph is established. Eliminating redundant and contradictory information means unifying the knowledge from different sources. For example, when the description of the same medical concept in the textbook and courseware is slightly different, the same knowledge point is unified and integrated through rule-based filtering, knowledge merging and simplification, conflict detection and conflict resolution strategies.

[0033] The present invention uses a large language model to automatically process multi-source data, avoiding the tedious process of manually combing knowledge points one by one, and greatly improving the efficiency of knowledge graph construction. At the same time, the fusion of multi-source data enables the knowledge graph to contain richer entity, relationship and attribute information, such as combining theoretical knowledge in textbooks with teaching examples, so that the knowledge graph is significantly improved in the integrity of medical knowledge, providing a more solid foundation for subsequent intelligent applications. Eliminating redundant and contradictory information from knowledge from different sources can not only increase the breadth of knowledge, but also reduce resource waste and information confusion, improve the efficiency and accuracy of knowledge management, and enhance the performance and robustness of the entire system.

[0034] The knowledge graph representation learning submodule is used to use the graph neural network model to perform graph representation learning based on the constructed knowledge graph, map the entities and relationships in the knowledge graph to a high-dimensional vector space, obtain the embedding vectors of the entities and relationships, store these embedding vectors and provide the corresponding content to the large language model. Mapping the entities and relationships in the knowledge graph to a high-dimensional vector space includes the following steps:

[0035] A1. In offline mode, an initial high-dimensional vector is initialized for each entity and relationship in the knowledge graph through random initialization. The dimension of the vector is a hyperparameter. The dimension of the vector is selected according to the specific data scale, for example, it can be set to 500 dimensions, 1000 dimensions or even higher. The value of the initial vector is randomly generated within a set range, for example, the initial vector is randomly generated within the interval [-0.1, 0.1].

[0036] A2. Define the optimization goal: Define the optimization goal based on the translation model (such as the TransE model). For each triple (h, r, t) in the knowledge graph, it is expected that the head entity vector h plus the relationship vector r is as close to the tail entity vector t as possible, that is, h+r≈t; define a distance metric function (such as Euclidean distance or Manhattan distance) to measure the difference between h+r and t, that is, the objective function, and the optimization goal is to minimize the objective function; usually a margin parameter is added to increase constraints, such as max(0,||h+rt||-margin).

[0037] A3. Training and updating vectors: The objective function is optimized using the stochastic gradient descent (SGD) optimization algorithm. For each triple, the gradient is calculated based on the objective function, and the value of the vector is adjusted with the set learning rate (hyperparameter) so that the objective function gradually decreases, thereby updating the entity vector and the relationship vector. The translation model training process traverses multiple triples in the knowledge graph and performs multiple rounds of iterations until the objective function converges or reaches the predetermined number of training rounds. In order to improve the training effect during the training process, negative sampling technology is usually used, that is, for each real triple, a negative sample is randomly generated (for example, the tail entity is replaced with other random entities) and included in the training, so that the translation model can better distinguish between real and false relationships and enhance the quality of vector representation.

[0038] A4. Evaluation and tuning: After training, the performance of the translation model needs to be evaluated with evaluation indicators to determine the quality of the vector representation. The evaluation indicators include mean rank (MR), mean reciprocal rank (MRR), etc. According to the evaluation results, the hyperparameters of the translation model (such as vector dimension, learning rate, margin, etc.) are adjusted to retrain the translation model for better performance.

[0039] Through the above steps, continuous iteration and optimization, we can eventually obtain vector representations that map entities and relationships in the knowledge graph to high-dimensional vector space. These vectors can capture the semantic information of entities and relationships to a certain extent, which helps to perform efficient retrieval and reasoning in the knowledge graph. The graph neural network model can accurately understand the relationship between entities and vectors in the knowledge graph through learning, providing strong support for the application of subsequent systems. Different graph representation learning modules may differ in specific implementation details, but the overall idea is roughly the same.

[0040] As the core module of this solution, the knowledge retrieval and content generation module is used to answer relevant medical questions using the knowledge graph in the online mode of the system; the graph neural network model is used to understand the problem and retrieve the subgraphs in the knowledge graph that are similar to the problem, and then the large language model generates the corresponding answer based on the retrieval results of the graph neural network model. First, the feature vector of the problem is calculated using the question feature vector calculation submodule in the knowledge retrieval and content generation module, for example, the words in the question are converted into word vectors, and then these word vectors are combined into a feature vector that represents the semantics of the entire question through natural language processing (NLP). Then, the knowledge graph retrieval submodule is used to retrieve entity nodes or subgraphs similar to the problem in the knowledge graph based on the feature vector of the problem; finally, the subgraph pruning optimization submodule is used to prune and optimize the retrieved subgraphs, remove branches and nodes whose correlation with the problem is lower than the threshold, and retain the core and most relevant knowledge structure. The content generation submodule provides the content corresponding to the most relevant knowledge to the large language model for fusion and generates the corresponding answer;

[0041] Specifically, the specific retrieval process of the knowledge graph retrieval submodule includes: reasoning about the feature vector of the question through the graph neural network model, and then finding the embedding vectors (embedding) of entities and relationships related to the question based on the learned knowledge graph information through the graph neural network model and calculating the vector similarity, locating the knowledge nodes related to the question and their surrounding associated subgraph structures. These subgraphs contain medical knowledge entities and relationships that may be related to the answer to the question.

[0042] The pruning optimization is to use a word vector model or other semantic representation methods to calculate the semantic similarity between the question and the nodes and relationships in the retrieved subgraph. For the semantic representation of the node, the name, attributes and other information of the node can be converted into a vector form, and then the similarity with the feature vector of the question is calculated. For example, for a disease node, its name, pathogenesis description and other text information can be represented by a vector through a word vector model, and then the cosine similarity is calculated with the feature vector of the question. For the calculation of the semantic similarity of the relationship, it can be based on the definition and description of the relationship type. For example, the "treatment" relationship is more relevant to the semantics related to treatment in the question, while the "composition" relationship may have a low correlation in some treatment problems. According to the semantic similarity score, a similarity threshold is set, and for nodes and relationships, the parts with similarities below the threshold are regarded as having a low correlation with the question and pruned. For example, when answering questions about a disease diagnosis method, the physiological process nodes (low semantic similarity) in the subgraph that are not related to the disease diagnosis and their related relationships can be pruned, and the nodes and relationships such as symptoms and examination methods related to the diagnosis are retained.

[0043] After pruning optimization, the retrieval-augmented generation (RAG) technology is used to achieve accurate answers, that is, the large language model is fine-tuned using the relevant information retrieved by the graph neural network model. The fine-tuned large language model uses its understanding and generation capabilities of natural language combined with the received professional knowledge to organize the entities, relationships and attribute information in the retrieved knowledge graph subgraph into a smooth, accurate and medically logical answer text. The answers generated by the large language model are pushed to the application connected to the northbound interface for output, so that the application can present the answers to the user or perform further processing, such as showing them to students in teaching applications and providing them to doctors for reference in medical auxiliary diagnosis applications.

[0044] The present invention optimizes the retrieval algorithm of RAG (retrieval augmented generation) technology, improves the accuracy and speed of retrieving information from medical knowledge graphs, and integrates graph neural networks into the retrieval process, utilizing its mining of knowledge graph structural information to optimize retrieval strategies, achieving more accurate information retrieval and efficient integration, thereby improving the quality of intelligent question-answering and knowledge application services.

[0045] In this embodiment, when a student asks a question or a teacher involves a certain knowledge point in the teaching process, the knowledge retrieval and content generation module first uses the graph neural network model to search in the constructed medical knowledge graph. For example, if the teacher is explaining "treatment methods for heart disease", the system finds the "heart disease" node in the knowledge graph, and then obtains the treatment method nodes related to it and the relationships connecting them (such as drug treatment, surgical treatment, etc.). According to these relationships and node information, detailed explanation content is extracted from data sources such as textbooks and teaching materials, such as the mechanism of action, scope of application, side effects, etc. of various drugs in drug treatment, common surgical methods, surgical risks, etc. for surgical treatment, and redundant and conflicting information is eliminated. The knowledge point content obtained from the knowledge graph is then input into the large language model, and the powerful language organization ability of the large language model is used to polish and integrate the content, so that its expression is clearer, smoother, and more logical, which is convenient for students to understand. For example, the scattered knowledge points of drug action mechanism are organized into a well-organized text to explain how drugs act on the physiological process of the heart and why this effect can treat heart disease. During their after-school learning process, students can ask the system medical questions at any time. The system, like an intelligent teaching assistant, can give timely and accurate answers to help students understand difficult knowledge and promote independent learning.

[0046] This solution performs pruning optimization on the retrieved knowledge graph subgraphs, effectively removing nodes and relationships in the knowledge graph subgraphs that are less relevant to the question, retaining the most valuable information for subsequent large language model answer generation, and improving the efficiency and accuracy of intelligent question-answering services. Secondly, the large language model is deeply integrated into the knowledge retrieval and content generation modules, and the graph neural network model's ability to understand the knowledge graph is used in the answering process, which significantly improves the accuracy and professionalism of intelligent question-answering. When answering highly professional questions such as the pathogenesis of rare diseases in medicine, the large language model can avoid erroneous or ambiguous answers due to insufficient general data training based on the precise entity relationships and attribute information in the knowledge graph, and give accurate answers that conform to the medical professional knowledge system, greatly improving the reliability of intelligent question-answering in the medical professional field.

[0047] The close integration of the graph neural network model and the large language model gives the system stronger reasoning capabilities. When dealing with complex medical problems, the graph neural network model can conduct in-depth analysis of the knowledge structure in the knowledge graph, dig out potential logical relationships, and provide a strong reasoning basis for the large language model to generate answers. For example, when answering questions involving the correlation and differential diagnosis of multiple diseases, the system can use the graph neural network model to sort out the correlations and differences in symptoms, causes, treatments, and other aspects of different diseases in the knowledge graph. The large language model then conducts comprehensive reasoning and expression based on this, thereby giving a more logical and in-depth answer. Existing technologies often find it difficult to achieve such an effect in complex medical knowledge reasoning.

[0048] The knowledge application service module in this system includes an intelligent question-answering submodule, an automatic question-setting submodule, and an answer-judgment submodule.

[0049] The intelligent question-answering submodule is used to receive questions about medical courses requested by students and use the relevant medical questions as input to the knowledge retrieval and content generation module to finally get the answers to the corresponding questions;

[0050] Similarly, the automatic question generation submodule is also part of the knowledge application service module, which is used to receive medical teachers' requests for question generation, and use the keywords in the request as the input of the knowledge retrieval and content generation module, and use the graph neural network model to learn and understand the knowledge graph to automatically generate various types of medical questions. First, design a question template based on the knowledge graph: analyze the relationship between knowledge points and concepts and nodes in the medical knowledge graph through the graph neural network model, and design different types of question templates according to their characteristics; make each question template correspond to specific knowledge points and question type requirements, and make the questions fully cover the key content in the knowledge graph. For example, for disease-related knowledge points, you can design a multiple-choice question template such as "The following symptoms about [disease name] are wrong ()", a fill-in-the-blank question template such as "The main pathogenesis of [disease name] is [specific mechanism content]", and a short-answer question template such as "Briefly describe the diagnostic method of [disease name]". Each template corresponds to specific knowledge points and question type requirements to ensure that the question can fully cover the key content in the knowledge graph. Use the relationships in the knowledge graph to build more complex question templates. For example, comprehensive analysis questions are designed based on the therapeutic relationship between drugs and diseases, the correlation between diseases and symptoms, etc. For example, "It is known that drug A can be used to treat disease B. Common symptoms of disease B are C, D, and E. Please analyze how drug A relieves these symptoms by acting on human physiological processes." This type of question requires students to understand the relationship between multiple knowledge points and conduct comprehensive reasoning.

[0051] Secondly, generate questions based on the question template: For the variable part in the question template (such as disease name, symptoms, drugs, etc.), randomly select and fill in according to the entities and relationships in the knowledge graph. For example, in the multiple-choice question template, a disease is randomly selected from the knowledge graph, and then the correct and wrong options are selected from the relevant symptoms of the disease to combine. In order to ensure the rationality of the difficulty of the questions, different random selection probabilities are set according to the importance of the knowledge points and the learning stage of the students. For example, for beginners, common diseases and typical symptoms are more likely to be selected as the content of the questions; for advanced learners, the proportion of questions on rare diseases or complex symptom relationships is increased.

[0052] Logical constraints and rationality checks: When generating questions, add logical constraints to ensure the rationality of the questions in terms of medical knowledge. For example, when generating questions about drug dosage, the numerical range should be determined based on the conventional use range and safe dosage limit of the drug to avoid unreasonable dosage settings. At the same time, the generated questions are checked for rationality, such as checking the uniqueness of the answers (for multiple-choice questions), the accuracy of the question description (avoiding grammatical errors or ambiguity), etc. If it is found that the question does not meet the requirements, re-fill the parameters or adjust the question structure.

[0053] The automatic question generation submodule also stratifies the questions: determine the indicators for measuring the difficulty of the questions, including the depth of knowledge points (involving concepts or in-depth physiological mechanisms, etc.), the complexity of knowledge associations (single knowledge point or a combination of multiple knowledge points), and the reasoning steps required to answer the questions; assign a difficulty level to each question based on the question difficulty index, which is divided into three difficulty levels: elementary, intermediate, and advanced; analyze the students' learning level and knowledge mastery based on their learning history and answering situation, and provide them with questions of corresponding difficulty levels. For students with slower learning progress, give priority to choosing from elementary difficulty questions and focus on consolidating knowledge; for students with strong learning ability, gradually increase the proportion of intermediate and advanced difficulty questions to challenge their knowledge application and reasoning ability. At the same time, based on the types of wrong questions and weak knowledge points of students, generate relevant questions in a targeted manner for intensive training to help students improve their learning effects.

[0054] Based on the deep mining of medical knowledge graph and intelligent algorithm module, the present invention can provide personalized knowledge application services for different users. In the automatic question application, personalized questions can be customized for each student based on the student's learning history, answering status and other data, using the knowledge point association and difficulty level information in the knowledge graph to meet the needs of students with different learning levels and learning progress. In the intelligent question-answering application, it can also provide targeted tutoring content and learning suggestions based on the student's questioning habits and knowledge weaknesses. This personalized service helps to improve the pertinence and effectiveness of medical education, which is a functional feature generally lacking in the existing technology.

[0055] Furthermore, the present invention also provides a question answer interpretation service, which is used to receive the student's answer content as the input of the knowledge retrieval and content generation module, and perform semantic recognition based on the understanding of the knowledge graph by the graph neural network model, and refer to the standard answer of the question, and interpret and score the answer content in combination with the large language model. . After receiving the answer result, the system automatically identifies the answer content. For objective questions (multiple-choice questions, fill-in-the-blank questions, judgment questions, etc.), the answer result is compared with the reference answer using the large language model, and the correction result is directly given; for subjective questions (analysis questions, discussion questions, short-answer questions, etc.), the answer result is analyzed by the large language model and compared with the knowledge in the knowledge graph, and the accuracy, completeness and logic of the answer are analyzed, and the corresponding score and comments are given. In addition, the system can also perform learning analysis based on the student's answer situation, including statistics on wrong question types, knowledge weaknesses, and answering habits, and generate analysis results. Teachers can adjust teaching strategies accordingly, and students can review and strengthen learning in a targeted manner to improve learning efficiency.

[0056] like Figure 2 and Figure 3 As shown, the specific implementation process of this system includes two parts: offline and online.

[0057] The offline learning model is constructed. Its main tasks are to construct the knowledge graph of the course, train the graph neural network and learn the representation of the knowledge graph. The process is as follows: construct the knowledge graph of this medical course based on the data of the medical course (including textbooks, courseware, videos and other medical teaching materials), and then use the graph neural network to perform graph representation training on the knowledge graph of this medical course, and then calculate the embedding vectors of each entity and relationship in the knowledge graph of this medical course. Finally, these embedding vectors are stored in the database for subsequent online retrieval.

[0058] The main task of the online retrieval generation mode is to retrieve course-related knowledge from the questions input by the upper-level knowledge application, generate reply answers and feed them back to the upper-level knowledge application for output. The specific process is as follows: receive questions requested by the upper-level knowledge application as system input, calculate the feature vector of the question based on natural language processing technology, and then retrieve local related knowledge and global related knowledge from the stored knowledge graph embedding vector according to the similarity with the feature vector of the question, send the retrieved local related knowledge and global related knowledge to the large language model for fusion to generate answer content, and finally generate answer content and feed it back to the upper-level knowledge application for output.

[0059] The present invention realizes the integrated integration of multiple medical knowledge application services such as intelligent question and answer, automatic question setting, and answer interpretation. In the medical education scenario, teachers and students can complete multiple tasks such as teaching assistance, learning testing and evaluation in the same system without switching between multiple software or platforms with different functions, which greatly improves the convenience and efficiency of teaching and learning. For example, teachers can use the automatic question setting function to generate homework or test papers during lesson preparation, conduct interactive teaching in class with the help of the intelligent question and answer function, and understand students' learning situation through the student answer interpretation function after class. All operations are completed in a coherent system environment, reducing the inconvenience and time waste caused by platform switching and data transmission.

[0060] In summary, the present invention, through innovative technical solutions, demonstrates advantages significantly superior to existing technologies in many aspects such as medical knowledge graph construction, intelligent question and answer, and knowledge application services. It can better promote the intelligent development of the medical field and provide more efficient, accurate, and personalized support and services for medical education and research.

Claims

1. A knowledge graph construction, retrieval and application system for medical education, characterized by: include: Data collection and integration module, used to collect and integrate various medical expertise to form a medical knowledge material library; Intelligent algorithm module, used to build knowledge graph based on medical knowledge material library; Use graph neural network models to learn and understand knowledge graphs to provide structured knowledge support for large language models; The knowledge retrieval and content generation module is used to search the knowledge graph based on the input content, provide the retrieved knowledge to the large language model, and obtain detailed explanations of relevant knowledge points, case analysis, and extended knowledge; The knowledge application service module is used to provide medical education services based on the learning and understanding of the knowledge graph by the graph neural network model and the integration of semantic information in the knowledge graph by the large language model; Among them, knowledge application services include intelligent question and answer services, automatic question setting services, and answer interpretation services.

2. A knowledge graph construction, retrieval and application system for medical education according to claim 1, characterized in that: The collection and integration of various medical expertise specifically includes: Collect various medical textbooks and convert them into digital format using optical character recognition, image recognition, and text extraction; Collect video and audio materials of medical teachers' lectures, use speech recognition technology to convert the video and audio materials into text, and associate them with the corresponding timeline; Obtain medical teachers' lectures and handouts, and extract text, charts, and formulas; An index system is established to associate knowledge points in textbook content, teacher lecture videos and audios, teacher lecture courseware and handouts; the index system is indexed based on the dimensions of the medical discipline system and knowledge classification; after content extraction of the collected textbooks, teacher lecture videos and audios, teacher lecture courseware and handouts, the extracted content is structuredly annotated according to the dimensions in the index system to form a medical knowledge material library.

3. A knowledge graph construction, retrieval and application system for medical education according to claim 1, characterized in that: The intelligent algorithm module includes: The knowledge extraction submodule is used to automatically extract knowledge from the medical knowledge library using a large language model, including medical entities, relationships, and attribute information; The knowledge fusion submodule is used to fuse the knowledge extracted from different sources from the medical knowledge library and eliminate redundant and contradictory information to obtain a knowledge graph; The knowledge graph representation learning submodule is used to use the graph neural network model to perform graph representation learning based on the constructed knowledge graph, map the entities and relationships in the knowledge graph to a high-dimensional vector space, obtain the embedding vectors of the entities and relationships, and provide the corresponding content to the large language model.

4. A knowledge graph construction, retrieval and application system for medical education according to claim 3, characterized in that: The specific process of the knowledge fusion submodule to carry out fusion processing and eliminate redundant and contradictory information is as follows: Data alignment and mapping: matching and corresponding knowledge from different sources extracted from the medical knowledge library; Knowledge integration and reconstruction: after data alignment, knowledge from different sources is merged and optimized, and a unified, complete and enhanced knowledge graph is established by building a unified knowledge structure, supplementing missing information and strengthening knowledge associations; The specific methods of eliminating redundant information in the knowledge fusion submodule include rule-based filtering, knowledge merging and streamlining; The specific methods used by the knowledge fusion submodule to eliminate contradictory information include conflict detection and conflict resolution strategies.

5. A knowledge graph construction, retrieval and application system for medical education according to claim 3, characterized in that: The knowledge graph representation learning submodule maps entities and relationships in the knowledge graph to a high-dimensional vector space, including the following steps: A1. Initialize an initial vector for each entity and relationship in the knowledge graph by random initialization in the offline mode of the system; the dimension of the vector is selected according to the specific data scale, and the value of the initial vector is randomly generated within a set range; the dimension of the vector is a hyperparameter; A2. Define the optimization goal: Define the optimization goal based on the translation model. For each triple (h, r, t) in the knowledge graph, we expect the head entity vector h plus the relationship vector r to be as close as possible to the tail entity vector t, i.e. h+r≈t. Define a distance metric function to measure the difference between h+r and t, i.e. the objective function. The optimization goal is to minimize the objective function. A3. Training and updating vectors: The objective function is optimized using the stochastic gradient descent optimization algorithm. For each triple, its gradient is calculated according to the objective function, and the value of the vector is adjusted with the set learning rate so that the objective function gradually decreases. The translation model training process traverses multiple triples in the knowledge graph and performs multiple rounds of iterations until the objective function converges or reaches the predetermined number of training rounds. During the training process, for each real triple, a negative sample is randomly generated and included in the training. A4. Evaluation and tuning: The performance of the translation model is evaluated using evaluation indicators, and the hyperparameters of the translation model are adjusted according to the evaluation results to retrain the translation model; the evaluation indicators include average ranking and average reciprocal ranking.

6. A knowledge graph construction, retrieval and application system for medical education according to claim 1, characterized in that: The knowledge retrieval and content generation modules include: The question feature vector calculation submodule is used to convert the words in the question into word vectors in the system online mode, and then combine these word vectors into a feature vector that represents the semantics of the entire question through natural language processing technology; The knowledge graph retrieval submodule is used to retrieve entity nodes or subgraphs similar to the question in the knowledge graph based on the feature vector of the question; The subgraph pruning optimization submodule is used to prune and optimize the retrieved entity nodes or subgraphs, remove branches and nodes whose relevance to the question is lower than the threshold, and retain the core and most relevant knowledge structure; The content generation submodule is used to provide the content corresponding to the most relevant knowledge to the large language model for fusion and generate corresponding answers; Among them, the specific retrieval process of the knowledge graph retrieval submodule includes: The feature vector of the problem is inferred through the graph neural network model, and then the graph neural network model is used to find the embedded vectors of entities and relationships related to the problem based on the learned knowledge graph information, and the vector similarity is calculated; the knowledge nodes related to the problem and their surrounding associated subgraphs are located.

7. A knowledge graph construction, retrieval and application system for medical education according to claim 1, characterized in that: Knowledge application service modules include: The intelligent question-answering submodule is used to receive questions about medical courses requested by students and use the relevant medical questions as input to the knowledge retrieval and content generation module to finally get the answers to the corresponding questions; The automatic question generation submodule is used to receive medical teachers' requests for questions, use the question keywords in the request as input to the knowledge retrieval and content generation module, and use the graph neural network model to learn and understand the knowledge graph to automatically generate various types of medical questions; The answer interpretation submodule is used to receive students' answers as input to the knowledge retrieval and content generation module. It performs semantic recognition based on the graph neural network model's understanding of the knowledge graph, and interprets and scores the answers with reference to the standard answers to the questions in combination with the large language model.

8. A knowledge graph construction, retrieval and application system for medical education according to claim 7, characterized in that: The specific method of pruning optimization is: Use word vector models or other semantic representation methods to calculate the semantic similarity between the question and the nodes and relationships in the subgraph; set a similarity threshold based on the semantic similarity score and prune the nodes and relationships whose similarity is lower than the threshold; After pruning optimization, the large language model is fine-tuned using the relevant information retrieved by the graph neural network model. The fine-tuned large language model uses its ability to understand and generate natural language to organize the entities, relationships, and attribute information in the retrieved knowledge graph subgraphs into a fluent, accurate, and medically logical answer text.

9. A knowledge graph construction, retrieval and application system for medical education according to claim 7, characterized in that: The specific method of automatically generating questions in the automatic question generation submodule includes: B1. Design a question template based on knowledge graph: The knowledge points and concepts in the medical knowledge graph are analyzed through the graph neural network model, and different types of question templates are designed according to their characteristics; each question template corresponds to specific knowledge points and question type requirements, and the questions fully cover the key content in the knowledge graph; B2. Generate a topic based on the topic template: Randomly select and fill in the variable part of the question template based on the graph neural network model's understanding of the entities and relationships in the knowledge graph; during the random selection and filling process, set different random selection probabilities based on the importance of the knowledge point and the student's learning stage; In the process of generating questions, add logical constraints, check the rationality of the question, the uniqueness of the answer, and the accuracy of the question description. If the question does not meet the requirements, re-fill the parameters or adjust the question structure to generate the final question; B3. Layer the questions: Determine the indicators for measuring the difficulty of questions, including the depth of knowledge points, the complexity of knowledge associations, and the reasoning steps required to answer the questions; assign a difficulty level to each question based on the question difficulty indicators, and divide it into three difficulty levels: elementary, intermediate, and advanced; analyze students' learning level and knowledge mastery based on their learning history and answering situation, and provide them with questions of corresponding difficulty levels.

10. A knowledge graph construction, retrieval and application system for medical education according to claim 7, characterized in that: The specific method of the question-answering submodule to identify and determine the answer based on the knowledge graph is as follows: After receiving the answer results, for objective questions, the large language model is used to compare the answer results with the reference answers, and the grading results are given directly; for subjective questions, the large language model is used to perform text analysis on the answer results and compare them with the knowledge in the knowledge graph, analyze the accuracy, completeness and logic of the answers, and give corresponding scores and comments; learning analysis is conducted based on the students' answers, including statistics on wrong question types, knowledge weaknesses, and answering habits, to generate analysis results.

Citation Information

Cited By

  • Digital course generation method and system based on large language model

    CN120471029A

  • A method and system for generating digital courses based on a large language model

    CN120471029B

  • Teacher lesson preparation assisting method and related device based on knowledge graph and large language model

    CN120561272A

  • Knowledge question and answer rapid processing method and system based on artificial intelligence

    CN120596639A

  • Medical subject knowledge service-oriented data processing method and electronic equipment

    CN121615743A