A multi-drug resistant bacteria knowledge question and answer method and system based on a medical knowledge graph

By constructing a knowledge-based question-and-answer system for multidrug-resistant bacteria based on medical knowledge graphs and combining it with a large language learning model, the problem of obtaining information on multidrug-resistant bacteria was solved, achieving efficient integration and sharing of information and providing precise clinical treatment support.

CN118964565BActive Publication Date: 2025-11-18SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411027102.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-11-18
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

The treatment of multidrug-resistant bacterial infections is hampered by the lack of comprehensive and up-to-date information, difficulties in updating the knowledge of medical personnel, and the absence of effective integration and sharing mechanisms for existing technologies, leading to challenges in clinical decision-making.

Method used

We will construct a knowledge-based question-and-answer system for multidrug-resistant bacteria based on medical knowledge graphs. By combining this system with a large language learning model, we will build a knowledge graph and database for multidrug-resistant bacteria to achieve structured integration and interoperability of information, and provide efficient and accurate prevention and control solutions.

Benefits of technology

It provides comprehensive and systematic decision support for medical personnel, reduces the cognitive burden of knowledge acquisition and application, and assists in the precise treatment of multidrug-resistant bacterial infections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118964565B_ABST
    Figure CN118964565B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-drug resistant bacteria knowledge question and answer method and system based on medical knowledge graph, which comprises the following steps: constructing original multi-drug resistant bacteria data document;Construct multi-drug resistant bacteria knowledge graph;Generate local multi-drug resistant bacteria knowledge document;General large language learning model is trained, and multi-drug resistant bacteria large language learning model is obtained;Construct local multi-drug resistant bacteria database;Search in local multi-drug resistant bacteria database based on user query information, obtain knowledge information;Construct first prompt information, and input into multi-drug resistant bacteria large language learning model, output answer information.The system comprises original data acquisition module, multi-drug resistant bacteria relationship extraction module, general large language learning model, model training module, local database construction module, search module and answer output module.Using the application can provide more efficient, accurate multi-drug resistant bacteria prevention and control scheme.The application can be widely applied in the field of medical health technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical and health technology, and in particular to a question-and-answer method and system for multidrug-resistant bacteria based on medical knowledge graphs. Background Technology

[0002] Multidrug-resistant organism (MDRO) infections have become a significant challenge to global public health. Treatment of MDRO infections is increasingly difficult due to the rapid spread of resistant strains and the inappropriate use of antibiotics. Currently, healthcare professionals rely on the latest research findings and clinical experience to develop effective treatment plans, which places high demands on their professional knowledge and information access capabilities. However, medical information is scattered across different healthcare institutions and systems, lacking effective integration and sharing mechanisms, making it difficult for healthcare professionals to obtain comprehensive and up-to-date treatment information. Furthermore, while medical knowledge is rapidly evolving, traditional methods limit clinicians' access to new knowledge, making it difficult to keep up with the latest research findings. Therefore, clinical healthcare personnel are limited by their diagnostic capabilities for bacterial infections, their knowledge level, and their experience in administering medication, hindering their ability to make effective clinical decisions and advance practical clinical practice.

[0003] With the rapid development of cutting-edge technologies such as 5G communication, the Internet of Things and artificial intelligence, it has become possible to build intelligent question-and-answer systems for disease knowledge based on existing medical knowledge information. However, existing technologies still face the problem of a lack of the latest treatment knowledge for multidrug-resistant bacteria. Summary of the Invention

[0004] To address the aforementioned technical problems, the present invention aims to provide a knowledge-based question-and-answer method and system for multidrug-resistant bacteria based on medical knowledge graphs. This system combines medical knowledge graph technology with large language learning models to provide more efficient and accurate prevention and control solutions for multidrug-resistant bacteria.

[0005] The first technical solution adopted in this invention is: a question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs, comprising the following steps:

[0006] Construct the original multidrug-resistant bacteria data document;

[0007] Based on the multidrug-resistant bacteria relationship extraction model, knowledge relationships are extracted from the original multidrug-resistant bacteria data documents to obtain a multidrug-resistant bacteria knowledge graph.

[0008] A general large language learning model is constructed, which takes the multidrug-resistant bacteria knowledge graph and structured text information as input to generate local multidrug-resistant bacteria knowledge documents;

[0009] The general large language learning model is trained based on the knowledge graph of multidrug-resistant bacteria to obtain the multidrug-resistant bacteria large language learning model.

[0010] A local multidrug-resistant bacteria database is constructed based on the aforementioned local multidrug-resistant bacteria knowledge document;

[0011] Based on the user's query information, a search is performed in the local multidrug-resistant bacteria database to obtain knowledge information;

[0012] Based on the knowledge information and user query information, a first prompt is constructed, and the first prompt is input into the multidrug-resistant bacteria large language learning model to output the answer information.

[0013] Furthermore, the multidrug-resistant bacteria relationship extraction model includes an input layer, an input representation layer, a named entity recognition layer, a classifier layer, and an output layer.

[0014] Furthermore, the step of extracting knowledge relationships from the original multidrug-resistant bacteria data documents based on the multidrug-resistant bacteria relationship extraction model to obtain a multidrug-resistant bacteria knowledge graph specifically includes:

[0015] The original multidrug-resistant bacteria data document is preprocessed to obtain the structured text information and unstructured text information;

[0016] The unstructured text information is input into the input representation layer for semantic enrichment to obtain dynamic word vectors;

[0017] The dynamic word vectors are input into the named entity recognition layer for entity recognition to obtain the recognized entities.

[0018] The identified entities are input into the classifier layer, and the relationships between the identified entities are predicted by combining the context information of the identified entities to obtain the triplet relationship pairs of multidrug-resistant bacteria.

[0019] A knowledge graph of multidrug-resistant bacteria was constructed based on the ternary relationships of multidrug-resistant bacteria.

[0020] Furthermore, the step of constructing a general large language learning model, using the multidrug-resistant bacteria knowledge graph and structured text information as input, to generate a local multidrug-resistant bacteria knowledge document specifically includes:

[0021] The second prompt information is obtained by pre-setting prompts for the knowledge graph of multidrug-resistant bacteria and the structured text information respectively;

[0022] The second prompt information is generated by using a general large language learning model to obtain a local multidrug-resistant bacteria knowledge document.

[0023] Furthermore, the step of training the general large language learning model based on the multidrug-resistant bacteria knowledge graph to obtain the multidrug-resistant bacteria large language learning model specifically includes:

[0024] The triplet relationships of the multidrug-resistant bacteria knowledge graph are preset to obtain a third prompt message;

[0025] Based on a general large language learning model, the third prompt information is used to generate text, resulting in multidrug-resistant bacteria question-and-answer data pairs.

[0026] The question-and-answer data pairs of the multidrug-resistant bacteria were used as training data to fine-tune the general large language learning model, resulting in the multidrug-resistant bacteria large language learning model.

[0027] Furthermore, the step of constructing a local multidrug-resistant bacteria database based on the local multidrug-resistant bacteria knowledge document specifically includes:

[0028] Content extraction is performed on the local multidrug-resistant bacteria knowledge document to obtain the multidrug-resistant bacteria text content;

[0029] The text content of the multidrug-resistant bacteria is segmented to obtain the multidrug-resistant bacteria text sentences;

[0030] A query document is constructed based on the text of the multidrug-resistant bacteria;

[0031] The query document is converted into a text vector based on a vector model.

[0032] An index is created based on the text vector to obtain the text index;

[0033] The text vectors and text indexes are stored in a local database to obtain a local multidrug-resistant bacteria database.

[0034] Furthermore, the step of searching the local multidrug-resistant bacteria database based on user query information to obtain knowledge information specifically includes:

[0035] Based on the vector model, the user query information is vectorized into text to obtain the question query vector;

[0036] Based on the query vector, relevant text indexes are searched in the local multidrug-resistant bacteria database to obtain knowledge information.

[0037] The second technical solution adopted in this invention is: a question-and-answer system for multidrug-resistant bacteria based on medical knowledge graphs, comprising:

[0038] The raw data acquisition module is used to construct raw multidrug-resistant bacteria data documents;

[0039] The multidrug-resistant bacteria relationship extraction module is used to extract knowledge relationships from the original multidrug-resistant bacteria data document to obtain a multidrug-resistant bacteria knowledge graph.

[0040] A general-purpose large language learning model, taking the multidrug-resistant bacteria knowledge graph and structured text information as input, generates a local multidrug-resistant bacteria knowledge document;

[0041] The model training module trains the general large language learning model based on the multidrug-resistant bacteria knowledge graph to obtain the multidrug-resistant bacteria large language learning model.

[0042] The local database construction module constructs a local multidrug-resistant bacteria database based on the local multidrug-resistant bacteria knowledge document.

[0043] The search module searches the local multidrug-resistant bacteria database based on user query information to obtain knowledge information.

[0044] The answer output module constructs a first prompt based on the knowledge information and user query information, inputs the first prompt into the multidrug-resistant bacteria large language learning model, and outputs the answer information.

[0045] The beneficial effects of the method and system of this invention are as follows: The method constructs a medical atlas of multidrug-resistant bacteria, structurally integrating discrete medical information to form an interconnected and shared knowledge network, providing comprehensive and systematic decision support for clinical medical staff. Based on the medical atlas, a local multidrug-resistant bacteria knowledge base document is constructed, enabling synchronous updates of medical knowledge information and maintaining consistency with the latest research results and clinical data. Combining the advantages of medical knowledge graphs and large language learning models, medical personnel can interact with the system using natural language, conveniently querying key information such as the causes, susceptible populations, treatment plans, and theoretical basis of multidrug-resistant bacteria. This reduces the cognitive burden on medical personnel in knowledge acquisition and application, assisting them in making more accurate and effective clinical treatment choices for the prevention, diagnosis, and treatment of multidrug-resistant bacteria. Attached Figure Description

[0046] Figure 1 This is a flowchart of the steps of a knowledge question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs according to the present invention;

[0047] Figure 2 This is a structural block diagram of a knowledge question-and-answer system for multidrug-resistant bacteria based on medical knowledge graphs according to the present invention;

[0048] Figure 3 This is a schematic flowchart of a question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs according to the present invention.

[0049] Figure 4 This is a flowchart illustrating the local construction process of multidrug-resistant bacteria knowledge documents using a medical knowledge graph-based question-and-answer method.

[0050] Figure 5 This is a schematic diagram of the multidrug-resistant bacteria relationship extraction model structure of a multidrug-resistant bacteria knowledge question-answering method based on medical knowledge graph of the present invention;

[0051] Figure 6 This is a schematic diagram of the general large language learning model structure of a knowledge question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs according to the present invention;

[0052] Figure 7 This is a flowchart of the training process for a general large language learning model based on a medical knowledge graph-based question-and-answer method for multidrug-resistant bacteria.

[0053] Figure 8 This is a flowchart illustrating the construction process of a local multidrug-resistant bacteria database based on a medical knowledge graph-based question-and-answer method according to the present invention. Detailed Implementation

[0054] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adapted according to the understanding of those skilled in the art.

[0055] Reference Figure 1 and Figure 3 This invention provides a question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs, the method comprising the following steps:

[0056] S1. Construct the original multidrug-resistant bacteria data document;

[0057] Specifically, the original multidrug-resistant bacteria data documents include, but are not limited to, clinical guidelines, academic papers, and case reports. These documents may exist in various file formats, such as doc, pdf, and txt, and they are the cornerstone of constructing a multidrug-resistant bacteria knowledge graph.

[0058] S2. Based on the multidrug-resistant bacteria relationship extraction model, knowledge relationships are extracted from the original multidrug-resistant bacteria data documents to obtain a multidrug-resistant bacteria knowledge graph.

[0059] Specifically, refer to Figure 5 The multidrug-resistant bacteria relationship extraction model includes an input layer, an input representation layer, a named entity recognition layer, a classifier layer, and an output layer.

[0060] The input layer is the starting point of the entire multidrug-resistant bacteria relationship extraction model, responsible for receiving the text data to be processed. In the multidrug-resistant bacteria relationship extraction model, this text data includes medical literature, clinical reports, medical records, etc., and contains information about multidrug-resistant bacteria.

[0061] S2.1 Perform text preprocessing on the original multidrug-resistant bacteria data document to obtain the structured text information and unstructured text information;

[0062] Specifically, refer to Figure 4 Using OCR technology and NLP text extraction algorithms, text information is extracted from the structured and unstructured data in the original multidrug-resistant bacteria data document, resulting in structured and unstructured text information. During the text extraction process, preprocessing operations are also performed on the original multidrug-resistant bacteria data document, including noise removal, format standardization, and language correction, to improve data quality.

[0063] S2.2. Input the unstructured text information into the input representation layer for semantic enrichment to obtain dynamic word vectors;

[0064] Specifically, the input representation layer is mainly used to process the input multidrug-resistant bacteria text. In a specific embodiment of this invention, the input representation layer uses the BERT model. The BERT model can generate dynamic word vectors to understand the complex context in multidrug-resistant bacteria text and obtain deep semantic information, solving the problem of polysemy. For medical terminology in multidrug-resistant bacteria text, the BERT model can obtain richer semantic representations, improving the ability of the multidrug-resistant bacteria relation extraction model to capture the features of multidrug-resistant bacteria text. The BERT model utilizes NLP word segmentation technology, adding a "[CLS]" classification marker and a "[SEP]" segmentation marker to the beginning and end of the text information sequence X, respectively. For a text information sequence X of fixed length n, its expression is as follows:

[0065] X = (x1, x2, ..., x n )

[0066] x1 = CLS

[0067] x n =SEP

[0068] Where, x n This represents the nth text input.

[0069] The BERT model outputs a fixed-length text information sequence vector, the expression of which is as follows:

[0070] V e =BERT(x1,x2,…,xn)=(e1,e2,…,e n )

[0071] Among them, V e This represents the text sequence vector output by the BERT model; e nThe mapping value corresponding to the nth text input.

[0072] When the length of the input text information sequence is less than the fixed length of the BERT model, the output of the BERT model will be padded with blanks.

[0073] S2.3. Input the dynamic word vector into the named entity recognition layer for entity recognition to obtain the recognized entities;

[0074] Specifically, the named entity recognition layer is mainly responsible for identifying medical entities with specific meanings from the input dynamic word vectors, such as MDRO names, treatment names, and susceptible populations. The named entity recognition layer uses the concept of a global pointer. Assuming that an input sequence X of length n contains only one type of multidrug-resistant bacterial entity to be identified, and each multidrug-resistant bacterial entity to be identified is a subsequence of sequence X and can be nested (i.e., the entity parts overlap), then an input sequence X of length n has... The problem of identifying multidrug-resistant bacterial entities is transformed into identifying multiple candidate entity fragments from... The problem of multi-label classification of k multidrug-resistant bacterial entities selected from different candidate entity fragments.

[0075] Furthermore, if the input sequence X has m types of multidrug-resistant bacterial entities that need to be identified, a separate classification head can be created for each type of multidrug-resistant bacterial entity.

[0076] The above process can be modeled as the dynamic word vector V e A linear transformation yields the first and last vectors of the sequence, expressed as follows:

[0077] p i,α =w p,α e i +b p,α

[0078] q i,α =w q,α e i +b q,α

[0079] P=(p 1,α ,p 2,α ,…,p n,α )

[0080] Q = (q 1,α ,q 2,α ,…,q n,α )

[0081] α∈{1,2,…,m}

[0082] Where, p i,αq represents the probability that the i-th position in the sequence is the head vector of the α-th type entity; i,α e represents the probability that the i-th position in the sequence is the tail vector of the α-th type entity; i w represents the mapping value corresponding to the i-th text input; p,α The weight parameter represents the weight of the α-th class entity in the sequence header vector; w q,α b represents the weight parameter of the α-th class entity in the tail vector of the sequence; p,α b represents the bias parameter of the α-th type entity in the sequence head vector; q,α The bias parameter represents the α-th type of entity in the sequence tail vector; α represents the α-th type of multidrug-resistant bacterial entity; m represents the number of entity types; P represents the sequence head vector; Q represents the sequence tail vector.

[0083] Based on the first and last vectors, a continuous segment X can be defined. [i:j] The expression for a score of entity category α is as follows:

[0084]

[0085] i∈{1,2,…,n}

[0086] j∈{1,2,…,n}

[0087] Among them, s α (i,j) represents the score of entity category α for the continuous segment from i to j.

[0088] To enhance the identification of multidrug-resistant bacteria, the model also uses positional encoding on top of the original linear transformation to enhance the information available at each position. Specifically, a transformation matrix R is added to the head vector and tail vector of the sequence. i ,get:

[0089]

[0090]

[0091] Among them, R j-i R represents the transformation matrix from position i to j; j q represents the change matrix offset by j positions; j,α This represents the probability that the j-th position in the sequence is a tail vector of the α-th type entity.

[0092] Thus, for each score value s α (i,j) can explicitly contain relative information about other positions. Note that s α (i,j) can be computed in parallel, and ideally can achieve a time complexity of nearly O(1).

[0093] S2.4 Input the identified entities into the classifier layer and combine the context information of the identified entities to predict the relationship between the identified entities, and obtain the triplet relationship pairs of multidrug-resistant bacteria;

[0094] Specifically, the classifier layer is the decision part of the relation extraction model. It predicts the relationships between multidrug-resistant bacterial entities based on the identified entities and their contextual information. In a specific embodiment of the invention, a convolutional neural network is used to classify and predict entity relationships. Considering the score s obtained in step S2.3... α (i,j) is a The classification problem of candidate entities may result in severe class imbalance. To better achieve the identification and classification of multidrug-resistant bacteria, this invention employs a generalized form of "softmax + cross entropy," with the loss function defined as follows:

[0095]

[0096] Ω = {(i,j)|1≤i≤j≤n}

[0097] U α ={(i,j)|type(X) [ i :j] )∈α}

[0098] V α =Ω-U α

[0099] Where (i,j) represents candidate entity fragment X [i:j] ;U α V represents the positive sample label, which is the first and last set of all entities in the sample whose category is α; α α represents the negative sample label, which is the first and last set of all entities in the sample whose category is not α; Ω represents the entire sample set.

[0100] The loss function described above does not exhibit class imbalance because it does not transform multi-label classification into multiple binary classification problems. Instead, it transforms the target class score into a pairwise comparison of the non-target class scores and automatically balances the weight of each item by using the exponential and logarithmic forms of the function.

[0101] After entities and relations are identified, the model will generate a series of triplet relation pairs (s i ,e i1 ,o i ,e i2 ,r i ), where s i Represents the subject of the triple, e i1 The entity type of the subject, oi Denotes the object of the triple, e i2 The entity type of the object, r i It represents the relationship between entities.

[0102] S2.5 Construct a knowledge graph of multidrug-resistant bacteria based on the ternary relationships of multidrug-resistant bacteria.

[0103] S3. Construct a general large language learning model, using the multidrug-resistant bacteria knowledge graph and structured text information as input, to generate a local multidrug-resistant bacteria knowledge document;

[0104] Specifically, refer to Figure 4 First, a multidrug-resistant organism (MDRO) relation extraction model is used to extract structured MRO triplet relation pairs from unstructured MRO text. Then, a general large language learning model is constructed to efficiently generate local MRO knowledge documents. The entire process is intelligent and automated. The structured text information consists of high-quality MRO triplet relation pairs manually annotated by clinical medical staff.

[0105] S3.1. Preset prompts for the multidrug-resistant bacteria knowledge graph and structured text information respectively to obtain the second prompt information;

[0106] Specifically, the second prompt message is: "Assuming you are a professional medical worker, based on the input multidrug-resistant bacteria ternary relation (s)..." i ,e i1 ,o i ,e i2 ,r i ), where s i Represents the subject of the triple, e i1 The entity type of the subject, o i Denotes the object of the triple, e i2 The entity type of the object, r i This indicates the relationship between entities. The resulting text should provide a professional description of this triple relationship, without abbreviations.

[0107] S3.2. Based on the general large language learning model, the second prompt information is generated into text to obtain a local multidrug-resistant bacteria knowledge document.

[0108] Specifically, refer to Figure 6 The general large language learning model in a specific embodiment of the present invention is based on the Transformer architecture. The general large language learning model integrates a self-attention mechanism and a feedforward neural network to achieve deep language understanding.

[0109] When text data is input into the model, it is first fed into the word embeddings layer, a large-scale embedding matrix with dimensions (130528, 4096). This layer is responsible for converting each word in the vocabulary into its 4096-dimensional numerical representation, providing the model with a raw, parsable text form. Subsequently, the input data flows through a list of 28 identical GLMBlocks, which are responsible for further processing and transformation of the word embedding output. Each GLMBlock contains an input layer normalization module (Input LayerNorm), a self-attention module, a post-attention layer normalization module (Post-Attention LayerNorm), and a feed-forward network.

[0110] The input layer normalization is used at the beginning of each processing layer to normalize the input data, which helps maintain the stability of model training and speeds up convergence.

[0111] The self-attention module, a core feature of the Transformer architecture, enables the model to consider all other words in the text sequence while processing any given word. The self-attention module includes Rotary Embedding; Query, Key, Value; and Dense Connection layers.

[0112] The rotation embedding is an innovative positional encoding method that effectively encodes positional information into a self-attention mechanism.

[0113] The query, key, and value elements are derived from the input vector through a linear transformation and are used to calculate attention weights and integrate information from other words.

[0114] The densely connected layer further processes the output of the self-attention mechanism to form the final self-attention result.

[0115] The self-attention post-level normalization module refers to the process where, after self-attention processing, the data undergoes another layer of normalization to prepare it for processing in the next layer.

[0116] The feedforward network is located after the self-attention mechanism. The feedforward network contains gated linear units (GLUs) to further process the data and extract features.

[0117] After continuous processing through all GLMBlock layers, the output of the generalized large language learning model undergoes final layer normalization (Final LayerNorm), a step that ensures the overall stability of the model's output. Next, the data is fed into the generalized large language learning model head (LM Head), a linear transformation layer responsible for converting the model's output into a probability distribution for the next word. Finally, based on this probability distribution, the generalized large language learning model selects the word with the highest probability as the next output in the sequence, thereby generating the text and obtaining the local multidrug-resistant bacteria knowledge document.

[0118] S4. Train the general large language learning model based on the multidrug-resistant bacteria knowledge graph to obtain the multidrug-resistant bacteria large language learning model.

[0119] S.4.1 Pre-set hints for the triplet relationships in the multidrug-resistant bacteria knowledge graph to obtain third hint information.

[0120] Specifically, refer to Figure 7 In a specific embodiment of this invention, a pre-constructed multidrug-resistant bacteria knowledge graph is used to collaboratively train and generate a multidrug-resistant bacteria large language learning model. Its core is to transform the triplet relationship pairs in the multidrug-resistant bacteria knowledge graph into multidrug-resistant bacteria question-and-answer data pairs. The third prompt information is: "Assuming you are a professional medical worker, based on the input multidrug-resistant bacteria triplet relationship pair (s... i ,e i1 ,o i ,e i2 ,r i ), where s i Represents the subject of the triple, e i1 The entity type of the subject, o i Denotes the object of the triple, e i2 The entity type of the object, r i This represents the relationship between entities. Generate 5 question-answer pairs from this triple; the questions and answers must be closely related to the triple, and expansion is prohibited.

[0121] S4.2. Based on the general large language learning model, the third prompt information is used to generate text, and multidrug-resistant bacteria question-and-answer data pairs are obtained.

[0122] Specifically, the general large language learning model can convert triplet relationship pairs into multidrug-resistant bacteria question-and-answer data pairs based on the third prompt information.

[0123] S4.3. The question-and-answer data pairs of the multidrug-resistant bacteria are used as training data to fine-tune the general language learning model to obtain the multidrug-resistant bacteria language learning model.

[0124] Specifically, multidrug-resistant organism (MDRO) question-and-answer (MRA) data pairs will be used as part of the training data to help the general-purpose large language learning model learn how to answer MRA-related questions. The general-purpose MRA is pre-trained on a wide range of text data, possessing basic language understanding and generation capabilities. Then, the generated MRA question-and-answer data pairs and potentially other relevant medical data are used to fine-tune the general-purpose MRA. This fine-tuning process involves further training the model on specific tasks to better adapt it to the MRA-related medical domain. After fine-tuning, the general-purpose MRA transforms into a MRA-specific MRA model. This MRA MRA MRA model can more accurately understand and generate information related to MRA. Further training is needed after fine-tuning to optimize the model's performance. This includes adjusting the model's hyperparameters and improving training strategies. Once optimized, the MRA MRA MRA model is deployed throughout the system to provide MRA-related question-and-answer services. In summary, the training process for the multidrug-resistant bacteria large language learning model is a systematic process involving multiple steps such as data preparation, model selection, fine-tuning, training, evaluation, and optimization, and it can be dynamically iterated. Through this process, the system can be trained to produce a model that can accurately understand and answer questions related to multidrug-resistant bacteria, providing high-quality consultation services for medical professionals and patients.

[0125] S5. Construct a local multidrug-resistant bacteria database based on the local multidrug-resistant bacteria knowledge document;

[0126] S5.1 Extract content from the local multidrug-resistant bacteria knowledge document to obtain the multidrug-resistant bacteria text content;

[0127] Specifically, refer to Figure 8 Extract text-related content about multidrug-resistant bacteria from local multidrug-resistant bacteria knowledge documents in various file formats.

[0128] S5.2. Perform text segmentation on the text content of the multidrug-resistant bacteria to obtain the text sentences of the multidrug-resistant bacteria;

[0129] Specifically, the extracted text content of multidrug-resistant bacteria is broken down into smaller units, such as a sentence or paragraph, using text segmentation technology. This step helps to organize multidrug-resistant bacteria knowledge information into a format that is easy to process and retrieve.

[0130] S5.3 Construct a query document based on the text of the multidrug-resistant bacteria;

[0131] Specifically, the segmented text about multidrug-resistant bacteria is further constructed into query documents, which will be used in the retrieval process of the question-answering system. Constructing query documents includes tagging keywords, concepts, and entities.

[0132] S5.4. Based on the vector model, perform text conversion on the queried document to obtain a text vector;

[0133] Specifically, converting query documents into text vectors can capture the semantic information of the text. The choice of vectorization model is crucial for subsequent retrieval and analysis. In this step, a specific embodiment of the present invention selects a suitable vectorization model, M3E, to convert the text into a high-quality vector representation.

[0134] S5.5. Create an index based on the text vector to obtain the text index;

[0135] Specifically, the generated text vectors are used to create an index, which helps optimize the response time for querying MDRO-related knowledge information and improve retrieval efficiency.

[0136] S5.6. Store the text vector and the text index in a local database to obtain a local multidrug-resistant bacteria database.

[0137] Specifically, a suitable database storage solution is selected to store the vectorized text and indexes. In a specific embodiment of this invention, Qdrant and Qdrant are vector databases specifically designed for storing and retrieving vector data. These vector databases are suitable for many applications, such as similarity search, recommendation systems, and natural language processing. The data in the database is represented in vector form. After the local multidrug-resistant bacteria database is built, updating the local multidrug-resistant bacteria knowledge document will synchronize the local multidrug-resistant bacteria database.

[0138] S6. Based on the user query information, search the local multidrug-resistant bacteria database to obtain knowledge information;

[0139] S6.1. Based on the vector model, the user query information is vectorized into text to obtain the question query vector;

[0140] Specifically, the vectorization model in step S5.4 is used to convert the user-input query information into a question query vector.

[0141] S6.2. Based on the question query vector, search for relevant text indexes in the local multidrug-resistant bacteria database to obtain knowledge information.

[0142] Specifically, based on the query vector, relevant medical knowledge about multidrug-resistant bacteria is searched in the local multidrug-resistant bacteria database. The search method is index lookup, and the top K most relevant knowledge entries are returned. The expression is as follows:

[0143] Context = concat[q1,q2,…,q k ]

[0144] Where Context represents knowledge information; q k This represents the k-th piece of knowledge information.

[0145] S7. Construct a first prompt based on the knowledge information and user query information, and input the first prompt into the multidrug-resistant bacteria big language learning model to output the answer information.

[0146] Specifically, the first prompt message is: "The known knowledge information is Context. Based on the above knowledge information, answer the user's question concisely and professionally. If you cannot get the answer from it, please say 'The question cannot be answered according to the knowledge base information' or 'Insufficient relevant information has been provided.' Fabricated elements are not allowed in the answer. When the answer involves literature, cite it according to the format of the literature. Please use Chinese for the answer. The question is the query vector 'Query'." Then, the first prompt message is input into the multidrug-resistant bacteria big language learning model. The multidrug-resistant bacteria big language learning model generates a comprehensive and systematic answer based on the multidrug-resistant bacteria knowledge graph and information from the local multidrug-resistant bacteria database.

[0147] Finally, in a specific embodiment of the present invention, the system will continuously collect new medical data on multidrug-resistant bacteria and feedback from clinical medical staff, and continuously optimize its performance and knowledge base through iterative training.

[0148] Reference Figure 2 This invention provides a knowledge-based question-and-answer system for multidrug-resistant bacteria based on medical knowledge graphs, comprising:

[0149] The raw data acquisition module is used to construct raw multidrug-resistant bacteria data documents;

[0150] The multidrug-resistant bacteria relationship extraction module is used to extract knowledge relationships from the original multidrug-resistant bacteria data document to obtain a multidrug-resistant bacteria knowledge graph.

[0151] A general-purpose large language learning model, taking the multidrug-resistant bacteria knowledge graph and structured text information as input, generates a local multidrug-resistant bacteria knowledge document;

[0152] The model training module trains the general large language learning model based on the multidrug-resistant bacteria knowledge graph to obtain the multidrug-resistant bacteria large language learning model.

[0153] The local database construction module constructs a local multidrug-resistant bacteria database based on the local multidrug-resistant bacteria knowledge document.

[0154] The search module searches the local multidrug-resistant bacteria database based on user query information to obtain knowledge information.

[0155] The answer output module constructs a first prompt based on the knowledge information and user query information, inputs the first prompt into the multidrug-resistant bacteria large language learning model, and outputs the answer information.

[0156] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0157] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs, characterized in that, Includes the following steps: Construct the original multidrug-resistant bacteria data document; Based on the multidrug-resistant bacteria relationship extraction model, knowledge relationships are extracted from the original multidrug-resistant bacteria data documents to obtain a multidrug-resistant bacteria knowledge graph. A general large language learning model is constructed, which takes the multidrug-resistant bacteria knowledge graph and structured text information as input to generate local multidrug-resistant bacteria knowledge documents; The general large language learning model is trained based on the knowledge graph of multidrug-resistant bacteria to obtain the multidrug-resistant bacteria large language learning model. A local multidrug-resistant bacteria database is constructed based on the aforementioned local multidrug-resistant bacteria knowledge document; Based on the user's query information, a search is performed in the local multidrug-resistant bacteria database to obtain knowledge information; Based on the knowledge information and user query information, a first prompt is constructed, and the first prompt is input into the multidrug-resistant bacteria large language learning model to output the answer information; The multidrug-resistant bacteria relationship extraction model includes an input layer, an input representation layer, a named entity recognition layer, a classifier layer, and an output layer. The step of extracting knowledge relationships from the original multidrug-resistant bacteria data documents based on the multidrug-resistant bacteria relationship extraction model to obtain a multidrug-resistant bacteria knowledge graph specifically includes: The original multidrug-resistant bacteria data document is preprocessed to obtain the structured text information and unstructured text information; The unstructured text information is input into the input representation layer for semantic enrichment to obtain dynamic word vectors; The dynamic word vectors are input into the named entity recognition layer for entity recognition to obtain the recognized entities. The identified entities are input into the classifier layer, and the relationships between the identified entities are predicted by combining the context information of the identified entities to obtain the triplet relationship pairs of multidrug-resistant bacteria. A knowledge graph of multidrug-resistant bacteria was constructed based on the ternary relationships of multidrug-resistant bacteria.

2. The question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs according to claim 1, characterized in that, The step of constructing a general large language learning model, using the multidrug-resistant bacteria knowledge graph and structured text information as input, to generate local multidrug-resistant bacteria knowledge documents specifically includes: The second prompt information is obtained by pre-setting prompts for the knowledge graph of multidrug-resistant bacteria and the structured text information respectively; The second prompt information is generated by using a general large language learning model to obtain a local multidrug-resistant bacteria knowledge document.

3. The question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs according to claim 1, characterized in that, The step of training the general large language learning model based on the multidrug-resistant bacteria knowledge graph to obtain the multidrug-resistant bacteria large language learning model specifically includes: The triplet relationships of the multidrug-resistant bacteria knowledge graph are preset to obtain a third prompt message; Based on a general large language learning model, the third prompt information is used to generate text, resulting in multidrug-resistant bacteria question-and-answer data pairs. The question-and-answer data pairs of the multidrug-resistant bacteria were used as training data to fine-tune the general large language learning model, resulting in the multidrug-resistant bacteria large language learning model.

4. The question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs according to claim 1, characterized in that, The step of constructing a local multidrug-resistant bacteria database based on the local multidrug-resistant bacteria knowledge document specifically includes: Content extraction is performed on the local multidrug-resistant bacteria knowledge document to obtain the multidrug-resistant bacteria text content; The text content of the multidrug-resistant bacteria is segmented to obtain the multidrug-resistant bacteria text sentences; A query document is constructed based on the text of the multidrug-resistant bacteria; The query document is converted into a text vector based on a vector model. An index is created based on the text vector to obtain the text index; The text vectors and text indexes are stored in a local database to obtain a local multidrug-resistant bacteria database.

5. The question-and-answer method for multidrug-resistant bacteria based on medical knowledge graphs according to claim 1, characterized in that, The step of searching the local multidrug-resistant bacteria database based on user query information to obtain knowledge information specifically includes: Based on the vector model, the user query information is vectorized into text to obtain the question query vector; Based on the query vector, relevant text indexes are searched in the local multidrug-resistant bacteria database to obtain knowledge information.

6. A question-and-answer system for multidrug-resistant bacteria based on medical knowledge graphs, characterized in that, A method for performing a question-and-answer session on multidrug-resistant bacteria based on a medical knowledge graph as described in claim 1, comprising: The raw data acquisition module is used to construct raw multidrug-resistant bacteria data documents; The multidrug-resistant bacteria relationship extraction module is used to extract knowledge relationships from the original multidrug-resistant bacteria data document to obtain a multidrug-resistant bacteria knowledge graph. A general-purpose large language learning model, taking the multidrug-resistant bacteria knowledge graph and structured text information as input, generates a local multidrug-resistant bacteria knowledge document; The model training module trains the general large language learning model based on the multidrug-resistant bacteria knowledge graph to obtain the multidrug-resistant bacteria large language learning model. The local database construction module constructs a local multidrug-resistant bacteria database based on the local multidrug-resistant bacteria knowledge document. The search module searches the local multidrug-resistant bacteria database based on user query information to obtain knowledge information. The answer output module constructs a first prompt based on the knowledge information and user query information, inputs the first prompt into the multidrug-resistant bacteria large language learning model, and outputs the answer information.

Citation Information

Patent Citations

  • Knowledge graph generation type question answering method and system based on large language model

    CN117033608A

  • College scientific research management question and answer system combining knowledge graph and large language model

    CN117609436A