Intelligent Medical Literature Q&A System and Method Based on RAG and LLM Technologies
By combining RAG and LLM technologies, an intelligent question-and-answer system for medical literature has been built, which solves the problems of long reading time and low answer accuracy of medical literature, and achieves efficient and accurate information acquisition, improving user experience and research efficiency.
Patent Information
- Application Number
- CN202410635924.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-05-21
AI Technical Summary
In the prior art, reading and understanding of medical literature requires a long time to be invested, and the generation results of large language models in the professional field lack authenticity and accuracy, resulting in inconvenience and inefficiency of users to obtain information.
Combining RAG and LLM technologies, we build an intelligent question-and-answer system for medical literature based on RAG and LLM. We combine local knowledge base with large language models, use embedding models and GPT to generate answers, and limit the generation of answers through vector search and propt templates to ensure the accuracy and professionalism of answers.
It improves the convenience and efficiency of users to obtain medical information, supports bilingual translation and simultaneous recitation, improves user experience, and promotes the work efficiency of medical researchers.
Smart Images

Figure CN118364088B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent medical literature question - answering system and method, and particularly to an intelligent medical literature question - answering system and method based on RAG (Retrieval - Augmented Generation) and LLM (Large Language Model) technologies, belonging to the technical field of natural language processing. Background Art
[0002] In modern medical practice, medical guidelines and drug instructions are key documents to ensure medical safety and improve treatment effects. They provide important information for doctors on disease management, drug prescription, and patient care, and also enable patients to understand drug characteristics, usage, and dosage in more detail. However, some challenges also arise. A single literature has a large amount of content, resulting in a long time required for reading and understanding, which leads to the problem of long reading time.
[0003] In the past few years, significant progress has been made in artificial intelligence and natural language processing technologies. In particular, GPT (Generative Pre - trained Transformer) series models can understand and generate natural language. However, when facing problems in professional fields, the generated results of large language models may lack authenticity and accuracy and may produce hallucinations. To enhance the ability of large language models to handle problems in professional fields, the present invention combines RAG with the LLM large language model, and proposes an intelligent medical literature question - answering system that combines a local knowledge base with LLM, based on RAG and LLM technologies, to improve the quality and stability of the generated text. This combination has broad application prospects in the field of NLP, such as text summarization, dialogue systems, machine translation, etc. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent medical literature question - answering system and method based on RAG and LLM technologies.
[0005] Technical Solution 1:
[0006] An intelligent medical literature question - answering system based on RAG and LLM technologies, including an embedding model and a GPT large model;
[0007] The embedding model consists of an input layer, a word embedding layer, an encoder, and an output layer cascaded in sequence;
[0008] Input layer: The input natural language text is segmented into more than one token, and each token is arranged in sequence to form a sequence. Each token in the sequence is converted into a corresponding ID using a vocabulary table, and the ID is an integer; output the corresponding ID sequence;
[0009] Embedding layer: maps the ID sequence into word vectors of a fixed dimension, and the word vectors are represented by word vector E word_embedding , the position information of the token in the sequence E position_embedding , the pairwise sentence discrimination word vector E segment_embedding ; the output vector of the word embedding layer is:
[0010] E output = E word_embedding + E position_embedding + E segment_embedding ;
[0011] Encoder: composed of more than 1 encoding layers with the same structure but different parameters in series; each encoding layer includes a self-attention mechanism layer and a feed-forward neural network layer;
[0012] The self-attention mechanism layer processes the output vector of the word embedding layer as follows:
[0013] Step Z1: Generate query vector, key vector and value vector: the output vector of the embedding layer generates query vector, key vector and value vector through the corresponding weight matrix;
[0014] Step Z2: Calculate the attention score: each query vector calculates the dot product with all key vectors, and the attention score representing the similarity between the query and the key is obtained;
[0015] Step Z3: Calculate the attention weight: apply the softmax function to the attention score of each query vector and normalize it into a probability form, representing the attention weight corresponding to the key vector;
[0016] Step Z4: Calculate the attention vector: multiply the attention weight by the value vector to obtain the attention vector;
[0017] Feed-forward neural network layer: processes the attention vector as follows:
[0018] Step Q1: The first linear transformation layer expands the dimension of the attention vector by n times to obtain an extended attention vector;
[0019] Step Q2: Process the extended attention vector with a non-linear activation function;
[0020] Step Q3: The second linear transformation layer shrinks the dimension of the extended attention vector to 1 / n;
[0021] The processing step of the output layer is: extract the output vector of the last encoding layer and input it into a fully connected layer for linear transformation, and the result of the linear transformation will be processed by an activation function and then output.
[0022] Furthermore, the word vector representation E word_embedding, the position information E of the token in the sequence position_embedding Obtained through training and learning.
[0023] Furthermore, the encoder is composed of 24 encoding layers with the same structure but different parameters connected in series.
[0024] Furthermore, the similarity is divided by a scaling factor to obtain the attention score;
[0025] Even further, the scaling factor is the square root of the dimension of the key vector.
[0026] Technical solution two:
[0027] An embedding model training method for the intelligent question-answering system described in Technical Solution One, including the following steps:
[0028] Step 1: Data collection, which consists of the following specific steps.
[0029] Step 1-1: Data file format conversion: Convert the collected Chinese and English data into data files in a preset format;
[0030] Step 1-2: Data cleaning: Filter out the useless information in the data file to generate natural language text;
[0031] Step 2: Preparation of pre-training data for the Embedding model: The natural language text in the pre-training dataset includes medication references and clinical guidelines;
[0032] Step 3: Pre-training: Set the initial learning rate, batch size, and number of epochs;
[0033] Step 4: Preparation of fine-tuning data for the Embedding model: Organize the fine-tuning data into the form of text pairs, including query, pos, and neg; where query is the question, pos is the positive label, and neg is the negative label;
[0034] Step 5: Fine-tuning the model: Use the text pair data processed in Step 4 to fine-tune the pre-trained model in Step 3.
[0035] Furthermore, when fine-tuning the model, an instruction is added to the query for the retrieval task, and the AdamW optimizer is used.
[0036] Technical solution three:
[0037] A question-answering method for the intelligent question-answering system described in Technical Solution One,
[0038] First, construct a local knowledge base based on the medical field. Then, use the method of vector retrieval to screen out k answers that are closest to the user's question as a reference basis. Combine the question and the relevant answers to form a prompt and input it into the model. Through the understanding and analysis of the large language model, generate the final answer. And if the question exceeds the scope of the knowledge base, the model is prohibited from automatically generating answers to mislead users.
[0039] Furthermore, when collecting and preprocessing data, collect high-quality drug instructions and clinical guidelines and convert them into a unified format.
[0040] The specific steps for constructing the local knowledge base are as follows:
[0041] First, create a knowledge base. The knowledge base consists of more than 1 knowledge vector.
[0042] Then, upload more than one medical-related document, select an appropriate loader to load the document according to the source file type; and split the document into smaller chunks.
[0043] Use the Embedding model to convert the sliced text data into a high-dimensional vector representation to obtain the knowledge base vectors; store the obtained knowledge base vectors in the vector library.
[0044] The user inputs the question query, concatenate the instruction in front of the question query, and use the Embedding model to convert the instruction plus query into a question vector.
[0045] Calculate the cosine similarity between the question vector and the knowledge vectors in the knowledge base, sort the knowledge vectors according to the similarity from high to low, and return the text chunks corresponding to the first k knowledge vectors.
[0046] Input the text corresponding to the k knowledge vectors and the input question into the prompt template to construct the prompt input information, and input it into the GPT large model. After being polished by GPT, output the result of the large model.
[0047] Furthermore, the prompt template is adjusted according to the result of the large model, and the obtained optimized prompt template is: prompt_template = "
Instruction
Known Information
Question
[0048] Adopting the above technical solution, the beneficial effects achieved by the present invention are:
[0049] The present invention combines RAG with medical literature, improving the convenience for users to obtain content. Users only need to upload a document and enter the corresponding question, and the system will provide a high-quality content. At the same time, it also improves the efficiency of users to acquire knowledge. It also supports bilingual translations and simultaneous recitation, greatly enhancing the user experience. It realizes the application practice of large language models in the medical field, improves the work efficiency of medical researchers, helps them obtain the required information faster, and thus promotes innovation and development in the medical field. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is the structure diagram of the embedding model in Embodiment 1 of the present invention;
[0051] Figure 2 It is the flowchart of the intelligent medical literature Q&A system in Embodiment 2 of the present invention;
[0052] Figure 3 It is the training loss curve in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] Embodiment 1:
[0054] An intelligent medical literature Q&A system based on RAG and LLM technologies, including an embedding model and a GPT large model;
[0055] The embedding model is composed of an input layer, a word embedding layer, an encoder, and an output layer connected in cascade;
[0056] Input layer: The input natural language text is segmented into more than one token, and each token is arranged in sequence to form a sequence. A CLS marker is added at the beginning of the sequence, and a SEP marker is added at the end of the sequence. Then each token is converted into the corresponding ID using the vocabulary, and the ID is an integer; the corresponding ID sequence is output;
[0057] Embedding layer: Map the ID sequence to a word vector with a fixed dimension, and the word vector is represented by the word vector E word_embedding , the position information of the token in the sequence E position_embedding , the pairwise sentence discrimination word vector E segment_embedding ; The output vector of the word embedding layer is:
[0058] E output = E word_embedding + E position_embedding + E segment_embedding
[0059] The word vector representation E word_embedding , the position information of the token in the sequence E position_embeddingObtained through training and learning; the paired-sentence discrimination word vector E segment_embedding The vector representation used to distinguish the two sentences in a sentence pair. For the input of a single sentence, the segment embedding is all 0. For paired sentences, 0 and 1 are used respectively to distinguish them.
[0060] Encoder: It is composed of 24 encoding layers with the same structure but different parameters connected in series. Each encoding layer includes a self-attention mechanism layer and a feed-forward neural network layer.
[0061] Self-attention mechanism layer: Allows the model to calculate and weight the relationships of all tokens in the input sequence. This is a non-linear feature extraction process that achieves bidirectional context understanding. The specific calculation process is as follows:
[0062] Step Z1: Generate query vectors, key vectors, and value vectors: The output vectors of the embedding layer generate query vectors, key vectors, and value vectors through corresponding weight matrices; in this embodiment, the weight matrices are all implemented by fully connected layers;
[0063] Step Z2: Calculate the attention scores: Each query vector calculates the dot product with all key vectors to represent the similarity between the query and the key. Usually, it is divided by a scaling factor to obtain the attention scores. This factor is the square root of the dimension of the key vector, aiming to stabilize the training process and prevent gradient disappearance or gradient explosion.
[0064] Step Z3: Calculate the attention weights: Apply the softmax function to the attention scores of each query vector to normalize them into probability form, representing the attention weights corresponding to the key vectors.
[0065] Step Z4: Calculate the attention vectors: Multiply the attention weights by the value vectors to obtain the weighted sum, which is the attention vector, that is, the output of the self-attention layer.
[0066] Feed-forward neural network layer: The purpose is to deepen the network structure and help the model learn more complex representations. Usually includes two linear transformation layers and a non-linear activation function. The specific steps are as follows:
[0067] Step Q1: The first linear transformation layer expands the dimension of the attention vector by n times, where n is a preset multiple, to obtain the expanded attention vector;
[0068] Step Q2: Apply a non-linear activation function to process the expanded attention vector;
[0069] Step Q3: The second linear transformation layer reduces the dimension of the expanded attention vector to 1 / n;
[0070] In this embodiment, the dimension of the attention vector is 1024 and n is taken as 4. The non-linear activation function used in this embodiment is Gaussian Error Linear Unit, abbreviated as GELU. The non-linear activation function adds non-linearity to the network, enhancing the network's expressive ability, so that it can handle more complex tasks.
[0071] The processing steps of the output layer are as follows: The output vector of the last encoding layer is extracted and input into a fully connected layer for linear transformation. The result of the linear transformation will be processed by an activation function and then output. In this embodiment, the Tanh function is used for non-linear processing.
[0072] Embodiment 2: An embedding model training method for the intelligent question-answering system described in Embodiment 1, including the following steps:
[0073] Step 1: Data collection, which consists of the following specific steps.
[0074] Step 1-1: Data file format conversion: Convert the collected Chinese and English data into data files in a preset format. The Chinese and English data include disease questions, medication references, and clinical guidelines. In this embodiment, the literature type pdf data needs to be converted into txt format. For text type pdf files, use the python tool Fitz library to convert them into txt files. For pdf files containing pictures, use an ocr model to convert them, extract the text content, and convert it into a txt file.
[0075] Step 1-2: Data cleaning: Convert the text to a unified UTF-8 encoding to ensure consistent text format. Filter out the useless information in the data file to generate natural language text; the useless information includes headers, footers, labels, special symbols, garbled characters, and duplicate content.
[0076] Step 2: Preparation of pre-training data for the Embedding model: The data text units in the pre-training dataset include content such as medication references and clinical guidelines. In this embodiment, a total of 5 million data text units are used.
[0077] Step 3: In the pre-training stage, in this embodiment, the initial learning rate is set to 2e -5 , and the data parallel method is adopted. The batch size is set to 16 and the number of epochs is set to 2. By observing the loss curve, when the number of epochs is set to 2, the loss curve tends to be flat.
[0078] Step 4: Data preparation for fine-tuning the Embedding model: The fine-tuning data needs to be organized into text pairs, for example {"query":"What are the methods for children to do gastroscopy","pos":["There are two common methods for children to do gastroscopy: 1. Painful gastroscopy, gastroscopy is done without anesthesia; 2. Painless gastroscopy, gastroscopy is done under anesthesia. The common method in China does not recommend anesthesia. Because children have many side effects when anesthetized, including the side effects of the anesthetics themselves and the recovery after anesthesia. Without anesthesia, that is, without using anesthetics, the current common method is to grab the child and do a gastroscopy. Although this situation recovers quickly, it also has disadvantages. Because children are afraid of doing the examination. In the case of no anesthesia, it is very painful to do the examination, although the process is more It is relatively short-lived, but it will cause certain psychological shadows and pressure in the future. The comprehensive pros and cons are that these two methods have their own advantages and disadvantages. For children with good cooperation, it is recommended not to use anesthesia. For children who are uncooperative or even extremely uncooperative, it is recommended to do gastroscopy and related examinations under anesthesia. "],"neg":["Precautions before gastroscopy are as follows: 1. According to the doctor's judgment, the patient's condition should meet the examination indications; 2. Keep an empty stomach on the morning of the examination, and perform blood tests before the examination; 3. Some medications need to be stopped before the examination. For example, those who take aspirin, Panax notoginseng and other blood-activating and blood-stasis-removing drugs, or anticoagulants should try to stop taking the drugs for one week before the examination under the guidance of a doctor to ensure the safety of the examination; because suspicious lesions may be found during the examination, biopsies are required. If anticoagulants are taken, more bleeding may occur. "]}
[0079] Among them, query is the question, pos is the positive label, neg is the negative label, and other content modules are randomly selected as negative labels.
[0080] Step 5: Fine-tuning stage, use the text pair data processed in step 4 to fine-tune the pre-model in step 3. In order to improve the general ability of semantic vectors in multi-task scenarios, instructions are added to the query of the retrieval task in fine-tuning. For English, the instruction is "Represent this sentence for searching relevant passages:" For Chinese, the instruction is "Generate a representation for this sentence to retrieve relevant articles:", using the AdamW optimizer with a learning rate of 1e -5 The contrast loss temperature is 0.01. During the fine-tuning process, the model effect is evaluated by viewing the loss curve and calculating the accuracy and MRR indicators of each round of model in the test samples.
[0081] (1) Accuracy, that is, the number of hits in the test set among the topk recall results of each user is +1, and the total number of hits is divided by the total number of people.
[0082] (2) MRR metric: Among the top k recall results for each user, sort the recall list by score. For the items in the list that match the test set, record the position of the item in the list, and use the reciprocal of this position as the score. Add up the scores of all users and then divide by the total number of users. For example, if the MRR value for top 5 is 0.2, it means that the average hit position in the recall list for each user is the 5th. The formula is as follows:
[0083]
[0084] Among them, N represents the total number of users, and p i represents the position of the true access value of the i-th user in the recommendation list. If the value does not exist in the recommendation list, then pi -> ∞.
[0085] Example 3:
[0086] A question-answering method for the intelligent question-answering system described in Example 1. First, construct a local knowledge base based on the medical field. Then, use the method of vector retrieval to screen out k answers that are closest to the user's question as a reference. Combine the question and the relevant answers to form a prompt and input it into the model. Through the understanding and analysis of the large language model, generate the final answer. And if the question exceeds the scope of the knowledge base, prohibit the model from automatically generating answers to mislead users. The specific implementation steps are as follows:
[0087] Step 1: Data collection and preprocessing. Collect 20,000 high-quality drug instructions and clinical guidelines each. Convert the collected documents into a unified format, such as txt format. Delete irrelevant information, such as advertisements and copyright statements, and fix encoding problems to ensure that all texts use a unified character encoding, such as UTF-8.
[0088] Step 2: Construct a local knowledge base. The specific steps are as follows:
[0089] Step 2-1: First, create a knowledge base. Create different knowledge bases according to different fields. The knowledge base consists of more than 1 knowledge vector. In this example, two knowledge bases, namely drug instructions and clinical guidelines, are created.
[0090] Step 2-2: Upload more than one medical-related document. The supported document formats for uploading are pdf, md, word, and txt formats. Batch uploading of documents and single document uploading are supported.
[0091] Step 2-3: Load the document. Select a suitable loader according to the source file type. The role of the loader is to convert the formatted text into an unformatted string for subsequent processing.
[0092] Step 2-4: Split the document. Since a single document often exceeds the model context limit, the document needs to be split into smaller chunks. In this embodiment, the RecursiveCharacterTextSplitter tool is used. This tool recursively splits the document by different characters, taking into account the length of the split text and overlapping characters. RecursiveCharacterTextSplitter defaults to using the four special symbols ["\n\n", "\n", "", ""] as markers for splitting the text. The maximum length of the string to be cut can be set through chunk_size. In this embodiment, chunk_size is set to 512, and overlap is set to the number of overlapping characters between the two segments of the string to maintain the semantic coherence in the string. In this embodiment, overlap is set to 20.
[0093] Step 2-5: Vectorize the document. Use the fine-tuned Embedding model in Embodiment 2 to convert the cut text data into a high-dimensional vector representation to obtain the knowledge base vector.
[0094] Step 2-6: Store the vector. Store the obtained knowledge base vector in the vector library. In this embodiment, we use the faiss retrieval technology. Faiss is an open-source clustering and similarity search library by the Facebook AI team, which provides efficient similarity search and clustering for dense vectors, supports searches for vectors at the billion level, and is currently the most mature approximate nearest neighbor search library.
[0095] Step 3: The user inputs a question query, such as "What are the methods for a child to have a gastroscopy?".
[0096] Step 4: Concatenate the instruction in front of the question query, "Generate a representation for this sentence for retrieving relevant articles:", and use the fine-tuned Embedding model in Embodiment 2 to convert the instruction plus query into a question vector.
[0097] Step 5: Calculate the cosine similarity between the question vector and the knowledge vectors in the knowledge base, sort the knowledge vectors in descending order of similarity, and return the texts corresponding to the top k knowledge vectors. The cosine similarity algorithm refers to using the cosine value of the angle between two vectors in a vector space as a measure of the difference between two individuals. The closer the cosine value is to 1 and the closer the angle is to 0, the more similar the two vectors are. The closer the cosine value is to 0 and the closer the angle is to 90 degrees, the less similar the two vectors are.
[0098] Step 6: Input the texts corresponding to the k knowledge vectors and the input question into the prompt template to construct the prompt input information, and then input it into the GPT large model. After being polished by GPT, the result of the large model is output. In the prompt template, it is restricted that answers unrelated to the question are not allowed to be generated to prevent misleading users.
[0099] Step 7: Adjust the prompt template according to the result of the large model to obtain the optimized prompt template: prompt_template = "
Instruction
Known Information
Question
[0100] Collect user feedback information. Add feedback buttons for satisfaction and dissatisfaction after each answer returned by the intelligent medical literature Q&A system. After the system has been online for some time, manually review the user feedback information. Use the information that users are satisfied with as positive examples and the information that users are dissatisfied with as negative examples to iteratively optimize the retrieval effect of the embedding model.
Claims
1. An intelligent medical literature Q&A system based on RAG and LLM technologies, characterized in that: It includes an embedding model and a GPT large model; The embedding model consists of an input layer, a word embedding layer, an encoder, and an output layer connected in series in sequence; Input layer: The input natural language text is segmented into more than one token, and each token is arranged in sequence to form a sequence. Each token in the sequence is converted into a corresponding ID using a vocabulary table, and the ID is an integer; an ID sequence is output. Embedding layer: mapping the ID sequence into word vectors of a fixed dimension, where the word vectors are represented by word vectors , the position information of the token in the sequence , the paired sentence discrimination word vector ; The output vector of the word embedding layer is: ; Encoder: It is composed of more than 1 encoding layer with the same structure but different parameters connected in series; each encoding layer includes a self-attention mechanism layer and a feed-forward neural network layer; The self-attention mechanism layer processes the output vector of the word embedding layer as follows: Step Z1: Generate query vectors, key vectors, and value vectors: The output vector of the embedding layer generates query vectors, key vectors, and value vectors through corresponding weight matrices; Step Z2: Calculate attention scores: Each query vector calculates the dot product with all key vectors to represent the attention score indicating the similarity between the query and the key; Step Z3: Calculate attention weights: Apply the softmax function to the attention scores of each query vector to normalize them into a probability form, representing the attention weights corresponding to the key vectors; Step Z4: Calculate attention vectors: Multiply the attention weights by the value vectors to obtain attention vectors; Feed-forward neural network layer: Processes the attention vectors as follows: Step Q1: The first linear transformation layer expands the dimension of the attention vector by n times, where n is a preset multiple, to obtain an expanded attention vector; Step Q2: Apply a non-linear activation function to process the expanded attention vector; Step Q3: The second linear transformation layer reduces the dimension of the expanded attention vector to 1 / n; The processing steps of the output layer are: Extract the output vector of the last encoding layer and input it into a fully connected layer for linear transformation, and the result of the linear transformation will be processed by an activation function and then output.
2. The intelligent medical literature Q&A system based on RAG and LLM technologies according to claim 1, wherein: Word vector representation and the position information of the token in the sequence are obtained through training and learning.
3. The intelligent medical literature Q&A system based on RAG and LLM technologies according to claim 1, characterized in that: The encoder is composed of 24 encoding layers with the same structure but different parameters connected in series.
4. The intelligent medical literature Q&A system based on RAG and LLM technologies according to claim 2, wherein: The similarity is divided by a scaling factor to obtain the attention score.
5. The intelligent medical literature Q&A system based on RAG and LLM technologies according to claim 4, wherein: The scaling factor is the square root of the dimension of the key vector.
6. An embedding model training method for the intelligent question answering system described in claim 1, including the following steps: Step 1: Data collection, which consists of the following specific steps: Step 1-1: Data file format conversion: Convert the collected Chinese and English data into data files in a preset format; Step 1-2: Data cleaning: Filter out useless information in the data files to generate natural language text; Step 2: Prepare pre-training data for the Embedding model: The natural language text in the pre-training dataset includes medication references and clinical guidelines; Step 3: Pre-training: Set the initial learning rate, batch size, and number of epochs; Step 4: Prepare fine-tuning data for the Embedding model: Organize the fine-tuning data into text pair forms, including query, pos, and neg; where query is the question, pos is the positive label, and neg is the negative label; Step 5: Fine-tune the model: Use the text pair data processed in Step 4 to fine-tune the pre-trained model in Step 3.
7. The embedding model training method according to claim 6, wherein: When fine-tuning the model, an instruction was added to the query for the retrieval task, and the AdamW optimizer was used.
8. A question-answering method for the intelligent question-answering system according to claim 1, characterized in that: First, a local knowledge base based on the medical field is constructed, and then k answers closest to the user's question are selected as a reference basis through vector retrieval. The question and the relevant answers are formed into a prompt and input into the model. Through the understanding and analysis of the large language model, the final answer is generated. And if the question exceeds the scope of the knowledge base, the model is prohibited from automatically generating answers to mislead users.
9. A question-and-answer method according to claim 8, characterized in that: When collecting and preprocessing data, high-quality drug instructions and clinical guidelines are collected and converted into a unified format; The specific steps for constructing the local knowledge base are as follows: First, a knowledge base is newly created, and the knowledge base consists of more than 1 knowledge vector; Then, upload more than one medical-related document, select an appropriate loader to load the document according to the source file type; and divide the document into small chunks; Use the Embedding model to convert the cut text data into a high-dimensional vector representation to obtain the knowledge base vector; store the obtained knowledge base vector in the vector library; The user inputs a question query, concatenate an instruction before the question query, and use the Embedding model to convert the instruction plus query into a question vector; Calculate the cosine similarity between the question vector and the knowledge vectors in the knowledge base, sort the knowledge vectors according to the similarity from high to low, and return the texts corresponding to the top k knowledge vectors; Input the texts corresponding to the k knowledge vectors and the input question into the prompt template to construct the prompt input information, and input it into the GPT large model, and output the large model result after being polished by GPT.
10. A question-answering method according to claim 9, characterized in that: The prompt template is adjusted according to the large model result, and the obtained optimized prompt template is: prompt_template = "【Instruction】Answer the question professionally according to the known information; if the answer cannot be obtained from it, please say \"The question cannot be answered according to the known information\", and it is not allowed to add fabricated content to the answer, give corresponding explanations to the answer, and the answer should be in Chinese; \n\n【Known Information】{context} \n\n【Question】{question}"; context is the text corresponding to the top k knowledge vectors retrieved, and question is the question input by the user; Collect user feedback information. Add satisfied and dissatisfied feedback buttons after each answer returned by the intelligent medical literature question-answering system. After it has been on the line for a period of time, manually review the user's feedback information. Use the information that the user is satisfied with as positive examples and the dissatisfied information as negative examples to iteratively optimize the retrieval effect of the embedding model.
Citation Information
Patent Citations
Word vector generation method, system and equipment for electronic medical record and storage medium
CN117195877A
Traditional Chinese medicine question and answer method and device based on long document retrieval enhancement generation and medium
CN117828050A