Medical knowledge service methods, devices, electronic equipment and media
By building a medical discipline knowledge base and multi-level evaluation and scoring, the problems of low efficiency and low quality of literature acquisition in the medical discipline field have been solved, and efficient and accurate knowledge acquisition has been achieved.
Patent Information
- Application Number
- CN202411996077.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In the existing technology, the efficiency and quality of literature acquisition in the field of medical disciplines are low, and auxiliary tools are insufficient in convenience, accuracy and efficiency, which makes knowledge acquisition difficult.
Build a medical discipline knowledge base, including a document library, a document entity library, an evidence library, a map library and an index library. By pre-judging user query questions, determining the subject area and query method, and utilizing the interaction between the medical discipline knowledge base and the large language model, conduct multi-level evaluation and scoring of query and feedback results to ensure the accuracy and quality of the results.
It improves the efficiency and accuracy of medical knowledge acquisition, reduces workload and complexity, and ensures the comprehensiveness and timeliness of feedback results.
Smart Images

Figure CN119903129B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technology, and in particular to a medical subject knowledge service method, device, electronic equipment and medium. Background Art
[0002] As the aging population continues to grow, chronic diseases such as Alzheimer's disease (AD) and cancer are placing a heavy burden on patients, their families, and society. To better address these issues, researchers in the medical field need to efficiently access relevant literature and knowledge about advancements in their fields.
[0003] Currently, the most important means for researchers in this field to access literature and knowledge of field progress is manual acquisition and reading, which is time-consuming, inefficient, and low-quality. Existing auxiliary tools also have significant shortcomings in terms of convenience, accuracy, and efficiency. Therefore, in the context of the knowledge explosion, the inefficiency and low quality of accessing medical literature, evidence, knowledge, and ideas has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a medical discipline knowledge service method, device, electronic equipment and medium to improve the efficiency and accuracy of obtaining medical discipline-related knowledge, reduce workload and complexity, and improve the effectiveness of knowledge services.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present invention provides a medical discipline knowledge service method, comprising: obtaining medical literature, and constructing a medical discipline knowledge base corresponding to each subject area based on the medical literature; wherein the medical discipline knowledge base comprises: a medical discipline literature base, a medical discipline literature entity base, a medical discipline evidence base, a medical discipline atlas base, and a medical discipline index base; obtaining a query question input by a user, and pre-judging the query question to determine the subject area and query method to which the query question belongs; based on the query method, querying the medical discipline knowledge base corresponding to the subject area to which the query question belongs and / or interacting with a medical discipline large language model to obtain a result of the query question; wherein the result of the query question comprises one of the following: a first query result, a first feedback result, a second query result, and a second feedback result based on the second query result; if the first query result or If the first feedback result is obtained, the first query result or the first feedback result is fed back to the user; if the second feedback result is obtained, the second feedback result is evaluated and scored for relevance, and when the score result of the relevance evaluation score of the second feedback result meets the first preset condition, the second feedback result is evaluated and scored for consistency; when the score result of the relevance evaluation score of the second feedback result meets the second preset condition, the third feedback result is obtained based on the medical discipline large language model, and the third feedback result is evaluated and scored for secondary relevance; when the score result of the secondary relevance evaluation score of the third feedback result meets the first preset condition, the third feedback result is evaluated and scored for consistency; when the score result of the consistency evaluation score of the second feedback result or the third feedback result meets the third preset condition, the second feedback result or the third feedback result is fed back to the user.
[0007] In a second aspect, the present invention provides a medical discipline knowledge service device, comprising: a medical discipline knowledge base construction module, for acquiring medical literature, and constructing a medical discipline knowledge base corresponding to each subject field based on the medical literature; wherein the medical discipline knowledge base includes: a medical discipline literature library, a medical discipline literature entity library, a medical discipline evidence library, a medical discipline atlas library, and a medical discipline index library; a pre-judgment module, for acquiring a query question input by a user, and pre-judging the query question to determine the subject field and query method to which the query question belongs; a query and interaction module, for querying the medical discipline knowledge base corresponding to the subject field to which the query question belongs and / or interacting with a medical discipline large language model based on the query method to obtain the result of the query question; wherein the result of the query question includes one of the following: a first query result, a first feedback result, a second query result, and a second feedback result based on the second query result; A feedback and evaluation module is used to feed back the first query result or the first feedback result to the user if a first query result or a first feedback result is obtained; if a second feedback result is obtained, perform a relevance evaluation and scoring on the second feedback result, and when the score result of the relevance evaluation score of the second feedback result meets the first preset condition, perform a consistency evaluation and scoring on the second feedback result; when the score result of the relevance evaluation score of the second feedback result meets the second preset condition, obtain a third feedback result based on the medical discipline large language model, and perform a secondary relevance evaluation and scoring on the third feedback result; when the score result of the secondary relevance evaluation score of the third feedback result meets the first preset condition, perform a consistency evaluation and scoring on the third feedback result; when the score result of the consistency evaluation score of the second feedback result or the third feedback result meets the third preset condition, feed back the second feedback result or the third feedback result to the user.
[0008] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of any one of the methods provided in the first aspect above.
[0009] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes the steps of any one of the methods provided in the first aspect.
[0010] The present invention brings the following beneficial effects:
[0011] The above-mentioned medical subject knowledge service method, device, electronic device and medium provided by the present invention first obtain medical literature and build a medical subject knowledge base corresponding to each subject field based on the medical literature; secondly, obtain the query question input by the user, and pre-judge the query question to determine the subject field and query method to which the query question belongs; then, based on the query method, query the medical subject knowledge base corresponding to the subject field to which the query question belongs and / or interact with the medical subject large language model to obtain the result of the query question; finally, if a first query result or a first feedback result is obtained, the first query result or the first feedback result is fed back to the user; if a second feedback result is obtained, the second feedback result is processed accordingly. The method comprises the following steps: first, performing a relevance evaluation and scoring on the second feedback result, and when the scoring result of the relevance evaluation score of the second feedback result meets the first preset condition, performing a consistency evaluation and scoring on the second feedback result, and when the scoring result of the relevance evaluation score of the second feedback result meets the second preset condition, obtaining a third feedback result based on the medical discipline big language model, and performing a secondary relevance evaluation and scoring on the third feedback result; when the scoring result of the secondary relevance evaluation score of the third feedback result meets the first preset condition, performing a consistency evaluation and scoring on the third feedback result; and when the scoring result of the consistency evaluation score of the second feedback result or the third feedback result meets the third preset condition, feeding back the second feedback result or the third feedback result to the user.
[0012] The above method can ensure the comprehensiveness and timeliness of knowledge in different subject areas by constructing a real-time updated medical subject knowledge base in different subject areas, including a medical subject literature base, a medical subject literature entity base, a medical subject evidence base, a medical subject atlas base, and a medical subject index base; by pre-judging the query questions input by the user, determining the subject area and query method to which it belongs, and then querying and / or interacting based on the medical subject knowledge base and / or medical subject large language model corresponding to the subject area to which it belongs, it can improve the accuracy and quality of the feedback results on the basis of improving service efficiency; finally, by performing relevance evaluation scoring, secondary relevance evaluation scoring and consistency evaluation scoring on the query results and feedback results, the results that meet the requirements are fed back to the user, further improving the accuracy and quality of obtaining medical subject-related knowledge.
[0013] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0014] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the preferred embodiments are specifically listed below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 A flowchart of a medical subject knowledge service method provided by an embodiment of the present invention;
[0017] Figure 2 A schematic diagram of the structure of a medical subject knowledge service device provided by an embodiment of the present invention;
[0018] Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] At present, in the context of knowledge explosion, the low efficiency and low quality of obtaining medical-related literature, evidence, knowledge, ideas, etc. has become a problem that needs to be solved urgently.
[0021] Based on this, the embodiments of the present invention provide a medical subject knowledge service method, device, electronic device and medium, which can improve the efficiency and accuracy of obtaining medical subject-related knowledge, reduce workload and complexity, and improve the effectiveness of knowledge services.
[0022] To facilitate understanding of this embodiment, a medical subject knowledge service method disclosed in an embodiment of the present invention is first introduced in detail. The method can be executed by an electronic device such as a smart phone, a computer, a tablet computer, etc. Figure 1 The flowchart of a medical subject knowledge service method shown in FIG. 1 illustrates that the method mainly includes the following steps S101 to S107:
[0023] Step S101: Obtain medical literature and construct a medical subject knowledge base corresponding to each subject area based on the medical literature.
[0024] In one embodiment, in order to support intelligent knowledge services for different subject areas, it is necessary to build a knowledge base of relevant research results in different subject areas, that is, a medical subject knowledge base in different subject areas (such as: Alzheimer's disease field, tumor field, etc.), including: a medical subject literature library, a medical subject literature entity library, a medical subject evidence library, a medical subject atlas library, and a medical subject index library.
[0025] In order to ensure the integrity of the content of the medical knowledge base, in this embodiment, multiple databases, such as PubMed, Embase, Cochrane, Wanfang Medical, etc., can be searched and simultaneously documents in various medical fields can be obtained. For the obtained documents, a machine learning method is used to accurately distinguish the documents in different subject areas, and then the documents in each subject area are respectively processed by extracting metadata, processing based on rules (pre-set), constructing classification models for processing, and constructing information extraction models for processing to construct medical subject document libraries in various subject areas; and based on the pre-constructed entity information extraction model, the documents are entity extracted to construct medical subject document entity libraries in various subject areas; for each document in the medical subject document library in each subject area, the above-constructed medical subject document entity library is combined to construct a medical subject atlas library in each subject area; for documents with RCT labels in their document types, evidence information is extracted based on the pre-constructed evidence extraction model to construct a medical subject evidence library in various subject areas; finally, a medical subject index library in various subject areas is constructed.
[0026] Step S102: obtaining a query question input by the user, and pre-judging the query question to determine the subject area and query method to which the query question belongs.
[0027] In one embodiment, the user service interface includes traditional keyword input, and users can enter query questions by entering keywords or express their query questions in natural language.
[0028] In order to ensure the quality of knowledge services and improve service efficiency, in this embodiment, a pre-judgment of the query questions input by the user can be performed to determine the subject field to which the query questions belong and the query method (that is, the method of answering the query questions). Then, according to the subject field to which the query questions belong, the corresponding medical subject knowledge base and / or medical subject large language model can be used to query and / or interact.
[0029] Step S103: Based on the query method, a medical subject knowledge base query corresponding to the subject area to which the query question belongs and / or a medical subject large language model interaction are performed to obtain a result of the query question.
[0030] In one embodiment, based on the above-mentioned query question, a corresponding medical discipline knowledge base query and / or medical discipline large language model interaction is performed to obtain the result of the query question; wherein the result of the query question includes one of the following: a first query result (corresponding to the result of the medical discipline knowledge base query), a first feedback result (corresponding to the result of the medical discipline large language model interaction), a second query result, and a second feedback result based on the second query result (corresponding to the result of the combination of the medical discipline knowledge base query and the medical discipline large language model interaction). The medical discipline knowledge base query process includes the following three links: querying the medical discipline literature library, querying the medical discipline literature entity library, and querying the medical discipline evidence library.
[0031] Step S104: If a first query result or a first feedback result is obtained, the first query result or the first feedback result is fed back to the user.
[0032] Step S105: If a second feedback result is obtained, a relevance evaluation score is performed on the second feedback result, and when the score result of the relevance evaluation score of the second feedback result meets the first preset condition, a consistency evaluation score is performed on the second feedback result. When the score result of the relevance evaluation score of the second feedback result meets the second preset condition, a third feedback result is obtained based on the medical discipline large language model, and a secondary relevance evaluation score is performed on the third feedback result.
[0033] Step S106: When the scoring result of the secondary relevance evaluation of the third feedback result meets the first preset condition, a consistency evaluation score is performed on the third feedback result.
[0034] Step S107: When the consistency evaluation score of the second feedback result or the third feedback result meets the third preset condition, the second feedback result or the third feedback result is fed back to the user.
[0035] In one embodiment, if a first query result is obtained through a query of the medical subject knowledge base, or a first feedback result is obtained through interaction with a medical subject large language model, the first query result or the first feedback result is directly fed back to the user; if a second feedback result is obtained through a query of the medical subject knowledge base and interaction with a medical subject large language model, the second feedback result of the obtained medical subject large language model is subjected to content inspection and filtering, as well as relevance evaluation and scoring.
[0036] If the relevance evaluation score of the second feedback result meets the first preset condition (the score is greater than or equal to the second threshold but less than the first threshold), the second feedback result is scored for consistency; if the relevance evaluation score of the second feedback result meets the second preset condition (the score is less than the second threshold), the third feedback result is obtained based on the entity list generated in the process of processing the user's query question and the constructed medical subject atlas library, as well as the medical subject large language model, and the third feedback result is scored for secondary relevance evaluation; otherwise (that is, the score is greater than or equal to the first threshold), the second feedback result is fed back to the user.
[0037] If the result of the secondary correlation evaluation score of the third feedback result meets the first preset condition, the third feedback result is evaluated and scored for consistency; if the score result of the secondary correlation evaluation score of the third feedback result meets the second preset condition, the medical discipline large language model interaction is re-performed to obtain a new second feedback result, and the above steps are repeated; otherwise, the third feedback result is fed back to the user; if the result of the consistency score of the second feedback result or the third feedback result meets the third preset condition (the score is greater than or equal to the third threshold), the second feedback result or the third feedback result is fed back to the user.
[0038] The above-mentioned medical discipline knowledge service method provided by the embodiment of the present invention can ensure the comprehensiveness and timeliness of knowledge in different disciplines by constructing a medical discipline knowledge base in different disciplines that is updated in real time and includes a medical discipline literature base, a medical discipline literature entity base, a medical discipline evidence base, a medical discipline atlas base, and a medical discipline index base; by pre-judging the query questions input by the user, determining the discipline field and query method to which it belongs, and then querying and / or interacting based on the medical discipline knowledge base and / or medical discipline large language model corresponding to the discipline field to which it belongs, it can improve the accuracy and quality of the feedback results on the basis of improving service efficiency; finally, by performing relevance evaluation scoring, secondary relevance evaluation scoring and consistency evaluation scoring on the query results and feedback results, the results that meet the requirements are fed back to the user, further improving the accuracy and quality of obtaining medical discipline-related knowledge.
[0039] The aforementioned step S101, i.e., constructing a medical knowledge base corresponding to each subject area based on medical literature, includes the following steps S1011 to S1016:
[0040] Step S1011: Classify medical literature according to subject areas to obtain initial medical subject literature libraries in different subject areas.
[0041] In one embodiment, after obtaining relevant medical literature, the medical literature is first classified to obtain literature in different subject areas. Specifically, machine learning methods can be used to distinguish the subject areas of the literature based on the title, abstract, keywords, journal name, etc., such as building a model based on BERT + tuning, to obtain an initial medical literature library in different subject areas.
[0042] Step S1012: For each document in the initial medical subject document library, the document is processed based on the pre-built metadata extraction model, knowledge rules, classification model and information extraction model, and a medical subject document library is constructed based on the content and metadata corresponding to each tag of each document.
[0043] In one embodiment, the medical discipline literature library framework can be constructed using a multi-level and multi-link approach. Overall, the first-level categories of the medical discipline literature library framework (the first-level categories of the knowledge system labels) include literature title information, content location information, literature type information, disease information, focus information, link information, measurement method information, intervention classification information, intervention information, evaluation standard information, evaluation indicator information, quality evaluation information, methodology evaluation information, comprehensive evaluation information, etc. Further, the first-level categories can be subdivided, for example, literature title information can be divided into title, abstract, keywords, author, affiliated institution, literature publication medium name (journal), impact factor, publication time, etc.; content location information can be further subdivided into title, abstract, keywords, introduction, background, method, key, experiment, discussion, outlook, etc.; literature type information can be further subdivided into clinical trial, randomized controlled trial (Randomized Controlled Trial), etc. Clinical Trial, RCT), meta-analysis, systematic review, review, guideline, and others; disease information can be further subdivided into AD, MCI, tumor, etc.; focus information can be further subdivided into calculation, reasoning, problem solving, decision-making, perception, memory, attention, visual space, execution, learning, language, etc. (taking AD as an example); measurement method information can be further subdivided into imaging, EEG, MEG, scale, etc. (taking AD as an example); link information can be further subdivided into screening, diagnosis, treatment, prognosis, etc.; intervention classification information can be further subdivided into traditional Chinese medicine, Western medicine, surgery, non-drug methods, etc.; intervention information can be further subdivided into Yizhi Jiannao Granules, Yizhi Drink, Rehmannia Yizhi Recipe, etc. (taking traditional Chinese medicine intervention as an example); evaluation standard information can be further subdivided into diagnostic criteria, efficacy evaluation table, etc.; evaluation indicator information can be further subdivided into Mini-mental State Examination (Mini-mental State Examination) Evaluation of report quality can be further divided into dimensions such as experimental design, data analysis, and data interpretation. Methodological evaluation can be further divided into random sequence generation method, allocation concealment, implementation bias, measurement bias, and follow-up bias. Comprehensive evaluation information can be further divided into quality assessment and recommendation strength (using grade as an example). Furthermore, these categories can be further subdivided depending on the specific situation.
[0044] For the literature in each subject area, according to the above knowledge system, each document in the literature library is processed by extracting metadata, processing based on rules (pre-set), processing based on a pre-built classification model, and processing based on a pre-built information extraction model. The corresponding label content of each component in the system is obtained and saved in the corresponding fields in the medical discipline literature library (including the paragraph and structural position, sentence and specific character position, etc., if any), to form a complete medical discipline literature library.
[0045] Step S1013: For each document in the medical subject literature library, entities in each document are extracted based on a pre-built entity information extraction model to construct a medical subject literature entity library.
[0046] In one embodiment, the entity categories in the medical subject literature entity library include: diseases, symptoms, examinations, tests, operations, scales, brain regions / parts, drugs, prescriptions, Chinese patent medicines, Chinese herbal medicines, surgeries, adverse reactions, data objects, processing algorithms and other categories. For each document in the medical subject literature library, based on a pre-built entity information extraction model (such as a seq2seq-based entity extraction model), the entities in the document are extracted and saved according to the document ID, entity ID, entity category, specific entity, location (including the paragraph and structural position, sentence and specific character position, etc.), related document data (including the document library knowledge system label and content corresponding to the document) and other fields to form a medical subject literature entity library.
[0047] Step S1014: For each document in the medical subject literature library, a relationship graph between the document and the entity is constructed based on the document and the medical subject literature entity library, and a triple graph is constructed based on the relationship between the entities in the document in the medical subject literature entity library, and a medical subject graph library is constructed based on the relationship graph and the triple graph.
[0048] In one embodiment, the medical discipline atlas library may be a two-layer structure, including: a first-layer relationship graph consisting of entities and documents, and a second-layer triple graph consisting of entities and relationships between entities within documents.
[0049] During specific implementation, for each document in the medical subject document library, the above-mentioned medical subject document entity library is combined to construct a medical subject atlas library. First, based on the document and the entity inside it, a relationship map between the entity and the document is constructed (the interaction between the document ID and the entity ID, that is, if there is an entity in the document, then this entity and the corresponding document establish a connection relationship, otherwise there is no connection relationship); then based on the relationship between the entities in the document, a triple map is constructed. Specifically, based on a pre-built relationship information extraction model (such as a relationship extraction model based on a pre-trained model such as Bert), the relationship between the entities in the document is extracted to construct a triple map (interaction between entity IDs). In order to ensure that the context of the triples can be distinguished, in this embodiment, attribute values can be added to the relationship of the triples (the value range includes the document ID-which can take multiple values, and the relevant document data includes the document library knowledge system label and content corresponding to the document); finally, the extracted map is saved in the medical subject atlas library (such as neo4j).
[0050] Step S1015: For each document in the medical subject library, if the document type tag content includes RCT, the evidence information in each document is extracted based on the pre-built evidence extraction model to build a medical subject evidence library; wherein the evidence information is information related to RCT.
[0051] In one embodiment, the medical discipline evidence library includes at least: patient status, diagnostic criteria (such as DSM-IV diagnostic criteria, etc.), grouping situation, grouping basis (such as random number table method, computer random number method, odd-even numbering method, etc.), sample size, demographic data (such as age distribution, gender distribution, etc.), intervention measures (such as placebo treatment, Western medicine treatment, Chinese medicine treatment, etc.), implementation methods (such as type of medicine, dosage, course of treatment), evaluation indicators, group intervention results, adverse events, experimental results, experimental significance, bias risk assessment, etc.
[0052] For each document in the medical discipline library, if the document type label includes RCT, evidence information is extracted based on a pre-built evidence extraction model (such as an evidence extraction model based on seq2seq), and saved according to the document ID, evidence ID, evidence field, evidence content, location (including the paragraph and structural position, sentence and specific character position, etc.), related document data (including the document's corresponding document library knowledge system label and content) and other fields to form a medical discipline evidence library.
[0053] Step S1016: construct sub-indexes of the medical subject literature library, the medical subject literature entity library, and the medical subject evidence library respectively, and integrate and associate the medical subject literature library, the medical subject literature entity library, and the medical subject evidence library based on the association IDs between the data of the medical subject literature library, the medical subject literature entity library, and the medical subject evidence library to construct an overall index to obtain a medical subject index library; wherein the sub-index and the overall index include: an inverted index and a vector index.
[0054] In one embodiment, a medical discipline index library is constructed. First, sub-indexes are constructed for the medical discipline document library, the medical discipline document entity library, and the medical discipline evidence library. Then, based on the association IDs (such as document IDs) between the data of the medical discipline document library, the medical discipline document entity library, and the medical discipline evidence library, the data of each library is integrated and associated to construct an overall index for the entire medical discipline. Secondly, for each sub-index and the overall index, an inverted index and a vector index are constructed respectively. In the process of constructing the vector index, the long text is first segmented (for example, segmented according to a length of 500 characters), and then the segmented text is vectorized (based on a pre-trained embedding model) and then stored separately. In order to further improve the efficiency and effect of retrieval, in this embodiment, in the process of constructing the vector index, for each segmented text block, a prompt method can be used based on a fine-tuned large model to input the entire document (to which the text block belongs) and the segmented text block into the large model to obtain a context summary for the segmented text block. The obtained context summary and the segmented text block are then connected to obtain the input text for constructing the index. In addition, for the medical discipline evidence database, a digital index is also constructed for the digital fields therein.
[0055] Therefore, a medical discipline knowledge base was constructed, which is specifically composed of a medical discipline literature library, a medical discipline literature entity library, a medical discipline atlas library, a medical discipline evidence library, and a medical discipline index library.
[0056] It is worth noting that the medical discipline knowledge base includes various parts such as the document library, document entity library, atlas library, evidence library, and index library, which will be updated in a timely manner based on the acquisition of source document data, thus ensuring the latest content of the entire medical discipline knowledge base system.
[0057] The aforementioned step S102, i.e., pre-judging the query question to determine the subject field and query method to which the query question belongs, includes: firstly judging the subject field to which the query question belongs to determine the subject field to which the query question belongs; then judging the objectivity and answerability of the query question. If the question is an objective question and an answerable question, the query method of the query question is determined based on a pre-built query method judgment model.
[0058] In one embodiment, a user-entered query is first pre-judged, specifically including: query objectivity, subject matter domain, and answerability. Objectivity refers to determining whether the query is objective and requires a true and definitive response (opposite questions include subjective questions, questions for which there is no consensus, and imaginary questions). Subject matter domain determination refers to determining the subject matter to which the user-entered query belongs. Answerability determination refers to determining whether the query can be correctly answered.
[0059] The above-mentioned judgments on the objectivity of the query question, the subject area to which the query question belongs, and the answerability of the query question can be completed by a pre-built judgment model (the judgment model is fine-tuned based on a pre-trained model such as T5). For questions that are judged by the judgment model to be non-objective or non-answerable, the user is directly prompted and then returned; for questions that are judged by the judgment model to be objective or answerable, the corresponding medical subject knowledge base and medical subject large language model (a base model such as llama is selected and obtained through steps such as fine-tuning across the entire medical field and fine-tuning in medical subjects) can be called according to the judgment result of the subject area to which the judgment model belongs to perform the following processing.
[0060] Furthermore, the query method of the user-entered query is determined: the method for answering the question is determined, namely, a method based on querying the medical knowledge base corresponding to the subject area of the query, a method based on interaction with the medical subject language model corresponding to the subject area of the query, or a method based on a combination of querying the medical subject knowledge base corresponding to the subject area of the query and the medical subject language model. Query method determination can be performed by a pre-built query method determination model (the model is fine-tuned based on a pre-trained model such as T5).
[0061] In one possible implementation, based on the query method, a medical subject knowledge base query corresponding to the subject area to which the query question belongs and / or a medical subject large language model interaction are performed, and the query results obtained include the following three situations:
[0062] Case 1: If the query method is an interactive combination of querying the medical subject knowledge base corresponding to the subject field to which the query question belongs and the medical subject large language model, the query question is preprocessed, and based on the preprocessed query question, the medical subject knowledge base corresponding to the subject field to which the query question belongs is queried to obtain a second query result, and based on the second query result, the first prompt content of the medical subject large language model corresponding to the subject field to which the query question belongs is constructed, and the first prompt content is input into the medical subject large language model to obtain a second feedback result based on the second query result.
[0063] Specifically, for questions that require a correct answer based on the medical subject knowledge base corresponding to the subject area to which the query question belongs and the medical subject large language model, the query question is first preprocessed, and then the medical subject knowledge base corresponding to the subject area to which the query question belongs is queried based on the preprocessed query question to obtain a second query result. Specifically, the following steps (1) to (8) are included:
[0064] Step (1): Input the query question into the medical subject language model corresponding to the subject area to which the query question belongs to perform problem decomposition to obtain a decomposed question, and then concatenate the decomposed question with the query question to obtain a concatenated question.
[0065] In the specific implementation, the user's query question Query is input into the medical discipline language model corresponding to the subject field to which the query question belongs. By requiring the user to solve the problem step by step, the medical discipline language model is prompted to decompose the problem and obtain the decomposed answer. 1 , and concatenate the decomposition problem and the query problem to obtain the concatenated problem Query con =(Query, Answer 1 ).
[0066] Step (2): Perform word segmentation and keyword extraction on the splicing problem to obtain a keyword list, and input the keyword list into a pre-built synonym language model to obtain a synonym list of the keyword.
[0067] In the specific implementation, a word segmenter (such as: stuttering) is used to query the splicing problem con Perform word segmentation and extract keywords (such as Textrank algorithm) to obtain keyword list Query syn , and input the keyword into the fine-tuned synonym language model (you can use a base model, such as llama), and get the synonym list of the keyword Query synon (Including keyword list Query syn ).
[0068] Step (3): Extract the entities of the query question based on the pre-built information extraction model to obtain an entity list.
[0069] In the specific implementation, the query question Query input by the user is extracted using the pre-established information extraction model to obtain the entity list Query entity .
[0070] Step (4): Based on the pre-built judgment model, the query condition judgment is performed on the query question to obtain the query condition result; wherein the query condition at least includes: the document type to which the query content belongs and the location of the query content in the document.
[0071] In the specific implementation, the query question Query input by the user is judged by using a pre-established judgment model (which can be based on a pre-trained model, such as the T5 model) to obtain the query condition result Query condition The query conditions at least include: the possible document type to which the query content may belong, and the possible location of the query content in the document, such as possible locations: abstract, keywords, introduction, background, methods, key, experiments, discussion, outlook, or possible document types: clinical trial, RCT, meta-analysis, systematic review, review, guideline, others, etc.
[0072] Step (5): Based on the synonym list of the keywords, the query condition results and the splicing problem, the first query logic and the second query logic are respectively constructed for the inverted index and the vector index of the medical subject literature library, and the inverted index query and the vector index query are respectively performed based on the first query logic and the second query logic to obtain the literature query results.
[0073] During specific implementation, a query is first performed on the medical subject literature library, including querying the inverted index of the medical subject literature library and querying the vector index of the medical subject literature library.
[0074] Query the inverted index of the medical subject literature library, and input the synonym list of the keyword Query synon And the query result Query condition , the logical relationship between keywords is logical OR, the logical relationship between query conditions is logical OR, and the logical relationship between keywords and query conditions is logical AND. The final first query logic is: And(Or(Query synon ), Or(Query condition )) Based on the first query logic, perform an inverted index query of the medical subject literature library to obtain the sorted topK (pre-set, such as 10) query results Search doc-inv .
[0075] Query the vector index of the medical subject literature library, and input the query as a splicing problem con and query condition results Query condition , the logical relationship between query conditions is logical OR, and the logical relationship between the splicing question and the query condition is logical AND, so the final second query logic is: And(Query con ,Or(Query condition)) Based on the second query logic, perform vector index query on the medical subject literature library and obtain the sorted topK (pre-set, for example, 10) query results. doc-vec In this embodiment, the query result Search doc-inv and search results doc-vec Together they constitute the literature search results.
[0076] Step (6): Construct a third query logic for the medical subject literature entity library based on the entity list and the query condition results, and perform a query on the medical subject literature entity library based on the third query logic to obtain an entity query result.
[0077] In the specific implementation, the query is performed on the medical subject literature entity library, and the input is the entity list Query entity , Query condition results condition , the logical relationship between entity lists is logical OR, the logical relationship between query conditions is logical OR, and the logical relationship between entity lists and query conditions is logical AND. The final third query logic is: And(Or(Query entity ), Or(Query condition )). Based on the third query logic, perform a query on the medical subject literature entity library to obtain the sorted topK (pre-set, such as 10) entity query results. entity (Including: document content obtained through document ID association, etc.).
[0078] Step (7): If the document type in the query condition result includes RCT, a fourth query logic for the medical discipline evidence library is constructed based on the entity list and the synonym list of the keyword, and the medical discipline evidence library is queried based on the fourth query logic to obtain the evidence query result.
[0079] In the specific implementation, the query condition result Query condition In the example, if the result of the RCT judgment in the document type information is true (indicated by 1), that is, the document type in the query condition result includes RCT, then the medical discipline evidence library is queried, and the input is the entity list Query entity 、Keyword synonym list Query synon , the logical relationship between entity list queries is logical OR, the logical relationship between keyword list queries is logical OR, and the logical relationship between entity list queries and keyword list queries is logical OR, then the final fourth query logic is: Or(Or(Query entity ), Or(Query synon)) Based on the fourth query logic, the medical discipline evidence database is searched to obtain the sorted topK (pre-set, such as 10) evidence query results Search evidence (Including: the names and content of each relevant field in the medical discipline evidence library, etc.).
[0080] Step (8): Input the query question, document query results, entity query results and evidence query results into a pre-built evaluation and scoring model to obtain a scoring result for each query result, and determine the second query result based on the scoring result.
[0081] In the specific implementation, for the query results obtained by the above query, each search result is evaluated and scored by the pre-built evaluation and scoring model, wherein the evaluation and scoring model consists of two sub-models Model q 、Model a Composition, the two model structures can be the same (such as siamese, etc.), but the training corpus is different. Each training corpus of the former is composed of (question 1: question 2) (the semantics of the two questions in the positive sample are similar, while those in the negative sample are different), and each training corpus of the latter is composed of (question: answer) (the content of the question and answer in the positive sample is matched, while that of the negative sample is not matched). The training goal of the former is to maximize the similarity index of similar (positive sample) questions, and the training goal of the latter is to maximize the correlation index between the positive question and its answer.
[0082] The input of the evaluation scoring model is the query question entered by the user, each query result mentioned above (including literature query results, entity query results and evidence query results), and the output scoring results are Score q (i.e. sub-model Model q Scoring results), Score a (i.e. sub-model Model a Scoring result of the query result), thereby obtaining the scoring result Score=W q *Score q +W a *Score a (Weight W q and W a can be determined in advance).
[0083] For each query result, you can sort them from large to small according to the above scoring results, and then intercept the results whose scoring is greater than or equal to the pre-set threshold L. min The first topL (pre-set) results are used as the second query result Search last If the second query result is lastIf the result is empty, step S103 is repeated. If the result is empty again, the user is directly fed back the result "unable to answer".
[0084] Case 2: If the query method is a method of querying the medical subject knowledge base corresponding to the subject field to which the query question belongs, then the medical subject knowledge base corresponding to the subject field to which the query question belongs is queried based on the query question to obtain the first query result.
[0085] Specifically, if the query mode judgment model determines that a correct answer can be obtained by simply querying the medical subject knowledge base corresponding to the subject area to which the query question belongs, the first query result is obtained by directly querying the medical subject knowledge base.
[0086] In one embodiment, the following methods may be used, including but not limited to:
[0087] First, the query question Query entered by the user is segmented using a word segmenter (such as Jieba) and keywords are extracted (such as using the Textrank algorithm) to obtain a keyword list Query syn1 , and input the keyword into the fine-tuned synonym language model to get the keyword synonym list Query synon1 (Including keyword list Query syn1 ). Secondly, for the query question Query input by the user, the pre-established information extraction model is used to extract entities and obtain the entity list Query entity1 Next, based on the synonym list of the keyword, Query synon1 and Entity List Query entity1 Construct query logic for medical subject literature library (based on keyword synonym list Query synon1 ), query logic for medical literature entity library (based on entity list Query entity1 ), and query logic for medical subject evidence base (keyword-based synonym list Query synon1 and Entity List Query entity1 Then, based on the query logic for the medical subject document library, a medical subject document library query is performed to obtain document query results, based on the query logic for the medical subject document entity library query is performed to obtain entity query results, and based on the query logic for the medical subject document evidence library query is performed to obtain evidence query results. Finally, the query question, document query results, entity query results, and evidence query results are input into a pre-built evaluation and scoring model (which may be the same as the evaluation and scoring model in the above example) to obtain a scoring result for each query result, and the query result is determined based on the scoring result.
[0088] In this embodiment, query logic is constructed, and queries are performed on the medical subject literature library, the medical subject literature entity library, and the medical subject evidence library. The query results obtained are merged and sorted, and the top M (pre-set) results are returned to the user. The specific implementation method is similar to the previous embodiment and will not be repeated here.
[0089] Case 3: If the query method is based on the interaction with the medical discipline large language model corresponding to the subject field to which the query question belongs, the query question is input into the medical discipline large language model corresponding to the subject field to which the query question belongs to obtain the first feedback result.
[0090] If the query judgment model determines that the correct answer can be obtained based on the medical discipline language model corresponding to the subject area to which the query question belongs, the query question is directly sent to the medical discipline language model to obtain the first feedback result Re.
[0091] In one embodiment, after obtaining the second query result in step S103, a first prompt for the medical discipline language model corresponding to the subject area of the query question can be constructed based on the second query result. The first prompt is then input into the medical discipline language model to obtain second feedback based on the second query result. Specifically, in this embodiment, the first prompt for interacting with the medical discipline language model includes: role, task, requirements, reference materials, related tips, etc.
[0092] In one embodiment, when constructing the first prompt content of the medical subject language model corresponding to the subject field to which the query question belongs based on the second query result, the following methods may be used, including but not limited to:
[0093] (1) Identify the role of the medical language model corresponding to the subject area of the query question. For example, you are an expert in the field of oncology.
[0094] (2) Clarify the task of this dialogue, including: listing the query questions input by the user (including question decomposition) and clarifying the task. For example, requiring the user to understand and respond to the knowledge in the subject area to which the query question belongs based on the user's query question and related reference materials.
[0095] (3) Clarify relevant requirements. For example, a word limit for responses, a requirement for factual answers—if you don’t know, just reply “I don’t know”, etc.
[0096] (4) List relevant information, including references related to the query question entered by the user.
[0097] In the specific implementation, the process of collating reference materials is as follows: First, for the second query result Search lastIf the query result corresponds to a document, the document content needs to be summarized according to the abstract extraction algorithm (such as clustering-based abstract extraction) to compress the content length so that the length does not exceed the limit (the allowed length of each reference (i.e., the preset length) is obtained based on the total length of the allowed references / the number of references). doc ); If the query result corresponds to pre-split content, and the content length does not exceed the above allowed length doc , no processing is performed; otherwise, compression is performed. If the search results correspond to evidence library content, the evidence content is combined into a natural language sentence using the format (evidence field: evidence content, ...). The query results are then sorted from highest to lowest according to their score, forming a list of references to serve the prompt.
[0098] (5) List relevant prompts (for example, keywords and entities related to the user's query question, etc.). Here, we mainly clarify the keyword list obtained in the early processing process. syn , and entity list query entity wait.
[0099] Thus, a complete first prompt is formed to interact with the medical subject language model corresponding to the subject area to which the query question belongs.
[0100] It should be noted that in the case where it is judged that "only a large language model of the medical discipline is needed to give a correct answer", the prompt content includes the first three items mentioned above.
[0101] The aforementioned step S105, i.e., performing relevance evaluation and scoring on the second feedback result, includes the following steps a1 to a3:
[0102] Step a1: Filter the second feedback result based on pre-set rules.
[0103] In one embodiment, the second feedback result Re 2th Content inspection and filtering are mainly completed according to pre-set rules. The rules restrict relatively clear knowledge: for example, in the field of AD, recommendations for AD drugs must be consistent with the guidelines, and discriminatory language and results must not appear. Content that does not comply with the rules will be handled accordingly.
[0104] Step a2: input the filtered second feedback result into the pre-built first relevance evaluation model to obtain the relevance evaluation score of the second feedback result.
[0105] In one embodiment, the filtered second feedback result is evaluated and scored. In this embodiment, the relevance evaluation score consists of two parts: on the one hand, it is given by the medical discipline language model that generates the second feedback result (by constructing a prompt consisting of roles, tasks - providing a relevance score between 0 and 100, user query questions, second feedback results and reference materials, etc., and sending it to the medical discipline language model to obtain the relevance evaluation score), specifically including: the relevance evaluation score Relevance between the second feedback result and the query question entered by the user mo-re-qu1 , and the relevance evaluation score between the second feedback result and the query result of the reference material listed in the prompt mo-re-se1 (wherein, the score is the weighted average of the relevance evaluation scores of the second feedback result and each query result), and then according to the pre-set weight (W mo-rq1 and W mo-rs1 ) The above two correlation evaluation scores are weighted and calculated to obtain the correlation evaluation score given by the medical discipline large language model:
[0106] Relevance model1 =W mo-rq1 *Relevance mo-re-qu1 +W mo-rs1 *Relevance mo-re-se1 .
[0107] On the other hand, the relevance evaluation score Relevance between the second feedback result and the query question input by the user is obtained by the pre-built first relevance evaluation model sub-model (which can be based on a pre-trained model, such as T5 tuning). s -re-qu1 , and the relevance evaluation score of the second feedback result and the query result of the reference material listed in the prompt s-re-se1 (Wherein, the score is a weighted average of the relevance evaluation scores of the second feedback result and each query result), and according to the pre-set weight (W s-rq1 and W s-rs1 ) The above two correlation evaluation scores are weighted and calculated to obtain the correlation evaluation score given by the first correlation evaluation model sub-model:
[0108] Relevance s1 =W s-rq1 *Relevance s-re-qu1 +W s-rs1 *Relevance s-re-se1 .
[0109] Finally, the relevance evaluation score of the second feedback result for the medical discipline large language model feedback is obtained:
[0110] Relevence1=(Relevance s1 +Relevance model1 ) / 2(During the evaluation process, obtain and record the second feedback result Re 2th Each sentence in the list and the corresponding reference number and content that exceeds the pre-set threshold).
[0111] Step a3: If the relevance evaluation score of the second feedback result is greater than or equal to the second threshold and less than the first threshold, the second feedback result is evaluated and scored for consistency; wherein the first threshold is greater than the second threshold; if the relevance evaluation score of the second feedback result is less than the second threshold, a third feedback result is obtained based on the medical discipline large language model, and a secondary relevance evaluation score is performed on the third feedback result; if the relevance evaluation score of the second feedback result is greater than or equal to the first threshold, the second feedback result is fed back to the user.
[0112] In one embodiment, the relevance evaluation score of the second feedback result given by the first relevance evaluation model and two pre-set thresholds (the first threshold Re max >Second threshold Re min ) is processed according to the situation. If the relevance evaluation score of the second feedback result is greater than or equal to the first threshold Re max , the second feedback result is directly fed back to the user; if the relevance evaluation score of the second feedback result is greater than or equal to the second threshold Re min But it is less than the first threshold Re max , then the second feedback result is evaluated and scored for consistency; if the relevance evaluation score of the second feedback result is less than the second threshold Re min , then the third feedback result is obtained based on the medical discipline language model, and the third feedback result is evaluated and scored for secondary relevance.
[0113] In a possible implementation, obtaining a third feedback result based on the medical discipline large language model and performing a secondary relevance evaluation and scoring on the third feedback result include the following steps b1 to b8:
[0114] Step b1: perform entity matching based on each entity in the entity list and the medical subject atlas library to obtain a matching entity list and a document set corresponding to the matching entity.
[0115] In the specific implementation, the matching method is used to query the entity list entityEach entity in the medical subject atlas is matched with the entity list to obtain a list of successfully matched entities and a collection of documents corresponding to the matching entities. entity If the matching entity list is empty, step S103 is repeated. If the matching entity list is empty again, a "cannot answer" result is directly fed back to the user.
[0116] Step b2: Based on the matching entity list, entities adjacent to the matching entity are extracted from the document collection to obtain entity triples, and the entity triples are filtered.
[0117] In the specific implementation, for the matching document set Docum entity ,Starting from the matching entity, extract the entities adjacent to the matching entity in each document to obtain entity triples, and filter the extracted entity triples to ensure that the entities in the entity triples must contain the corresponding entities.
[0118] Step b3: Repeat the entity triple extraction and filtering until a preset number of times is reached to obtain a first entity triple set.
[0119] In the specific implementation, the triple extraction is continued until the preset number (N) is reached, thereby obtaining the first entity triple set Triple entity .
[0120] Step b4: input the query question into the pre-built answer question pattern prediction model to predict the answer question pattern, and obtain a pattern library consisting of multiple answer question patterns; wherein the answer question pattern is the entity relationship of the answer question.
[0121] In the specific implementation, the answer question pattern prediction is performed. Specifically, the query question entered by the user is input into the pre-built answer question pattern prediction model (which can be fine-tuned based on the pre-trained model, such as the T5 model), and a pattern library consisting of multiple answer question patterns is obtained. query Among them, the question answering pattern is the entity relationship of answering questions in combination with key entities, such as (drug, treatment, disease).
[0122] Step b5: vectorize each entity relationship in the pattern library and each entity triple relationship in the first entity triple set based on the pre-built pre-trained model, perform similarity calculation on the entity relationships and entity triple relationships in the vectorized pattern library, and intercept a preset number of entity triples based on the similarity calculation results to obtain a second entity triple set.
[0123] In the specific implementation, for the pattern library Pattern queryEach entity relationship in the pattern (such as "treatment" in (drug, treatment, disease)) and the first entity triple set Triple entity The relationship between each entity triple in (such as "therapeutic drug" in (memantine, therapeutic drug, Alzheimer's disease)) is vectorized based on the pre-built pre-training model, and the similarity between the entity relationship of the vectorized pattern library and the relationship of the entity triple is calculated (for example, the cosine function is used to calculate the similarity), and sorted according to the size of the similarity score, intercepting the pre-selected number (P) of entity triplets to obtain the second entity triple set Triple filterd In this embodiment, in addition to similarity calculation, antonym calculation, implication calculation, etc. may also be performed according to the need of answering the question.
[0124] Step b6: construct a second prompt content of the medical discipline large language model based on the second entity triple set, and input the second prompt content into the medical discipline large language model to obtain a third feedback result.
[0125] In the specific implementation, the second prompt content prompt of the medical subject language model is constructed. The construction of the prompt is similar to the above steps, wherein the reference material here is the extracted second entity triple set Triple filterd , the triples are combined into natural language sentences in the form of (subject, relation, object). Specifically, the original text content associated with the triples can be included in the reference material as needed.
[0126] Then, the constructed second prompt content prompt is sent to the medical discipline language model, and the third feedback result Re returned by the medical discipline language model is obtained. 3th .
[0127] Step b7: filtering the third feedback result based on a pre-set rule, and inputting the filtered third feedback result into a pre-built second relevance evaluation model to obtain a relevance evaluation score of the third feedback result.
[0128] In specific implementation, the third feedback result is first checked and filtered based on pre-set rules, and then the filtered third feedback result is evaluated and scored based on the pre-built second relevance evaluation model. The evaluation and scoring results are composed of two parts: on the one hand, it is given by the medical discipline language model that generates the third feedback result (by providing a prompt consisting of roles, tasks - providing a relevance score between 0 and 100, user query questions, third feedback results and reference materials, etc., and sending it to the medical discipline language model to obtain the relevance evaluation score), specifically including: the relevance evaluation score Relevance for the third feedback result and the query question entered by the user mo-re-qu2 , and the relevance evaluation score Relevance between the third feedback result and the query result of the reference material listed in the prompt (determined according to the second entity triple set) mo-re-se2 (Wherein, the score is the weighted average of the relevance evaluation scores of the third feedback result and each query result), and then according to the pre-set weight (W mo-rq2 and W mo-rs2 ) The above two correlation evaluation scores are weighted and calculated to obtain the correlation evaluation score given by the medical discipline large language model:
[0129] Relevance model2 =W mo-rq2 *Relevance mo-re-qu2 +W mo-rs2 *Relevance mo-re-se2 .
[0130] On the other hand, the relevance evaluation score Relevance between the third feedback result and the query question input by the user is obtained by the pre-built second relevance evaluation model sub-model s-re-qu2 , and the relevance evaluation score Relevance between the third feedback result and the query result of the reference material listed in the prompt (determined according to the second entity triple set) s-re-se2 (Wherein, the score is the weighted average of the relevance evaluation scores of the third feedback result and each query result), and according to the pre-set weight (W s-rq2 and W s-rs2 ) The above two correlation evaluation scores are weighted and calculated to obtain the correlation evaluation score given by the correlation evaluation model:
[0131] Relevance s2 =W s-rq2 *Relevance s-re-qu2 +W s-rs2 *Relevance s-re-se2 .
[0132] Finally, the relevance evaluation score of the third feedback result for the medical discipline large language model feedback is obtained:
[0133] Relevence2=(Relevance s2 +Relevance model2 ) / 2(In the process, record the third feedback result Re 3th Each sentence in the list and the corresponding reference number and content that exceeds the pre-set threshold).
[0134] Step b8: If the relevance evaluation score of the third feedback result is greater than or equal to the first threshold, the third feedback result is fed back to the user; if the relevance evaluation score of the third feedback result is greater than or equal to the second threshold and less than the first threshold, the third feedback result is evaluated and scored for consistency; if the relevance evaluation score of the third feedback result is less than the second threshold, the first prompt content of the medical discipline large language model is reconstructed to obtain a new second feedback result, and the new second feedback result is evaluated and scored for relevance; if the relevance evaluation score of the new second feedback result is less than the second threshold, the second prompt content of the medical discipline large language model is reconstructed to obtain a new third feedback result; if the relevance evaluation score of the new third feedback result is less than the second threshold, the query question is preprocessed again, and based on the preprocessed query question, the medical discipline knowledge base corresponding to the subject field to which the query question belongs is queried and the medical discipline large language model is interacted with to obtain a new second feedback result.
[0135] In specific implementation, according to the scoring result of the third feedback result given by the second correlation evaluation model and the two thresholds set in advance (the first threshold Re max >Second threshold Re min If the relevance evaluation score of the third feedback result is greater than or equal to the first threshold Re max , the third feedback result is directly fed back to the user; if the relevance evaluation score of the third feedback result is greater than or equal to the second threshold Re min But it is less than the first threshold Re max , then the consistency evaluation score of the third feedback result is performed; if the relevance evaluation score of the third feedback result is less than the second threshold Re min, then the first prompt content of the medical discipline language model is returned to be reconstructed (at this time, the relevant prompt also includes the evaluation and scoring of the previous third feedback result content), a new second feedback result is obtained, and the new second feedback result is evaluated and scored for relevance. If the relevance evaluation score of the new second feedback result is less than the second threshold, the second prompt content of the medical discipline language model is reconstructed to obtain a new third feedback result, and the new third feedback result is evaluated and scored for a second time; if the relevance evaluation scores of the two (configurable) third feedback results are both less than the second threshold Re min , then return to step (1) to re-preprocess the query question (at this time, the relevant prompts also include the evaluation and scoring of the third feedback result content before), if the relevance evaluation scores of the third feedback results returned three times (configurable) are all less than the second threshold Re min , it will directly return to the user "cannot answer".
[0136] In a possible implementation, performing consistency evaluation and scoring on the second feedback result or the third feedback result includes the following steps c1 to c4:
[0137] Step c1: construct a key summary prompt based on the second feedback result or the third feedback result, and input the key summary prompt into the medical discipline large language model to obtain a key summary result of the second feedback result or the third feedback result.
[0138] In specific implementation, the second feedback result Re for the query question Query input by the user is obtained according to the above steps. 2th Or the third feedback result Re 3th , build a key summary prompt (including: role, task-requirement combined with user query questions, reference materials, second feedback results or third feedback results, etc. to generate key summary), and send the key summary prompt to the medical discipline language model to generate a key summary for the second feedback result Re 2th Or the third feedback result Re 3th Key points of the summary results sum .
[0139] Step c2: scoring the key summary results of the second feedback result or the third feedback result based on the pre-built consistency evaluation model to obtain a consistency evaluation score of the key summary results of the second feedback result or the third feedback result.
[0140] In specific implementation, the key points of the second feedback result or the third feedback result of the model feedback based on the pre-built consistency evaluation model are summarized as follows: sumThe consistency evaluation score is composed of two parts. On the one hand, it is given by the medical discipline language model that generates the key summary results. By constructing a prompt consisting of roles, tasks - providing a relevance score between 0 and 100, user query questions, second feedback results or third feedback results and key summary results, and sending it to the medical discipline language model, the consistency evaluation score of the key summary results is obtained, that is, the consistency evaluation score of the key summary results and the second feedback results Re 2th Or the third feedback result Re 3th Relevance of consistency evaluation score model -sum On the other hand, the second feedback result Re is obtained by the pre-built consistency evaluation model sub-model (which can be obtained based on the pre-trained model tuning, such as the T5 model) 2th Or the third feedback result Re 3th Relevance of consistency evaluation score with key summary results s-sum Finally, we get the consistency evaluation score of the summary results of the key points of the large language model for medical disciplines: Relevance sum =(Relevance s-sum +Relevance model-sum ) / 2.
[0141] Step c3: If the consistency evaluation score of the key summary result of the second feedback result is greater than or equal to the third threshold, the second feedback result is fed back to the user; if the consistency evaluation score of the key summary result of the second feedback result is less than the third threshold, the first prompt content of the medical discipline large language model is reconstructed to obtain a new second feedback result. If the relevance evaluation score of the new second feedback result is greater than or equal to the second threshold and less than the first threshold, the new second feedback result is scored for consistency evaluation. If the consistency evaluation score of the key summary result of the new second feedback result is less than the third threshold, the query question is preprocessed again, and based on the preprocessed query question, the medical discipline knowledge base corresponding to the subject field to which the query question belongs is queried and the medical discipline large language model is interacted to obtain a new second feedback result.
[0142] In the specific implementation, the consistency evaluation score of the summary result of the final second feedback result and the pre-set third threshold (Re sum-min ) is processed according to the situation: if the consistency evaluation score of the key summary result of the second feedback result is greater than or equal to the third threshold Re sum-min, the second feedback result is directly fed back to the user; otherwise, the first prompt content of the reconstructed medical discipline language model is returned (at this time, the relevant prompt also includes the consistency evaluation score for the previous feedback result), and a new second feedback result is obtained. If the relevance evaluation score of the new second feedback result is greater than or equal to the second threshold and less than the first threshold, the consistency evaluation score of the new second feedback result is performed. If the consistency evaluation scores of the key summary results of the second feedback results returned twice (configurable) are both less than the third threshold Re sum-min , then return to step (1) to re-preprocess the query question (at this time, the relevant prompts also include the consistency evaluation score for the previous feedback results). If the consistency evaluation scores of the key summary results of the new second feedback results returned three times (configurable) are all less than the third threshold Re sum-min , it will directly return to the user "cannot answer".
[0143] Step c4: if the consistency evaluation score of the key summary result of the third feedback result is greater than or equal to the third threshold, the third feedback result is fed back to the user; if the consistency evaluation score of the key summary result of the third feedback result is less than the third threshold, the first prompt content of the medical discipline large language model is reconstructed to obtain a new second feedback result; if the relevance evaluation score of the new second feedback result is less than the second threshold, the second prompt content of the medical discipline large language model is reconstructed to obtain a new third feedback result; if the relevance evaluation score of the new third feedback result is greater than or equal to the second threshold and less than the first threshold, the new third feedback result is scored for consistency evaluation; if the consistency evaluation score of the key summary result of the new third feedback result is less than the third threshold, the query question is preprocessed again, and based on the preprocessed query question, the medical discipline knowledge base corresponding to the subject field to which the query question belongs is queried and the medical discipline large language model is interacted to obtain a new second feedback result.
[0144] In the specific implementation, the consistency evaluation score of the results is summarized according to the key points of the third feedback result and the pre-set third threshold (Re sum-min ) is processed according to the situation: if the consistency evaluation score of the key summary result of the third feedback result is greater than or equal to the third threshold Re sum-min, the third feedback result is directly fed back to the user; otherwise, the first prompt content of the reconstructed medical discipline language model is returned (at this time, the relevant prompt also includes the consistency evaluation score for the previous feedback result), and a new second feedback result is obtained, and the new second feedback result is evaluated and scored for relevance; if the relevance evaluation score of the new second feedback result is less than the second threshold, the second prompt content of the medical discipline language model is reconstructed to obtain a new third feedback result, and the new third feedback result is evaluated and scored for secondary relevance. If the relevance evaluation score of the new third feedback result is greater than or equal to the second threshold and less than the first threshold, the new third feedback result is evaluated and scored for consistency. If the consistency evaluation scores of the key summary results of the third feedback results returned twice (configurable) are both less than the third threshold Re sum-min , then return to step (1) to re-preprocess the query question (at this time, the relevant prompts also include the consistency evaluation score for the previous feedback results). If the consistency evaluation score of the third feedback result returned three times (configurable) is less than the third threshold Re sum-min , it will directly return to the user "cannot answer".
[0145] Furthermore, for the second feedback result or the third feedback result that has passed the evaluation score, a reference mark is generated in the second feedback result or the third feedback result according to each sentence and the corresponding reference material number in the recorded second feedback result or the third feedback result, and the reference material content and the reference material number (i.e., the document ID) are listed after the second feedback result or the third feedback result, thereby generating the final feedback result Re last , to provide feedback to users.
[0146] The above method provided by the embodiment of the present invention has the following advantages and positive effects compared with the existing technology: (1) By constructing a real-time updated, multi-dimensional medical knowledge base including documents, entities, graphs, evidence, indexes, etc., it ensures that the medical knowledge in different disciplines is multi-dimensional, cross-granular, comprehensive and up-to-date; (2) Based on the knowledge structure of document tags, entities, graphs, evidence, etc. of the multi-dimensional knowledge base, and a fine-tuned large language model, a combination of knowledge and data-driven methods is adopted to ensure the automation and human-like effect of the medical knowledge service process in different disciplines; (3) Pre-judgment of the user's query query (including query method judgment), matching the complexity of the question and the answer method, ensuring the accuracy and effect on the basis of ensuring the efficiency of knowledge service; (4) by optimizing the user's query, optimizing the knowledge query, and optimizing the prompt construction and feedback results, the accuracy of the reply feedback results is further guaranteed; (5) For more complex questions, the present invention can be combined with the constructed medical discipline atlas library, and adopt a method combined with the prediction of the answer question pattern to extract more accurate and specific reference materials for feedback result generation, further ensuring the accuracy of the query result; (6) by adopting the method of optimizing the user's query and query The method of evaluating and scoring the correlation between the query results, the query questions of the user and the feedback results, the query results and the feedback results, and the feedback results and the key summary results ensures the self-consistency from the user's query to the feedback results and the key summary results. The method of giving the feedback result of "unable to answer" in the case of low scores is adopted to ensure the accuracy of the reply results. (7) Through step-by-step (link cycle-based) planning (based on the decomposition and processing of the problem, etc.), optimization (based on the evaluation of previous results), evaluation of interaction results and evaluation-based routing selection methods, the detailed decomposition and self-optimization of the knowledge service process are promoted, while maintaining On the premise of proving the efficiency of the process, the quality and effect of the knowledge service process are also guaranteed; (8) by adding reference marks and reference materials to the content of the feedback results, the interpretability of the feedback results is enhanced, and the basis of the feedback results is explained; (9) combined with the knowledge base and query process of the graph structure, the accuracy of answers to details and more advanced focus questions can be greatly improved; (10) Based on the above process, the present invention can be widely used in different medical disciplines to obtain literature, obtain evidence, read literature, obtain research ideas, research design, data processing, and generate results, etc., reducing the workload and complexity of researchers in the medical disciplines and improving efficiency, quality and effect.
[0147] Regarding the medical subject knowledge service method provided in the above embodiment, the present invention also provides a medical subject knowledge service device, see Figure 2The schematic diagram of the structure of a medical subject knowledge service device shown in FIG. shows that the device mainly includes the following parts:
[0148] The medical subject knowledge base construction module 201 is used to obtain medical literature and construct medical subject knowledge bases corresponding to various subject areas based on the medical literature; wherein the medical subject knowledge base includes: a medical subject literature library, a medical subject literature entity library, a medical subject evidence library, a medical subject atlas library, and a medical subject index library;
[0149] The pre-judgment module 202 is used to obtain the query question input by the user and pre-judgment the query question to determine the subject area and query method to which the query question belongs;
[0150] The query and interaction module 203 is configured to query the medical subject knowledge base corresponding to the subject area to which the query question belongs and / or interact with the medical subject large language model based on the query method to obtain a result of the query question; wherein the result of the query question includes one of the following: a first query result, a first feedback result, a second query result, and a second feedback result based on the second query result;
[0151] The feedback and evaluation module 204 is used to feed back the first query result or the first feedback result to the user if a first query result or a first feedback result is obtained; if a second feedback result is obtained, perform a relevance evaluation and scoring on the second feedback result, and when the score result of the relevance evaluation score of the second feedback result meets the first preset condition, perform a consistency evaluation and scoring on the second feedback result; when the score result of the relevance evaluation score of the second feedback result meets the second preset condition, obtain a third feedback result based on the medical discipline large language model, and perform a secondary relevance evaluation and scoring on the third feedback result; when the score result of the secondary relevance evaluation score of the third feedback result meets the first preset condition, perform a consistency evaluation and scoring on the third feedback result; when the score result of the consistency evaluation score of the second feedback result or the third feedback result meets the third preset condition, feed back the second feedback result or the third feedback result to the user.
[0152] The above-mentioned medical discipline knowledge service device provided by the embodiment of the present invention can ensure the comprehensiveness and timeliness of knowledge in different disciplines by constructing a medical discipline knowledge base in different disciplines that is updated in real time and includes a medical discipline literature base, a medical discipline literature entity base, a medical discipline evidence base, a medical discipline atlas base, and a medical discipline index base; by pre-judging the query questions input by the user, determining the discipline field and query method to which it belongs, and then querying and / or interacting based on the medical discipline knowledge base and / or medical discipline large language model corresponding to the discipline field to which it belongs, it can improve the accuracy and quality of the feedback results on the basis of improving service efficiency; finally, by performing correlation evaluation scoring, secondary correlation evaluation scoring and consistency evaluation scoring on the query results and feedback results, the results that meet the requirements are fed back to the user, further improving the accuracy and quality of obtaining medical discipline-related knowledge.
[0153] It should be noted that the implementation principle and technical effects of the device provided in the embodiment of the present invention are the same as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.
[0154] An embodiment of the present invention further provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above embodiments.
[0155] Figure 3 This is a structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 30, a memory 31, a bus 32 and a communication interface 33, wherein the processor 30, the communication interface 33 and the memory 31 are connected via the bus 32; the processor 30 is used to execute an executable module stored in the memory 31, such as a computer program.
[0156] The memory 31 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage. The system network element communicates with at least one other network element via at least one communication interface 33 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.
[0157] The bus 32 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 3 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0158] Among them, the memory 31 is used to store programs, and the processor 30 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 30 or implemented by the processor 30.
[0159] The processor 30 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits or software instructions in the processor 30. The processor 30 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or the like. The storage medium is located in the memory 31 , and the processor 30 reads the information in the memory 31 and completes the steps of the above method in combination with its hardware.
[0160] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the previous method embodiment. The specific implementation can be referred to the previous method embodiment and will not be repeated here.
[0161] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0162] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A medical knowledge service method, characterized in that: include: Acquire medical literature and construct medical subject knowledge bases corresponding to various subject areas based on the medical literature; wherein the medical subject knowledge bases include: a medical subject literature base, a medical subject literature entity base, a medical subject evidence base, a medical subject atlas base, and a medical subject index base; Obtaining a query question input by a user, and pre-judging the query question to determine the subject area and query method to which the query question belongs; Based on the query method, a medical subject knowledge base corresponding to the subject area to which the query question belongs is queried and / or a medical subject large language model is interacted to obtain a result of the query question; wherein the result of the query question includes one of the following: a first query result, a first feedback result, a second query result, and a second feedback result based on the second query result; If the first query result or the first feedback result is obtained, feeding back the first query result or the first feedback result to the user; If the second feedback result is obtained, a relevance evaluation score is performed on the second feedback result, and when the relevance evaluation score of the second feedback result meets the first preset condition, a consistency evaluation score is performed on the second feedback result. When the relevance evaluation score of the second feedback result meets the second preset condition, a third feedback result is obtained based on the medical discipline large language model, and a secondary relevance evaluation score is performed on the third feedback result. When the scoring result of the secondary relevance evaluation of the third feedback result meets the first preset condition, performing consistency evaluation and scoring on the third feedback result; When the consistency evaluation score of the second feedback result or the third feedback result meets a third preset condition, feeding back the second feedback result or the third feedback result to the user; A medical subject knowledge base corresponding to each subject field is constructed based on the medical documents, including: classifying the medical documents according to subject fields to obtain initial medical subject document bases in different subject fields; for each document in the initial medical subject document base, processing the document based on a pre-built metadata extraction model, knowledge rules, classification model and information extraction model, and constructing a medical subject document base according to the content and metadata corresponding to each tag of each document; for each document in the medical subject document base, extracting entities in each document based on a pre-built entity information extraction model to construct a medical subject document entity base; for each document in the medical subject document base, constructing a relationship map between documents and entities based on the documents and the medical subject document entity base, and constructing a three-dimensional map based on the relationships between entities in documents in the medical subject document entity base. Tuple map, and construct a medical discipline map library based on the relationship map and the triple map; for each document in the medical discipline document library, if the document type tag content of the document includes RCT, then extract the evidence information in each document based on the pre-built evidence extraction model to construct a medical discipline evidence library; wherein, the evidence information is information related to RCT; construct sub-indexes of the medical discipline document library, the medical discipline document entity library and the medical discipline evidence library respectively, and integrate and associate the medical discipline document library, the medical discipline document entity library and the medical discipline evidence library based on the association ID between the data of the medical discipline document library, the medical discipline document entity library and the medical discipline evidence library to construct an overall index to obtain a medical discipline index library; wherein, the sub-index and the overall index include: an inverted index and a vector index; Among them, for each of the sub-indexes and the overall index, an inverted index and a vector index are constructed respectively; in the process of constructing the vector index, the long text is first segmented, and then the segmented text is vectorized and then stored in the library respectively; in the process of constructing the vector index, for each segmented text block, based on the fine-tuned large model, a prompt method is used to input the entire document and the segmented text block into the large model to obtain a context summary for the segmented text block, and then the obtained context summary and the segmented text block are connected to obtain the input text for constructing the index.
2. The method according to claim 1, characterized in that The query methods include: a method based on querying the medical subject knowledge base corresponding to the subject field to which the query question belongs, a method based on interacting with the medical subject large language model corresponding to the subject field to which the query question belongs, and a method based on a combination of querying the medical subject knowledge base corresponding to the subject field to which the query question belongs and interacting with the medical subject large language model; Based on the query method, a medical subject knowledge base corresponding to the subject area to which the query question belongs is queried and / or a medical subject large language model is interacted to obtain a result of the query question, including: If the query method is a method of querying the medical subject knowledge base corresponding to the subject field to which the query question belongs, then querying the medical subject knowledge base corresponding to the subject field to which the query question belongs based on the query question to obtain a first query result; If the query method is a method of interacting with a large medical subject language model corresponding to the subject field to which the query question belongs, inputting the query question into the large medical subject language model corresponding to the subject field to which the query question belongs to obtain a first feedback result; If the query method is a method of interactively combining a query on the medical subject knowledge base corresponding to the subject field to which the query question belongs and a medical subject large language model, the query question is preprocessed, and a query on the medical subject knowledge base corresponding to the subject field to which the query question belongs is performed based on the preprocessed query question to obtain a second query result, and based on the second query result, a first prompt content of the medical subject large language model corresponding to the subject field to which the query question belongs is constructed, and the first prompt content is input into the medical subject large language model to obtain a second feedback result based on the second query result.
3. The method according to claim 2, characterized in that Preprocessing the query question, and querying the medical subject knowledge base corresponding to the subject field to which the query question belongs based on the preprocessed query question to obtain a second query result, including: Inputting the query question into a medical subject language model corresponding to the subject field to which the query question belongs to perform problem decomposition to obtain a decomposed question, and concatenating the decomposed question with the query question to obtain a concatenated question; Performing word segmentation and keyword extraction on the splicing problem to obtain a keyword list, and inputting the keyword list into a pre-built synonym language model to obtain a synonym list of the keywords; Extracting entities of the query question based on a pre-built information extraction model to obtain an entity list; Performing query condition judgment on the query question based on a pre-built judgment model to obtain a query condition result; wherein the query condition at least includes: the document type to which the query content belongs and the location of the query content in the document; Based on the synonym list of the keyword, the query condition result and the splicing problem, a first query logic and a second query logic are respectively constructed for the inverted index and the vector index of the medical subject literature library, and an inverted index query and a vector index query are respectively performed based on the first query logic and the second query logic to obtain a literature query result; Constructing a third query logic for the medical subject literature entity library based on the entity list and the query condition result, and performing a query on the medical subject literature entity library based on the third query logic to obtain an entity query result; If the document type in the query condition result includes RCT, constructing a fourth query logic for the medical discipline evidence library based on the entity list and the synonym list of the keyword, and performing a query on the medical discipline evidence library based on the fourth query logic to obtain an evidence query result; The query question, the document query result, the entity query result and the evidence query result are input into a pre-built evaluation and scoring model to obtain a scoring result for each query result, and a second query result is determined based on the scoring result.
4. The method according to claim 3, characterized in that The second feedback result is evaluated and scored for relevance, including: Filtering the second feedback result based on pre-set rules; Inputting the filtered second feedback result into the pre-built first relevance evaluation model to obtain a relevance evaluation score of the second feedback result; If the relevance evaluation score of the second feedback result is greater than or equal to a second threshold and less than a first threshold, then performing a consistency evaluation score on the second feedback result; wherein the first threshold is greater than the second threshold; If the relevance evaluation score of the second feedback result is less than the second threshold, obtaining a third feedback result based on the medical discipline large language model, and performing a secondary relevance evaluation and scoring on the third feedback result; If the relevance evaluation score of the second feedback result is greater than or equal to the first threshold, the second feedback result is fed back to the user.
5. The method according to claim 4, characterized in that Obtaining a third feedback result based on the medical discipline large language model, and performing a secondary relevance evaluation and scoring on the third feedback result, including: Perform entity matching based on each entity in the entity list and the medical subject atlas library to obtain a matching entity list and a document set corresponding to the matching entity; Based on the matching entity list, entities adjacent to the matching entity are extracted from the document collection to obtain entity triples, and the entity triples are filtered; Repeating entity triple extraction and filtering until a preset number of times is reached to obtain a first entity triple set; Input the query question into a pre-built answer question pattern prediction model to predict the answer question pattern, and obtain a pattern library consisting of multiple answer question patterns; wherein the answer question pattern is the entity relationship of the answer question; Vectorizing each entity relationship in the pattern library and each entity triple relationship in the first entity triple set based on a pre-built pre-trained model, performing similarity calculation on the vectorized entity relationships in the pattern library and the entity triple relationships, and intercepting a preset number of entity triples based on the similarity calculation results to obtain a second entity triple set; Constructing second prompt content of the medical discipline large language model based on the second entity triple set, and inputting the second prompt content into the medical discipline large language model to obtain a third feedback result; Filtering the third feedback result based on a pre-set rule, and inputting the filtered third feedback result into a pre-built second relevance evaluation model to obtain a relevance evaluation score of the third feedback result; If the relevance evaluation score of the third feedback result is greater than or equal to the first threshold, feeding back the third feedback result to the user; If the relevance evaluation score of the third feedback result is greater than or equal to the second threshold and less than the first threshold, performing a consistency evaluation and scoring on the third feedback result; If the relevance evaluation score of the third feedback result is less than the second threshold, the first prompt content of the medical discipline large language model is reconstructed to obtain a new second feedback result, and the new second feedback result is evaluated and scored for relevance. If the relevance evaluation score of the new second feedback result is less than the second threshold, the second prompt content of the medical discipline large language model is reconstructed to obtain a new third feedback result. If the relevance evaluation score of the new third feedback result is less than the second threshold, the query question is preprocessed again, and based on the preprocessed query question, the medical discipline knowledge base corresponding to the subject field to which the query question belongs is queried and interacted with the medical discipline large language model to obtain a new second feedback result.
6. The method according to claim 4 or 5, characterized in that Conducting consistency evaluation and scoring on the second feedback result or the third feedback result, including: Constructing a key summary prompt based on the second feedback result or the third feedback result, and inputting the key summary prompt into the medical discipline large language model to obtain a key summary result of the second feedback result or the third feedback result; Scoring the key summary results of the second feedback result or the third feedback result based on a pre-built consistency evaluation model to obtain a consistency evaluation score of the key summary results of the second feedback result or the third feedback result; If the consistency evaluation score of the key summary result of the second feedback result is greater than or equal to the third threshold, the second feedback result is fed back to the user; if the consistency evaluation score of the key summary result of the second feedback result is less than the third threshold, the first prompt content of the medical discipline large language model is reconstructed to obtain a new second feedback result, if the relevance evaluation score of the new second feedback result is greater than or equal to the second threshold and less than the first threshold, the new second feedback result is scored for consistency evaluation, if the consistency evaluation score of the key summary result of the new second feedback result is less than the third threshold, the query question is re-preprocessed, and based on the preprocessed query question, the medical discipline knowledge base corresponding to the subject field to which the query question belongs is queried and the medical discipline large language model is interacted to obtain a new second feedback result; If the consistency evaluation score of the key point summary result of the third feedback result is greater than or equal to the third threshold, the third feedback result is fed back to the user; if the consistency evaluation score of the key point summary result of the third feedback result is less than the third threshold, the first prompt content of the medical discipline large language model is reconstructed to obtain a new second feedback result. If the relevance evaluation score of the new second feedback result is less than the second threshold, the second prompt content of the medical discipline large language model is reconstructed to obtain a new third feedback result. If the relevance evaluation score of the new third feedback result is greater than or equal to the second threshold and less than the first threshold, the new third feedback result is scored for consistency evaluation. If the consistency evaluation score of the key point summary result of the new third feedback result is less than the third threshold, the query question is preprocessed again, and based on the preprocessed query question, the medical discipline knowledge base corresponding to the subject field to which the query question belongs is queried and the medical discipline large language model is interacted to obtain a new second feedback result.
7. A medical knowledge service device, characterized in that: include: A medical subject knowledge base construction module is used to obtain medical literature and construct medical subject knowledge bases corresponding to various subject areas based on the medical literature; wherein the medical subject knowledge base includes: a medical subject literature library, a medical subject literature entity library, a medical subject evidence library, a medical subject atlas library, and a medical subject index library; A pre-judgment module is used to obtain a query question input by a user and pre-judgment the query question to determine the subject area and query method to which the query question belongs; A query and interaction module, configured to query the medical subject knowledge base corresponding to the subject area to which the query question belongs and / or interact with the medical subject large language model based on the query method to obtain a result of the query question; wherein the result of the query question includes one of the following: a first query result, a first feedback result, a second query result, and a second feedback result based on the second query result; A feedback and evaluation module, configured to, if a first query result or a first feedback result is obtained, feed back the first query result or the first feedback result to the user; if a second feedback result is obtained, perform a relevance evaluation and scoring on the second feedback result, and when the score result of the relevance evaluation score of the second feedback result meets the first preset condition, perform a consistency evaluation and scoring on the second feedback result; when the score result of the relevance evaluation score of the second feedback result meets the second preset condition, obtain a third feedback result based on the medical discipline large language model, and perform a secondary relevance evaluation and scoring on the third feedback result; when the score result of the secondary relevance evaluation score of the third feedback result meets the first preset condition, perform a consistency evaluation and scoring on the third feedback result; when the score result of the consistency evaluation score of the second feedback result or the third feedback result meets the third preset condition, feed back the second feedback result or the third feedback result to the user; The medical discipline knowledge base construction module is specifically used to: classify the medical documents according to subject fields to obtain initial medical discipline document libraries in different subject fields; for each document in the initial medical discipline document library, process the document based on a pre-built metadata extraction model, knowledge rules, classification model and information extraction model, and construct a medical discipline document library according to the content and metadata corresponding to each tag of each document; for each document in the medical discipline document library, extract the entities in each document based on a pre-built entity information extraction model to construct a medical discipline document entity library; for each document in the medical discipline document library, construct a relationship map between documents and entities based on the documents and the medical discipline document entity library, and construct a triple map based on the relationship between entities in documents in the medical discipline document entity library, and A medical discipline graph library is constructed based on the relationship graph and the triple graph; for each document in the medical discipline document library, if the document type tag content of the document includes RCT, evidence information in each document is extracted based on a pre-constructed evidence extraction model to construct a medical discipline evidence library; wherein the evidence information is information related to RCT; sub-indexes of the medical discipline document library, the medical discipline document entity library and the medical discipline evidence library are constructed respectively, and based on the association ID between the data of the medical discipline document library, the medical discipline document entity library and the medical discipline evidence library, the medical discipline document entity library and the medical discipline evidence library are integrated and associated to construct an overall index to obtain a medical discipline index library; wherein the sub-index and the overall index include: an inverted index and a vector index; Among them, for each of the sub-indexes and the overall index, an inverted index and a vector index are constructed respectively; in the process of constructing the vector index, the long text is first segmented, and then the segmented text is vectorized and then stored in the library respectively; in the process of constructing the vector index, for each segmented text block, based on the fine-tuned large model, a prompt method is used to input the entire document and the segmented text block into the large model to obtain a context summary for the segmented text block, and then the obtained context summary and the segmented text block are connected to obtain the input text for constructing the index.
8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are executed.
Citation Information
Patent Citations
Question query method and device, electronic equipment and computer readable storage medium
CN118689967A
Knowledge question and answer accuracy improving method based on semantic elements
CN119202209A