Method and device for establishing clinical auxiliary diagnosis and treatment professional large model, and evaluation method
By deconstructing professional literature into knowledge points and mapping them into a knowledge space, and combining this with a questioning paradigm to establish a large-scale clinical auxiliary diagnosis and treatment model, the problem of large language models lacking authoritative support and inaccurate positioning in clinical auxiliary diagnosis and treatment is solved, and precise retrieval and recommendation based on evidence-based medicine is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV QILU HOSPITAL
- Filing Date
- 2024-09-13
- Publication Date
- 2026-06-05
AI Technical Summary
Existing large language models struggle to adhere to evidence-based medicine principles in clinical auxiliary diagnosis and treatment, providing answers that lack support from authoritative medical literature, and failing to accurately locate content and provide relevant recommendations.
By deconstructing professional literature into knowledge points and mapping them to a knowledge space, a questioning paradigm is constructed to constrain the coordinate mapping relationship between natural language expression and the knowledge space. This leads to the establishment of a large-scale professional model for clinical auxiliary diagnosis and treatment, enabling precise location retrieval and relevant recommendations of question content.
It enables the location of answer content and related recommendations based on authoritative medical literature, meets the requirements of evidence-based medicine, and improves the usability and accuracy of clinical auxiliary diagnosis and treatment systems.
Smart Images

Figure CN119204196B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology of large language models, and in particular to a method, evaluation method and device for establishing a large professional model for clinical auxiliary diagnosis and treatment. Background Technology
[0002] Large Language Models (LLMs) are language models based on deep neural network architectures. By learning from large amounts of text data, they predict the probability of the next word or phrase, enabling computers to better understand and generate human language. Applying large model technology to the field of clinical auxiliary diagnosis and treatment can create specialized large model systems that provide clinical diagnostic knowledge and question answering, offering convenient medical consultations for patients and providing assistance to doctors in clinical diagnosis and treatment.
[0003] Due to the scientific and standardized nature of clinical diagnosis and treatment, when applying large-scale model technology to the field of clinical auxiliary diagnosis and treatment to form a professional large-scale model system that provides clinical diagnosis and treatment knowledge Q&A, the following issues need to be addressed: 1) Adhering to the principle of "evidence-based medicine": Measures must be taken to address the illusion problem of large-scale models (i.e., the text generated by the model does not follow the original text or does not conform to the facts). Clinical diagnosis and treatment follow the principle of "evidence-based medicine," and all diagnostic and treatment actions must be based on solid evidence. This evidence is authoritative medical literature, not text predicted and generated by large-scale model algorithms. Therefore, professional large-scale models used for clinical auxiliary diagnosis and treatment must provide the source of the content they rely on in the medical literature in their generated answers to meet the requirements of "evidence-based medicine." Currently, large language models based on text prediction are difficult to apply clinically because the automatically generated content cannot provide sources. 2) Precise Content Location and Related Recommendations: Whether using keyword-based full-text search or LLM-based vectorized similarity matching, both methods suffer from limitations. Firstly, the lack of strict paradigm constraints in querying makes it difficult to pinpoint the specific question. Secondly, the search indexes built on fine-grained keywords or synonyms result in overly broad search results, failing to accurately locate corresponding content within specialized literature. After achieving precise content location, users further desire to explore similar content, necessitating the system's ability to provide relevant recommendations.
[0004] Therefore, how to establish a large-scale professional model for clinical auxiliary diagnosis and treatment is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] In view of the above problems, embodiments of this application provide a method, evaluation method and apparatus for establishing a large-scale clinical auxiliary diagnosis and treatment model, so as to overcome the above problems or at least partially solve the above problems.
[0006] A first aspect of this application discloses a method for establishing a large-scale clinical auxiliary diagnosis and treatment model, the method comprising:
[0007] Based on the knowledge characteristics of clinical diagnosis and treatment, professional literature is deconstructed into knowledge points, and these knowledge points are mapped to a knowledge space;
[0008] A questioning paradigm is constructed, which is used to constrain the coordinate mapping relationship between the natural language expression of the question and the knowledge space;
[0009] Based on the knowledge space and the questioning paradigm, a large-scale clinical auxiliary diagnosis and treatment model is established.
[0010] A second aspect of this application discloses a method for evaluating a large-scale clinical auxiliary diagnosis and treatment model, the method comprising:
[0011] A test dataset is constructed based on the priority order of the main dimensions of the knowledge space, the form of the question paradigm, the values of the question paradigm variables, and the synonymous expression of the question paradigm.
[0012] The test dataset is input into the clinical auxiliary diagnosis and treatment professional big model for processing to obtain the answer content of each test question in the test dataset;
[0013] The correctness of the answers is evaluated by business experts, and the confidence level of the large-scale clinical auxiliary diagnosis and treatment model is determined based on the error rate of the evaluation results.
[0014] A third aspect of this application discloses an apparatus for establishing a large-scale clinical auxiliary diagnosis and treatment model, the apparatus comprising:
[0015] The knowledge space mapping module is used to deconstruct professional literature into knowledge points based on the knowledge characteristics of clinical diagnosis and treatment, and map the knowledge points to the knowledge space;
[0016] The questioning paradigm construction module is used to construct questioning paradigms, which constrain the natural language expression of the question and the coordinate mapping relationship of the knowledge space.
[0017] The large model building module is used to build a large professional model for clinical auxiliary diagnosis and treatment based on the knowledge space and the questioning paradigm.
[0018] A fourth aspect of this application discloses an evaluation device based on a large-scale clinical auxiliary diagnosis and treatment model, the device comprising:
[0019] The data construction module is used to construct a test dataset based on the main dimensions of the knowledge space, the form of the question paradigm, the values of the question paradigm variables, and the priority order of the synonym expressions of the question paradigm.
[0020] The question processing module is used to input the test dataset into the clinical auxiliary diagnosis and treatment professional big model for processing, and obtain the answer content for each test question in the test dataset;
[0021] The response evaluation module is used to assess the correctness of the response content through business experts, and to determine the confidence level of the clinical auxiliary diagnosis and treatment professional big model based on the error rate of the evaluation results.
[0022] A fifth aspect of this application discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for establishing a large-scale clinical auxiliary diagnosis and treatment model as described in the first aspect of this application, or the steps of the method for evaluating a large-scale clinical auxiliary diagnosis and treatment model as described in the second aspect of this application.
[0023] The embodiments of this application have the following advantages:
[0024] In this embodiment, regarding the generation of answer content, professional literature is deconstructed into knowledge points based on the knowledge characteristics of clinical diagnosis and treatment, and these knowledge points are mapped to a knowledge space to achieve accurate location and retrieval of professional literature content. Furthermore, based on the positional relationships of knowledge points in the knowledge space (i.e., knowledge space coordinates), inclusion and adjacency relationships between knowledge points are also established, providing a foundation for the associated recommendation of knowledge points. Regarding the understanding of questions, a question paradigm is used to establish a coordinate mapping relationship between the natural language expression of the question and the knowledge space, thereby establishing a large-scale professional model for clinical auxiliary diagnosis and treatment based on the knowledge space and the question paradigm.
[0025] Thus, this large-scale clinical auxiliary diagnosis and treatment model can map common clinical diagnosis and treatment questions into knowledge space coordinates, and match knowledge points in professional literature based on these coordinates. This enables precise location and retrieval of the question content, and provides relevant recommendations based on the spatial relationships between knowledge points. Therefore, it solves the problems of "lack of authoritative medical literature support for answers" and "difficulty in accurately locating content and providing relevant recommendations" in the application of large language models in clinical auxiliary diagnosis and treatment. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1This is a flowchart illustrating the steps of a method for establishing a large-scale clinical auxiliary diagnosis and treatment model provided in an embodiment of this application.
[0028] Figure 2 This is a schematic diagram illustrating the dimensions of a knowledge space and the location of knowledge points, provided in an embodiment of this application.
[0029] Figure 3 This is a schematic diagram of a knowledge space construction method provided in an embodiment of this application;
[0030] Figure 4 This is a flowchart illustrating the steps of a question-answer generation method based on a large-scale clinical auxiliary diagnosis and treatment model provided in this application embodiment;
[0031] Figure 5 This is a flowchart of a question-answer generation method based on a large-scale clinical auxiliary diagnosis and treatment professional model provided in an embodiment of this application;
[0032] Figure 6 This is a flowchart illustrating the steps of an evaluation method for a large-scale clinical auxiliary diagnosis and treatment model provided in an embodiment of this application.
[0033] Figure 7 This is a schematic diagram of the structure of a device for establishing a large-scale clinical auxiliary diagnosis and treatment model provided in an embodiment of this application;
[0034] Figure 8 This is a schematic diagram of the structure of a large-scale clinical auxiliary diagnosis and treatment professional evaluation device provided in an embodiment of this application;
[0035] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0036] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0037] To overcome the problems of "lack of authoritative medical literature support (i.e., adherence to "evidence-based medicine" principles) and "difficulty in accurately locating content and providing relevant recommendations" in the application of large language models in clinical auxiliary diagnosis and treatment, this application proposes the following technical concept:
[0038] (1) Considering that the professional literature used to guide clinical auxiliary diagnosis and treatment mainly includes diagnosis and treatment guidelines and consensus compiled by experts in related fields, professional literature usually provides guidance on the admission examination, diagnosis, prognosis, treatment and other aspects of diagnosis and treatment for specific diseases. These aspects are intertwined with medical examinations, patient indicators, drugs and treatment plans, and there are also primary or secondary complication relationships between diseases. Therefore, whether keyword search or vector similarity matching is used, it is difficult to grasp the direction of the question and a large amount of content that the questioner does not care about will be retrieved, which will greatly reduce the usability of the system.
[0039] Therefore, in terms of generating response content, this application proposes a method for indexing document content based on knowledge space mapping. According to the knowledge characteristics of clinical diagnosis and treatment, professional documents are deconstructed into knowledge points, and the knowledge points are mapped to the knowledge space to achieve accurate location retrieval of professional document content. Furthermore, based on the positional relationship of knowledge points in the knowledge space, the inclusion and adjacency relationships between knowledge points are also established, providing a foundation for the association and recommendation of knowledge points.
[0040] Regarding the understanding of questions, this application's embodiments analyze and summarize common clinical diagnosis and treatment problems, and propose a questioning paradigm description method based on variable constraints. Through the questioning paradigm, a coordinate mapping relationship between the natural language expression of the question and the knowledge space is established.
[0041] A comprehensive clinical auxiliary diagnosis and treatment model is established based on knowledge space and question paradigm. This model maps common clinical diagnosis and treatment questions to knowledge space coordinates and matches knowledge points in professional literature based on these coordinates. This enables precise location and retrieval of question content and provides relevant recommendations based on the spatial relationships between knowledge points.
[0042] (2) Considering that different questions test different capabilities of the professional model, users must know the extent to which they can trust the professional model's answers before using it; that is, the confidence level of the answer is related to the type and manner of the question. This application's embodiments tested the confidence level of the professional model for clinical auxiliary diagnosis and treatment. To understand the confidence level of the professional model's answers to specific types of questions, it is necessary to analyze and categorize the professional model's capabilities, and then design corresponding assessment schemes for evaluation. The capabilities of the professional model are reflected in two aspects: first, the coverage of professional knowledge (depending on the input set of professional literature); and second, the ability to understand the question and answer it using professional knowledge.
[0043] This application embodiment tests the coverage of professional knowledge. Based on the variable value space in the question paradigm, a test dataset is generated. For answering ability, evaluation items are formulated, including: understanding of questions, knowledge point location, comprehensiveness of content, and correctness of reasoning. The test dataset is input into the established clinical auxiliary diagnosis and treatment professional big model, and business experts evaluate the answer content generated based on the search results. The confidence level of the clinical auxiliary diagnosis and treatment professional big model is calculated by the error rate.
[0044] The following sections, with reference to the accompanying drawings, provide a detailed description of the methods for establishing and evaluating the large-scale clinical auxiliary diagnosis and treatment model provided in the embodiments of this application, in sections 1.1, 1.2, 1.3, and 1.4.
[0045] 1.1 Methodology for establishing a large-scale clinical auxiliary diagnosis and treatment model:
[0046] Reference Figure 1 As shown, Figure 1 This is a flowchart illustrating the steps of a method for establishing a large-scale clinical auxiliary diagnosis and treatment model provided in an embodiment of this application. Figure 1 As shown, the method for establishing this large-scale clinical auxiliary diagnosis and treatment model may include steps S110 to S130:
[0047] Step S110: Deconstruct professional literature into knowledge points based on the knowledge characteristics of clinical diagnosis and treatment, and map the knowledge points to the knowledge space.
[0048] Step S120: Construct a questioning paradigm, which is used to constrain the coordinate mapping relationship between the natural language expression of the question and the knowledge space.
[0049] Step S130: Based on the knowledge space and the questioning paradigm, establish a large-scale clinical auxiliary diagnosis and treatment model.
[0050] In this embodiment, the knowledge space represents the knowledge coverage encompassed by various knowledge points within the clinical diagnosis and treatment field. Clinical diagnosis and treatment knowledge includes multiple elements, primarily including: disease (disease and its stages, complications, etc.), diagnosis and treatment (diagnosis, prognosis, treatment, follow-up, etc.), drugs (drug combinations used in treatment plans), examinations (microbiological examinations, imaging examinations, physical examinations, etc.), and patient indicators (age, body surface area, prognostic score). Each clinical diagnosis and treatment knowledge point involves a combination of these elements, and all knowledge points constitute the knowledge coverage.
[0051] Therefore, this application embodiment deconstructs professional literature into knowledge points based on the knowledge characteristics of clinical diagnosis and treatment, and maps the knowledge points to the knowledge space to achieve accurate location retrieval of professional literature content; furthermore, based on the positional relationship of knowledge points in the knowledge space, it also establishes the inclusion and adjacency relationships between knowledge points, providing a foundation for the association and recommendation of knowledge points.
[0052] In this embodiment of the application, the questioning paradigm refers to the way questions are described. By analyzing and summarizing common clinical diagnosis and treatment problems, a questioning paradigm is constructed to constrain the natural language expression of questions and the coordinate mapping relationship of the knowledge space. The questioning paradigm standardizes the natural language expression of knowledge point questions.
[0053] This leads to the establishment of a comprehensive clinical auxiliary diagnosis and treatment model based on knowledge space and question paradigms. This model maps common clinical diagnosis and treatment questions to knowledge space coordinates and matches these coordinates with knowledge points in professional literature, enabling precise retrieval of question content and providing relevant recommendations based on the spatial relationships between knowledge points. Therefore, it addresses the problems of "lack of authoritative medical literature support for answers" and "difficulty in accurately locating content and providing relevant recommendations" in the application of large language models in clinical auxiliary diagnosis and treatment.
[0054] The methods for constructing knowledge spaces and questioning paradigms will be explained in detail in Sections 1.1.1 and 1.1.2, respectively.
[0055] 1.1.1 Knowledge Space Construction Method:
[0056] Reference Figure 2 As shown, Figure 2 This is a schematic diagram illustrating the dimensions and knowledge point locations of a knowledge space according to an embodiment of this application. The knowledge space includes a tree-like multi-level classification tree with multiple dimensions. The multiple dimensions and the multi-level classifications under each dimension constitute a knowledge space coordinate system. The multiple dimensions include: disease, drug, medical record acquisition, patient indicators, and diagnosis and treatment process.
[0057] Specifically, the knowledge space includes: knowledge points, knowledge point locations, knowledge point indexes, and knowledge point relationships. A knowledge point refers to a specific category within a particular dimension (e.g., a disease dimension) in clinical diagnostic and treatment literature. In some embodiments, a knowledge point refers to a conceptual or operational statement that takes a specific category under another dimension as a premise. For example, the knowledge point "drugs for treating lymphoma" uses "lymphoma" as a premise.
[0058] The location of a knowledge point refers to its coordinates in the knowledge space (i.e., knowledge space coordinates). Knowledge space coordinates include primary coordinates and secondary coordinates. The dimension in which the topic of the knowledge point is located is taken as the primary dimension, and the other dimensions in which the prerequisites of the knowledge point are located are taken as conditional dimensions. Thus, the classification of the knowledge point in the primary dimension is defined as the primary coordinate, and the classification of the knowledge point in the conditional dimension is defined as the secondary coordinate. The primary coordinate and the secondary coordinate together determine the location of the knowledge point in the knowledge space.
[0059] A knowledge point index (i.e., a knowledge point vector index) is used to retrieve knowledge points within a knowledge space. For example, a large-scale clinical auxiliary diagnosis and treatment model can retrieve the knowledge points involved in a question based on the knowledge point index. The knowledge point index includes: the knowledge space coordinates of the knowledge points and the text vector index.
[0060] The relationships between knowledge points include: overlapping relationships, inclusion relationships, and adjacency relationships. After the knowledge contained in professional literature is deconstructed into knowledge points, the relationships between knowledge points can be obtained through the relationships between their coordinates in the knowledge space. Among them, overlapping relationships refer to knowledge points with completely identical coordinates; inclusion relationships refer to knowledge points with completely identical coordinates and classified as master-child relationships; and adjacency relationships refer to knowledge points with completely identical coordinates and classified as sibling nodes under the same master node.
[0061] In one specific implementation, step S110 above, "deconstructing professional literature into knowledge points according to the knowledge characteristics of clinical diagnosis and treatment, and mapping the knowledge points to a knowledge space," includes steps A1 to A4:
[0062] Step A1: Use each dimension as the root node of the tree-like multi-level classification tree to obtain the initial knowledge space coordinate system.
[0063] Step A2: Extract content and construct a document tree from the professional documents to obtain the document tree.
[0064] Step A3: Extract the knowledge space coordinates based on the text content of each document node in the document tree to form knowledge points in the knowledge space.
[0065] Step A4: Refine the initial knowledge space coordinate system based on the knowledge points, and establish an index for each knowledge point to obtain the knowledge space.
[0066] In this embodiment, multiple dimensions such as disease, drug, medical record acquisition, patient indicators, and diagnosis and treatment process are named as the root node of each tree-like multi-level classification tree to form an initial knowledge space coordinate system.
[0067] Among them, professional literature refers to diagnostic and treatment guidelines and consensus summarizing and compiled by experts in relevant fields. Professional literature usually provides guidance on the diagnosis and treatment of specific diseases, such as admission examination, diagnosis, prognosis, and treatment. A document tree is a tree composed of document nodes according to the paragraph and chapter structure of professional literature. The document tree is obtained by extracting content from professional literature and constructing the document tree based on the extracted content.
[0068] Specifically, step A2, "extracting content and constructing a document tree from the professional document to obtain a document tree," includes: converting the professional document into a target format file; extracting the text and table content from the target format file to obtain plain text paragraph content and table text; using the natural paragraphs of the professional document as document nodes, establishing a master-child document node relationship according to the chapter affiliation of the professional document; and adding the plain text paragraph content and the table text to the corresponding document nodes to obtain a document tree.
[0069] In practical applications, professional documents are usually in PDF (Portable Document Format) format. To facilitate the extraction of content from these documents, they are converted to a target format (e.g., HTML, where HTML stands for HyperText Markup Language). Tools like pdf2htmlex (which converts PDF files to HTML) or similar tools can be used to convert PDF documents to HTML. Similarly, tools like BeautifulSoup+cssutil (a parser for HTML documents and a Python library for parsing and manipulating Cascading Style Sheets) or similar tools can be used to extract the main text from the HTML file, obtaining plain text paragraphs. Simultaneously, OpenCV (an open-source computer vision library providing rich image processing and computer vision algorithms) or similar tools can be used to extract table data. This involves identifying the table borders in the HTML file, recognizing the column headers based on the borders, and then identifying the rows of data line by line to obtain the table text. After obtaining the plain text paragraph content and table text, a document tree is built based on the natural paragraphs of the professional literature. The document tree establishes a master-child document node relationship based on the chapter affiliation of the professional literature, thereby adding the plain text paragraph content and table text to the corresponding document nodes to obtain the document tree.
[0070] For example, the data structure of a document tree can be represented as:
[0071] {"id":"", / / String, the ID (identifier) of the document tree node; each node has a unique ID;
[0072] "parent":"", / / String, the ID of the parent node of the node. If it is the root node of the document tree, the parent node ID is an empty string;
[0073] "children":[""], / / String array, each node has multiple child nodes, storing the IDs of the child nodes; if there are no child nodes, the array length is 0;
[0074] "prev_node":"", / / String, the ID of the previous sibling node;
[0075] "next_node":"", / / String, the ID of the next sibling node;
[0076] "title":"", / / String, the topic of the current node;
[0077] "content":"", / / String, the text content of the current node;
[0078] "summary":"", / / String, a text summary of the current node;
[0079] "keyword":[""], / / Array of strings, the text keyword of the current node;
[0080] "kb_dimension":[""], / / String array, the knowledge dimension array of the current node}.
[0081] After obtaining the document tree, step A3 is executed: based on the text content of each document node in the document tree, knowledge space coordinates are extracted to form knowledge points in the knowledge space. In practice, the large model extracts knowledge space coordinates based on the text content to form knowledge points in the knowledge space.
[0082] The knowledge space coordinates include primary coordinates and secondary coordinates. Specifically, based on the text content of each document node in the document tree, the knowledge space coordinates are extracted, including: determining the dimension containing the topic of the text content as the primary dimension, and determining the category name of the text content on the primary dimension as the primary coordinate; determining the dimension containing the preconditions of the text content as the condition dimension, and determining the category name of the text content on the condition dimension as the secondary coordinate.
[0083] In some embodiments, step A3, "extracting knowledge space coordinates based on the text content of each document node in the document tree," may specifically include sub-steps A31 to A33:
[0084] Step A31: Use the multiple dimensions and one alternative dimension as classification options for the text content.
[0085] Step A32: If the text content is classified as the candidate dimension, determine the knowledge space coordinates as empty coordinates.
[0086] Step A33: When the text content is classified into the multiple main dimensions, perform the following steps: 1) When the text content involves a new category, add a new category name as a knowledge space coordinate in the knowledge space coordinate system; 2) When the text content involves an existing category with the same name, use the existing category with the same name as the knowledge space coordinate; 3) When the text content involves an existing category with the same identity but different category names, use the existing category name as the knowledge space coordinate.
[0087] In this embodiment, the knowledge space coordinates include primary coordinates and secondary coordinates, both determined according to steps A31 to A33 described above. Five dimensions—"disease, drug, medical record acquisition, patient indicators, and treatment process"—are used as classification options for the text content. Simultaneously, an alternative dimension is added as a classification option; this alternative dimension is a category other than the five dimensions. By setting the classification options, the clinical auxiliary diagnosis and treatment professional large model is required to select the content classification option to which the document node belongs. For text content classified as an alternative dimension, the knowledge space coordinates are determined to be empty coordinates, meaning that non-clinical knowledge whose topic is not within the knowledge space dimension is not included. However, for text content classified as a primary dimension (i.e., the topic of the text content is within the knowledge space coordinate system), the knowledge space coordinates are determined according to the three methods described in step A33.
[0088] After obtaining the knowledge points, step A4 is executed to refine the initial knowledge space coordinate system based on the knowledge points and establish an index for each knowledge point, thus obtaining the knowledge space. Refining the initial knowledge space coordinate system based on knowledge points means that, during the continuous generation of new knowledge points, the subcategories of each dimension of the knowledge space coordinate system are continuously improved based on the subcategories of the new knowledge points.
[0089] Specifically, establishing an index for each knowledge point (i.e., a knowledge point vector index) refers to: establishing the knowledge space coordinates and text vector indexes for each knowledge point. Specifically, establishing an index for each knowledge point includes: generating a text content vector based on the text content of each document node in the document tree; and establishing a knowledge point vector index for each document node in the document tree in a vector database based on the text content vectors.
[0090] Through the above implementation process, a knowledge space coordinate system was established for clinical diagnosis and treatment professional knowledge, consisting of five dimensions: disease, drug, medical record acquisition, patient indicators, and diagnosis and treatment process, as well as subcategories under each dimension. Professional documents were deconstructed into document trees, and knowledge space coordinates were extracted from the nodes of the document trees to form knowledge points, which were then mapped to the knowledge space.
[0091] For example, Figure 3 This is a schematic diagram of a knowledge space construction method provided in an embodiment of this application. The knowledge space is obtained by extracting content from patent documents, constructing knowledge, maintaining the knowledge space, and establishing an index. Specifically, content extraction refers to: converting professional documents into target format files (e.g., converting PDF professional documents into HTML files), thereby extracting the text content and table content of the target format files respectively, obtaining plain text paragraph content and table text; using the natural paragraphs of the professional documents as document nodes, establishing a master-child document node relationship according to the chapter affiliation of the professional documents; and adding the plain text paragraph content and table text to the corresponding document nodes to obtain a document tree.
[0092] Knowledge construction refers to: extracting knowledge space coordinates from the text content of each document node in the document tree to form knowledge points in the knowledge space; and generating text content vectors based on the text content of each document node in the document tree. Knowledge space maintenance refers to: refining the initial knowledge space coordinate system using knowledge points to obtain the knowledge space. Indexing refers to: creating knowledge point vector indexes for each document node in the document tree in a vector database based on the text content vectors. It should be noted that knowledge construction, knowledge space maintenance, and indexing are performed simultaneously; that is, the knowledge space is maintained and the knowledge point vector indexes are created during the knowledge point extraction process.
[0093] In this way, the knowledge space enables accurate location and retrieval of professional literature content; and, based on the positional relationship of knowledge points in the knowledge space, it also establishes the inclusion and adjacency relationships between knowledge points, providing a foundation for the association and recommendation of knowledge points.
[0094] 1.1.2 Questioning Paradigm Construction Method:
[0095] In one specific implementation, the "constructing a questioning paradigm" in step S120 above includes steps B1 and B2:
[0096] Step B1: Determine the topic of the knowledge points involved in the question as the main dimension of the paradigm, and determine the dimension of the preconditions of the knowledge points involved in the question as the conditional dimension of the paradigm.
[0097] Step B2: Construct a question paradigm based on the main dimension and the conditional dimension; wherein the question paradigm includes: a primitive paradigm containing variables and rule symbols, and an intermediate paradigm containing variables.
[0098] In the embodiments of the present application, the question paradigm is used to standardize the natural language expressions for asking knowledge points, and one or more variables are included in the question paradigm. Compared with the intermediate paradigm, the original paradigm contains rule symbols. If the rule symbols are removed from the original paradigm, then the question paradigm is the intermediate paradigm.
[0099] For example, for the original paradigm: What diseases can be differentially diagnosed with A001 for [1 available][1X101]?
[0100] Variable description: X101 is a lubricating word variable, and its value range is: "required", "should", and A001 is a disease name variable, and its value range is common disease names.
[0101] Rule description: The content enclosed in "[]" is optional content; the number immediately following "[" represents mutually exclusive groups, that is, at most one of all the contents in the same group.
[0102] Correspondingly, examples of the intermediate paradigm corresponding to the above original paradigm include:
[0103] A: What diseases can be differentially diagnosed with A001?
[0104] B: What diseases can be differentially diagnosed with A001 for X101?
[0105] C: What diseases can be differentially diagnosed with A001?
[0106] It should be noted that the question paradigm has synonymy. When the variable values of the question paradigm are the same, the two paradigms express the same meaning, and they are called "synonymous paradigms" with each other. Examples are as follows:
[0107] 1. What diseases can be differentially diagnosed with A001 for A001X101?
[0108] 2. What diseases can be differentially diagnosed with A001 for [1 available][1X101]?
[0109] In the embodiments of the present application, the question paradigm has a main dimension and / or a conditional dimension. The theme of the knowledge points involved in the question is called the main dimension of the paradigm; the prerequisite conditions of the knowledge points involved in the question are called the conditional dimension of the paradigm. For examples of question paradigms with different main dimensions and conditional dimensions, reference can be made to Table 3 in Section 1.4 below. The variable names in the question paradigm are composed of one capital letter representing the dimension / category and three Arabic numerals, a total of 126. For examples of variable names, reference can be made to Tables 4 to 13 in Section 1.4 below.
[0110] It is understandable that when a question paradigm includes a variable, the paradigm only includes the primary dimension and does not include the conditional dimension. For example, in the example question paradigm "What are the preceding diseases of A001?" in item 1 of Table 3, the dimension containing the variable A001 is the primary dimension, and the question paradigm does not include the conditional dimension.
[0111] When a question paradigm contains multiple variables, it includes a primary dimension and a conditional dimension. For example, in the question paradigm "[When] Y001 is used to treat A001[1][1], what are the complications?" in item 2 of Table 3, the dimension containing variable Y001 is the conditional dimension, and the dimension containing variable A001 is the primary dimension.
[0112] In one specific implementation, constructing a questioning paradigm based on the main dimension and the conditional dimension includes: when the questioning paradigm includes multiple variables, determining variable constraints based on the values of the multiple variables; and constructing the questioning paradigm based on the variable constraints, the main dimension, and the conditional dimension.
[0113] In this embodiment of the application, if a question paradigm has two or more variables, there may be constraints between the values of the variables. The variable constraints are determined based on the values of multiple variables, and the question paradigm is constructed based on the variable constraints. For example, taking the value constraints between the variable name "A001" and other variable names as an example, the variable constraints are shown in Tables 14 to 17.
[0114] In this embodiment, a questioning paradigm standardizes the expressions for knowledge point questions in a natural language manner. Each questioning paradigm includes one or more variables; if multiple variables are included in the same paradigm, variable constraints can be added between them. This allows for the establishment of a large-scale clinical auxiliary diagnosis and treatment model based on the knowledge space and the questioning paradigm. This model can map common clinical diagnosis and treatment questions to knowledge space coordinates, and match these coordinates with knowledge points in professional literature, thereby achieving precise location and retrieval of question content and providing relevant recommendations based on the spatial relationships between knowledge points.
[0115] 1.2 Question and Answer Generation Method Based on a Large-Scale Clinical Auxiliary Diagnosis and Treatment Model:
[0116] Reference Figure 4 As shown, Figure 4 This is a flowchart illustrating the steps of a question-answer generation method based on a large-scale clinical auxiliary diagnosis and treatment model provided in this application embodiment. This method is applied to the large-scale clinical auxiliary diagnosis and treatment model established by the aforementioned method. Figure 4 As shown, the large-scale clinical auxiliary diagnosis and treatment model can generate answers to questions by following steps S410 to S430.
[0117] Step S410: Obtain the target question raised by the user according to the questioning paradigm, and extract the first knowledge space coordinates from the target question. The questioning paradigm is used to constrain the natural language expression of the question and the coordinate mapping relationship of the knowledge space.
[0118] Step S420: Based on the first knowledge space coordinates, match knowledge points from the knowledge space to obtain target knowledge points that satisfy the first knowledge space coordinates.
[0119] Step S430: Based on the target knowledge points, generate the answer to the target question and provide the answer to the user.
[0120] In this embodiment, the clinical auxiliary diagnosis and treatment professional big data model extracts the first knowledge space coordinates from the target question based on the knowledge space coordinates of the current dialogue context (i.e., historical knowledge space coordinates). It then compares the first knowledge space coordinates with the knowledge point mapping formed by professional literature deconstruction in the knowledge space to find target knowledge points that satisfy the first knowledge space coordinates. Based on this, the clinical auxiliary diagnosis and treatment professional big data model summarizes and organizes the target knowledge points to form a summary text (i.e., the answer to the target question), and finally feeds the answer to the target question back to the user.
[0121] In this way, based on the natural language processing capabilities of the large-scale clinical auxiliary diagnosis and treatment model, combined with the question paradigm, the coordinate information of the corresponding knowledge space is extracted from the target question. Based on the knowledge space coordinates in the target question, knowledge points in professional literature are matched in the knowledge space, thereby achieving precise location and retrieval of the question content and ensuring the accuracy of the final generated question answer.
[0122] In one specific implementation, step S420 above, "matching knowledge points from the knowledge space according to the first knowledge space coordinates to obtain target knowledge points that satisfy the first knowledge space coordinates," includes steps F1 to F4:
[0123] Step F1: Based on the first knowledge space coordinates and the knowledge space coordinates of the knowledge points, perform vector retrieval routing to narrow the retrieval scope from all documents to documents whose root node's knowledge space coordinates intersect with the first knowledge space coordinates.
[0124] Step F2: Process the target question into a target question vector, and match the target question vector with the knowledge point vector index within the retrieval range to obtain a set of semantically similar knowledge points and the similarity of each knowledge point.
[0125] Step F3: Perform intersection calculation between the second knowledge space coordinates corresponding to each knowledge point and the first knowledge space coordinates to obtain a set of knowledge points that satisfy the first knowledge space coordinates.
[0126] Step F4: Based on the similarity of each knowledge point, sort each knowledge point in the knowledge point set, and select the top N knowledge points as the target knowledge points, where N is an integer greater than 0.
[0127] In this embodiment, document retrieval routing is performed based on the knowledge space coordinates of knowledge points, narrowing the retrieval range of the target question vector, thereby reducing the retrieval computation and shortening the retrieval response time. The knowledge space includes a knowledge point vector index. The target question is vectorized, and within the retrieval range, semantic similarity matching is performed between the target question vector and the knowledge point vector index to recall a set of semantically similar knowledge points and the similarity of each knowledge point. This filters the knowledge points by calculating the intersection of the second knowledge space coordinates and the first knowledge space coordinates corresponding to each knowledge point. Specifically, it determines whether the second knowledge space coordinates and the first knowledge space coordinates of each knowledge point overlap, and identifies overlapping knowledge points as the set of knowledge points that satisfy the first knowledge space coordinates. Then, the knowledge points are sorted according to the similarity of each knowledge point in step F2, and the top N knowledge points with the highest similarity ranking are selected as target knowledge points.
[0128] In this way, by using knowledge space coordinates to filter knowledge points with similar semantics, irrelevant information is removed, and accurate content positioning is achieved.
[0129] In one specific implementation, the method further includes the following steps E1 to E3:
[0130] Step E1: Obtain the set of references corresponding to the target knowledge point.
[0131] Step E2: Generate a set of recommended questions based on the knowledge space coordinates of the current dialogue context, where the knowledge space coordinates of the current dialogue context represent the historical knowledge space coordinates.
[0132] Step E3: Feed back the answers to the questions, the reference set, and the recommended question set to the user.
[0133] In this embodiment, the content organization based on knowledge points provides corresponding references and sources, adhering to the principles of evidence-based medicine in clinical diagnosis and treatment. Furthermore, based on the knowledge space coordinates of the current dialogue context, a set of recommended questions is generated (i.e., adjacent coordinates and parent coordinates are returned from the knowledge space coordinates of the current dialogue context as recommended questions) to recommend adjacent or related knowledge to the user. By using knowledge space coordinates instead of dialogue history in the dialogue context, this method significantly shortens the context length and reduces the performance requirements of large model runtime environments.
[0134] In one specific implementation, the knowledge space coordinates of the current dialogue context are maintained in the following manner: if the knowledge space coordinates of the current dialogue context overlap with the first knowledge space coordinates, the knowledge space coordinates of the current dialogue context remain unchanged; if the knowledge space coordinates of the current dialogue context do not overlap with the first knowledge space coordinates, the first knowledge space coordinates are used as the knowledge space coordinates of the current dialogue context.
[0135] In this way, during the process of generating the answer to the target question, the knowledge space coordinates of the current dialogue context are maintained, ensuring the accuracy of the knowledge space coordinates of the current dialogue context, thereby achieving precise positioning of the content.
[0136] For example, Figure 5 This is a flowchart illustrating a method for generating answers to questions based on a large-scale clinical auxiliary diagnosis and treatment model, as provided in this application embodiment. Specifically, it includes the following steps G1 to G9:
[0137] Step G1: In the context of the dialogue between the user and the professional clinical auxiliary diagnosis and treatment model, obtain the target question raised by the user according to the questioning paradigm.
[0138] Step G2: Send the target question to the Clinical Auxiliary Diagnosis and Treatment Professional Model, and require the Clinical Auxiliary Diagnosis and Treatment Professional Model to extract the first knowledge space coordinate (including extraction based on synonyms of the fine classification) Kn from the target question based on the knowledge space coordinate Kl of the current dialogue context; if the first knowledge space coordinate Kn overlaps with the knowledge space coordinate Kl of the current dialogue context, Kl is not changed; otherwise, Kn is used to replace Kl as the knowledge space coordinate of the current dialogue context.
[0139] Step G3: Based on the first knowledge space coordinates and the knowledge space coordinates of the knowledge points, perform vector retrieval routing (i.e., document retrieval routing) to narrow the retrieval scope from all documents to documents whose root node's knowledge space coordinates intersect with the first knowledge space coordinates.
[0140] Step G4: Vectorize the target question, perform semantic similarity matching on the knowledge point text vector index within the retrieval range, and recall the set of semantically similar knowledge points and the similarity of each knowledge point.
[0141] Step G5: Knowledge point filtering. The intersection of the second knowledge space coordinates and the first knowledge space coordinates corresponding to each knowledge point is calculated to obtain a set of knowledge points that satisfy the first knowledge space coordinates (filtering rule: the first knowledge space coordinates and the second knowledge space coordinates overlap). Then, according to the similarity of each knowledge point, each knowledge point in the knowledge point set is sorted, and the top N knowledge points in the sorting are taken as the target knowledge points, thus obtaining the set T of knowledge points with the highest similarity ranking.
[0142] Step G6: Input the knowledge point set T into the large-scale clinical auxiliary diagnosis and treatment model, and require the large-scale clinical auxiliary diagnosis and treatment model to summarize and organize it to form a summary text content S (i.e. generate the answer to the target question).
[0143] Step G7: Obtain the reference set R corresponding to the filtered knowledge points (i.e., the target knowledge points).
[0144] Step G8: Generate a set of recommended questions Q based on the knowledge space coordinates Kl of the current dialogue context. The recommendation rule is to return the adjacent coordinates and the parent coordinate from the knowledge space coordinates Kl of the current dialogue context as recommended questions.
[0145] Step G9: Return to the user the knowledge space coordinates Kl of the current dialogue context, the summary text content S, the reference set R, and the recommendation question set Q.
[0146] Through the above implementation process, when generating answers to user questions, knowledge space coordinates are extracted from the dialogue context and the question. Document retrieval routing is then performed based on the knowledge space coordinates of the knowledge points, narrowing the retrieval range of the target question vector. Within the retrieval range, knowledge space coordinates are used to filter knowledge points with similar semantics, removing irrelevant information and achieving rapid and accurate content location. Based on the organization of content according to knowledge points, corresponding references and sources can be provided, adhering to the evidence-based medicine principles of clinical diagnosis and treatment. In the dialogue context, using knowledge space coordinates instead of dialogue history can significantly shorten the context length and reduce the performance requirements of large model operating environments.
[0147] 1.3 Evaluation Methods for the Large-Scale Clinical Auxiliary Diagnosis and Treatment Model:
[0148] In this embodiment of the application, the confidence level of the large-scale clinical auxiliary diagnosis and treatment professional model established in Section 1.1 is tested from two aspects: the "coverage of professional knowledge" and the ability to "understand questions and answer them using professional knowledge".
[0149] Reference Figure 6 As shown, Figure 6 This is a flowchart illustrating the steps of an evaluation method for a large-scale clinical auxiliary diagnosis and treatment model provided in this application embodiment. Figure 6 As shown, the evaluation method for this large-scale clinical auxiliary diagnosis and treatment model may include steps S610 to S630:
[0150] Step S610: Construct a test dataset based on the priority order of the main dimension of the knowledge space, the form of the question paradigm, the value of the question paradigm variable, and the synonym expression of the question paradigm.
[0151] Step S620: Input the test dataset into the clinical auxiliary diagnosis and treatment professional big model for processing to obtain the answer content of each test question in the test dataset.
[0152] Step S630: The correctness of the answer is evaluated by business experts, and the confidence level of the clinical auxiliary diagnosis and treatment professional big model is determined based on the error rate of the evaluation results.
[0153] In this embodiment, a test dataset is constructed according to the priority order of the main dimension of the knowledge space, the form of the question paradigm, the value of the question paradigm variable, and the synonym expression of the question paradigm, ensuring the uniformity of the distribution of test questions in the test dataset. The confidence level of the large-scale clinical auxiliary diagnosis and treatment model is then measured based on this test dataset. Furthermore, business experts evaluate the content of the answers generated based on the search results, reach a consensus on the evaluation results, and calculate the confidence level of various aspects of the large-scale clinical auxiliary diagnosis and treatment model based on the error rate of the evaluation results. Thus, accurate and comprehensive confidence level testing is achieved.
[0154] The evaluation process involved business experts assessing the responses generated based on search results and reaching a consensus on the assessments. This included evaluations of various aspects such as question comprehension, knowledge point identification, content comprehensiveness, and reasoning correctness. Specifically, question comprehension refers to whether the relevant knowledge space dimensions were correctly extracted from the question and the context of the dialogue; knowledge point identification refers to whether the corresponding set of knowledge points was retrieved based on the question's knowledge space dimensions, and whether any irrelevant knowledge points existed; content comprehensiveness refers to whether all corresponding knowledge points were retrieved; and reasoning correctness refers to whether the content extracted and summarized from the retrieved knowledge points is logically sound. In this way, through these evaluations of question comprehension, knowledge point identification, content comprehensiveness, and reasoning correctness, a comprehensive test of the question-answering ability of the large-scale clinical auxiliary diagnosis and treatment model was conducted.
[0155] 1.3.1 Method for constructing the test dataset:
[0156] In some specific implementations, step S610 above, "constructing a test dataset according to the priority order of the main dimension of the knowledge space, the form of the question paradigm, the value of the question paradigm variable, and the synonymous expression of the question paradigm," includes steps C1 to C5:
[0157] Step C1: Determine the number of questions and the set of question paradigms.
[0158] The questioning paradigm set contains the questioning paradigms that need to be tested.
[0159] Step C2: Based on the number of questions and the set of question paradigms, determine the target main dimension and the first number of questions for each target main dimension.
[0160] In this embodiment, each questioning paradigm includes a main dimension, and different questioning paradigms correspond to different main questioning dimensions. To ensure the uniformity of the distribution of test questions, the target main dimension (i.e., the main dimension of the knowledge space to be tested) and the number of questions corresponding to each target main dimension (i.e., the first number of questions) are first determined based on the number of questions and the set of questioning paradigms.
[0161] Specifically, based on the number of questions and the set of question paradigms, the target main dimension and the first number of questions for each target main dimension are determined, including:
[0162] If the number of questions is less than the number of main dimensions in the question paradigm set, a main dimension is randomly selected from all main dimensions as the target main dimension, and the first number of questions for each target main dimension is determined to be 1.
[0163] If the number of questions is not less than the number of principal dimensions in the question paradigm set, all principal dimensions are taken as target principal dimensions, and the number of questions is evenly distributed to obtain the first number of questions for each target principal dimension.
[0164] For example, given the number of questions M, if M is less than the number of principal dimensions in the question paradigm set, it means that not all principal dimensions can be measured, and only some principal dimensions can be measured. In this case, a principal dimension is randomly selected from all principal dimensions as the target principal dimension, and the first number of questions Dia for each target principal dimension is set to 1. Otherwise, all principal dimensions are used as target principal dimensions, and the number of questions M is evenly distributed (i.e., evenly distributed according to divisibility, and then the remaining number is randomly distributed), finally obtaining the target principal dimension Di and the first number of questions Dia for each target principal dimension.
[0165] Step C3: Based on the target main dimension and the first number of questions, determine the target question paradigm under each target main dimension and the second number of questions for each target question paradigm.
[0166] In this embodiment of the application, each target main dimension corresponds to a different target question paradigm. In order to ensure the uniformity of the distribution of test questions, the target question paradigm under each target main dimension and the number of questions corresponding to each target question paradigm are determined according to the target main dimension and the first number of questions, i.e., the second number of questions.
[0167] Specifically, based on the target main dimension and the first number of questions, the target question paradigm under each target main dimension and the second number of questions for each target question paradigm are determined, including:
[0168] For all question paradigms of each target main dimension, only one representative paradigm is selected from all synonym paradigms to form a set of question paradigms to be extracted.
[0169] If the number of questions in the first question is less than the number of question paradigms in the set of question paradigms to be extracted, question paradigms are randomly extracted from the set of question paradigms to be extracted as target question paradigms under each target main dimension, and the number of questions in the second question of each target question paradigm is determined to be 1.
[0170] If the second number of questions is not less than the number of question paradigms in the set of question paradigms to be extracted, all question paradigms in the set of question paradigms to be extracted are taken as target question paradigms, and the first number of questions is evenly distributed to obtain the target question paradigm under each main dimension and the second number of questions for each target question paradigm.
[0171] For example, for all question paradigms under the target main dimension Di, only one representative paradigm is selected from all synonym paradigms to form the question paradigm set Pi to be extracted. If the first number of questions Din is less than the number of question paradigms in the question paradigm set Pi, it means that not all question paradigms in the question paradigm set Pi can be measured, and only a portion of the question paradigms can be measured. In this case, question paradigms are randomly selected from the question paradigm set Pi to be extracted as the target question paradigms under each target main dimension. Otherwise, the first number of questions Din is evenly distributed across each question paradigm in the question paradigm set Pi to be extracted, resulting in the target question paradigms under the target main dimension Di and the second number of questions Qin for each target question paradigm.
[0172] Step C4: Based on the target questioning paradigm and the second number of questions, determine the test questioning paradigm and the third number of questions for each test questioning paradigm, wherein the test questioning paradigms are all in the form of intermediate paradigms.
[0173] In this embodiment, the questioning paradigm includes a primary paradigm and an intermediate paradigm, and the content of the questions in the primary paradigm and the intermediate paradigm may be identical. Therefore, to ensure the uniformity of the distribution of test questions, a test questioning paradigm in the form of an intermediate paradigm is determined based on the target questioning paradigm and the second number of questions, as well as the number of questions corresponding to each test questioning paradigm, i.e., the third number of questions.
[0174] Specifically, based on the target questioning paradigm and the second number of questions, a test questioning paradigm and a third number of questions for each test questioning paradigm are determined, including:
[0175] When the test questioning paradigm is the original paradigm, intermediate paradigms are extracted from the intermediate paradigm set corresponding to the original paradigm to form the test questioning paradigm, and the second questioning quantity is evenly distributed to obtain the third questioning quantity for each test questioning paradigm.
[0176] When the test questioning paradigm is an intermediate paradigm, the target questioning paradigm is used as the test questioning paradigm, and the number of third questions in each test questioning paradigm is equal to the number of second questions.
[0177] For example, for each target question paradigm qi and its second number of questions Qin, if the target question paradigm qi contains rule symbols, it indicates that the target question paradigm qi is the original paradigm. Then, further, from the set of intermediate paradigms corresponding to the original paradigm, intermediate paradigms are extracted as test question paradigms ri (resulting in the test question paradigm set Ri), and the second number of questions Qin is evenly distributed, thus obtaining the test question paradigms and the third number of questions Rin for each test question paradigm. Otherwise, the target question paradigm qi is used as the test question paradigm ri, and the third number of questions Rin for each test question paradigm is equal to the second number of questions Qin.
[0178] Step C5: Determine the variable values for the questioning paradigm based on the test questioning paradigm and the third number of questions, and generate a test dataset based on the variable values for the questioning paradigm.
[0179] In this embodiment of the application, the variables in the question paradigm include different values. In order to ensure the uniformity of the distribution of test questions, it is necessary to determine the variable values of the question paradigm so as to generate the corresponding test dataset based on the variable values of the question paradigm.
[0180] Specifically, based on the test questioning paradigm and the number of third questions, the variable values for the questioning paradigm are determined, and a test dataset is generated based on the variable values for the questioning paradigm, including:
[0181] When there are variable constraints in the test question paradigm, the variable value pairs are used as the set of variable values to be extracted. The third number of variable pairs are extracted and used to replace the variable names in the test question paradigm to obtain the test dataset.
[0182] In the absence of variable constraints in the test question paradigm, the set of variable values to be extracted is the range of variable values. The third number of variable values are extracted and used to replace the variable names in the test question paradigm to obtain the test dataset.
[0183] For example, for each test question paradigm ri and the third number of generated questions Rin for each test question paradigm, if there are variable constraints, the variable value pairs are used as the set of variable values to be extracted; otherwise, the range of variable values is used as the set of variable values to be extracted. Thus, Rin variable values / variable value pairs are extracted to replace the corresponding variable names in the test question paradigm ri, generating the test dataset.
[0184] In this embodiment of the application, the test dataset is generated by sequentially following the priority order of the main dimension of the knowledge space, the form of the question paradigm, the value of the question paradigm variable, and the synonym expression of the question paradigm through steps C1 to C5, thereby ensuring the uniformity of the distribution of test questions in the test dataset.
[0185] The following example, with 200 questions, further illustrates the method for generating the test dataset, specifically including steps Q1 to Q5.
[0186] Step Q1: Determine the number of questions to generate for testing (200) and the set of question paradigms.
[0187] Step Q2: The number of main dimensions of the question paradigm is 5 (including: disease, drug, medical record acquisition, patient indicators, and treatment process). The number of questions is greater than the number of main dimensions of the question paradigm. The number of questions is evenly distributed on each main dimension to obtain the dimension set D ("disease", "drug", "medical record acquisition", "patient indicators", "treatment process"), and the number of the first questions under each main dimension Di is Din (200 / 5=40).
[0188] Step Q3: For all question paradigms under the main dimension Di (taking "disease" as an example), select only one representative paradigm for each synonym paradigm to form the question paradigm set Pi to be extracted. The question paradigm set Pi has 15 question paradigms. The first number of questions Din (=40) is greater than the number of question paradigms in the question paradigm set Pi (15). The first number of questions is evenly distributed on each question paradigm in the question paradigm set Pi to be extracted to obtain the target question paradigm set Qi extracted under the main dimension Di of disease, and the second number of questions Qin allocated to each target question paradigm (40 / 15: 3 for 10 target question paradigms and 2 for 5 target question paradigms).
[0189] Step Q4: For each target question paradigm qi in the target question paradigm set Qi (taking "How to diagnose A103[,][X101] for C102?" as an example), and the second number of questions to be generated Qin (the number extracted in this case is 3), since they contain rule symbols, paradigms are further extracted from the intermediate paradigm set corresponding to the original paradigms, and the number of questions is evenly distributed to obtain the test question paradigm set Ri, and the third number of questions Rin for each test question paradigm. For example, the extraction results are shown in Table 1:
[0190] Table 1. Examples of Test Question Paradigm Extraction Results
[0191]
[0192] Step Q5: Variable "X101" has no variable constraints; its value range ("needs" or "should") is used as the set of variable values to be extracted. Variables "A103" and "C102" have constraints; their value pairs are used as the set of variable values to be extracted. The results after extraction are shown in Table 2.
[0193] Table 2 shows examples of extraction results using variable value pairs as the set of variable values to be extracted.
[0194]
[0195] Finally, for the set of test question paradigms Ri, the generated test questions are as follows:
[0196] A: How to perform imaging examinations to diagnose hemophilia patients with inhibitory drugs?
[0197] B: How should pathological examination be performed to diagnose kidney damage in multiple myeloma?
[0198] C: How to perform serological testing to diagnose cytomegalovirus infection in patients with blood diseases who have received allogeneic hematopoietic stem cell transplantation?
[0199] 1.3.2 Testing process of the large-scale clinical auxiliary diagnosis and treatment model:
[0200] First, the test dataset is input into the large-scale clinical auxiliary diagnosis and treatment model for processing, yielding the answer content for each test question in the dataset. Then, the answers from the large-scale clinical auxiliary diagnosis and treatment model are evaluated manually, including assessments of the understanding of the question, identification of relevant knowledge points, comprehensiveness of content, and correctness of reasoning.
[0201] For example, the test dataset is sent to a large-scale clinical auxiliary diagnosis and treatment model to obtain the answer Ai corresponding to the i-th question Qi. Multiple experts then manually evaluate the model's answers based on evaluation item Ej and reach a consensus, obtaining the total number Eij (correct / incorrect) of "correct" evaluation results for each evaluation item. Based on the total number of questions and the accuracy rate of a specific evaluation item, the confidence level of the large-scale clinical auxiliary diagnosis and treatment model for each evaluation item j is obtained across N samples. If there are 200 samples, and for the evaluation item "Is the question understood correctly?", 180 of them correctly extract the knowledge space dimension information involved in the question, then the confidence level for question understanding across 200 samples is considered to be 180 / 200 = 90%.
[0202] It should be noted that all steps of the above-mentioned methods for establishing the large-scale clinical auxiliary diagnosis and treatment model, generating answers based on the large-scale clinical auxiliary diagnosis and treatment model, and evaluating the large-scale clinical auxiliary diagnosis and treatment model are implemented through electronic devices.
[0203] 1.4 The table mentioned above is illustrated below:
[0204] Table 3: Questioning Paradigms for Different Main Dimensions
[0205]
[0206] Table 4. Meaning of the alphanumeric values of variable names
[0207]
[0208] Table 5 lists disease-related variables that begin with "A", along with examples of variable names and values.
[0209]
[0210] Table 6 lists drug-related variables that begin with "B", along with examples of their values.
[0211]
[0212] Table 7 lists the variable names for the medical record retrieval class, all starting with "C". Examples of variable names and their values are provided.
[0213]
[0214] Table 8 lists the patient indicator variables, each beginning with "D," along with examples of variable names and values.
[0215]
[0216] Table 9 lists examples of prevention-related variables that begin with "E".
[0217]
[0218] Table 10 lists the variable names for the diagnosis and prognosis categories, all beginning with "F". Examples of variable names and their values are provided.
[0219]
[0220]
[0221] Table 11 lists the variable names for treatment and efficacy categories, all beginning with "G". Examples of variable names and their values are provided.
[0222]
[0223] Table 12 lists other categories of variable names that are alternative terms unrelated to medical terminology, beginning with "X". Examples of variable names and their values are provided.
[0224]
[0225] Note: "X912" is only used in the field of hematology when combined with words such as "examination," "diagnosis," and "prognosis," so it is not defined in the vector space. That is, in the "full scope of hematological diagnosis and treatment" of this project, "X912" is not studied, but only the results of "X912 examination" are used.
[0226] Table 13 lists the variable names for combination classes, indicating that the substitute word is composed of combinations of different types of substitute words. These names begin with "Y" and include examples of variable names and their values.
[0227]
[0228] Note: If a substitute word is combined with another substitute word from the "Other Category (Substitute Words Starting with X)" category, it is not included here. For example, "B103" is essentially "drug," but "X631" is used to enrich the options. "C101" is essentially "medical record retrieval," but "X912" is used to enrich the options. "F001" is essentially "diagnosis," but "X912" is used to enrich the options.
[0229] Table 14 Variable Constraints between Variable "A001" and Variable "C102"
[0230]
[0231] Table 15 Variable Constraints between Variable "A001" and Variable "D011"
[0232]
[0233] Table 16 Variable Constraints between Variable "A001" and Variables "F502" and "F503"
[0234]
[0235] Table 17 Variable Constraints between Variable "A001" and Variable "X902"
[0236]
[0237] This application also provides an apparatus for establishing a large-scale clinical auxiliary diagnosis and treatment model, referring to... Figure 7 As shown, Figure 7 This is a schematic diagram of a device for establishing a large-scale clinical auxiliary diagnosis and treatment model according to an embodiment of this application. The device includes:
[0238] The knowledge space mapping module 710 is used to deconstruct professional literature into knowledge points according to the knowledge characteristics of clinical diagnosis and treatment, and map the knowledge points to the knowledge space;
[0239] The questioning paradigm construction module 720 is used to construct questioning paradigms, which are used to constrain the natural language expression of the question and the coordinate mapping relationship of the knowledge space.
[0240] The large model building module 730 is used to build a large professional model for clinical auxiliary diagnosis and treatment based on the knowledge space and the questioning paradigm.
[0241] In one optional embodiment, the knowledge space includes a tree-like multi-level classification tree with multiple dimensions. The multiple dimensions and the multi-level classification under each dimension constitute a knowledge space coordinate system. The multiple dimensions include: disease, drug, medical record acquisition, patient indicators, and diagnosis and treatment process.
[0242] The knowledge space mapping module includes:
[0243] The initialization module is used to take each dimension as the root node of a tree-like multi-level classification tree to obtain the initial knowledge space coordinate system;
[0244] The document tree construction module is used to extract content from the professional documents and construct a document tree to obtain the document tree.
[0245] The coordinate extraction module is used to extract knowledge space coordinates based on the text content of each document node in the document tree, forming knowledge points in the knowledge space;
[0246] The index refinement module is used to refine the initial knowledge space coordinate system based on the knowledge points and establish an index for each knowledge point to obtain the knowledge space.
[0247] In one optional embodiment, the refined index module includes:
[0248] The vector generation module is used to generate a text content vector based on the text content of each document node in the document tree;
[0249] The index building module is used to build a knowledge point vector index for each document node of the document tree in the vector database based on the text content vector.
[0250] In one alternative embodiment, the document tree building module includes:
[0251] The format conversion module is used to convert the professional documents into target format files;
[0252] The content extraction module is used to extract the text content and table content of the target format file to obtain plain text format paragraph content and table text;
[0253] The node construction module is used to take the natural paragraphs of the professional document as document nodes and establish a master-child document node relationship according to the chapter affiliation of the professional document;
[0254] The content addition module is used to add the plain text format paragraph content and the table text to the corresponding document nodes to obtain a document tree.
[0255] In one optional embodiment, the knowledge space coordinates include primary coordinates and secondary coordinates;
[0256] The coordinate extraction module includes:
[0257] The principal coordinate extraction module is used to determine the dimension in which the topic of the text content is located as the principal dimension, and to determine the category name of the text content on the principal dimension as the principal coordinate.
[0258] The coordinate extraction module is used to determine the dimension in which the preconditions of the text content are located as the condition dimension, and to determine the category name of the text content on the condition dimension as the coordinate.
[0259] In one optional embodiment, the coordinate extraction module includes:
[0260] The classification option module is used to use the multiple dimensions and one alternative dimension as classification options for the text content;
[0261] The first knowledge space coordinate module is used to determine that the knowledge space coordinates are empty coordinates when the text content is classified as the candidate dimension;
[0262] The second knowledge space coordinate module is used to perform the following steps when the text content is classified into the multiple dimensions: when the text content involves a new category, add a new category name as a knowledge space coordinate in the knowledge space coordinate system; when the text content involves an existing category with the same name, use the existing category with the same name as the knowledge space coordinate; when the text content involves an existing category with the same name but different category names, use the existing category name as the knowledge space coordinate.
[0263] In one optional embodiment, the questioning paradigm construction module includes:
[0264] The first determining module is used to determine the topic of the knowledge points involved in the question as the main dimension of the paradigm, and to determine the dimension of the premises of the knowledge points involved in the question as the condition dimension of the paradigm.
[0265] A submodule is constructed to build a question paradigm based on the main dimension and the conditional dimension.
[0266] The questioning paradigm includes the following forms: the original paradigm containing variables and rule symbols, and the intermediate paradigm containing variables.
[0267] In an optional embodiment, the construction submodule is further configured to, when the questioning paradigm includes multiple variables, determine variable constraints based on the values of the multiple variables; and construct the questioning paradigm based on the variable constraints, the main dimension, and the conditional dimension.
[0268] In an optional embodiment, the device further includes a question-answer generation module for generating questions and answers using a large-scale clinical auxiliary diagnosis and treatment model. The question-answer generation module includes:
[0269] The question acquisition module is used to acquire the target question raised by the user according to the questioning paradigm, and extract the first knowledge space coordinates from the target question. The questioning paradigm is used to constrain the natural language expression of the question and the coordinate mapping relationship of the knowledge space.
[0270] The knowledge matching module is used to match knowledge points from the knowledge space based on the coordinates of the first knowledge space to obtain target knowledge points that satisfy the coordinates of the first knowledge space.
[0271] The answer generation module is used to generate the answer to the target question based on the target knowledge point, and then provide the answer to the user.
[0272] In an optional embodiment, the device further includes:
[0273] The reference acquisition module is used to acquire the reference set corresponding to the target knowledge point;
[0274] The recommendation question generation module is used to generate a set of recommendation questions based on the knowledge space coordinates of the current dialogue context, wherein the knowledge space coordinates of the current dialogue context represent the historical knowledge space coordinates.
[0275] The feedback module is used to provide the user with the answers to the questions, the reference set, and the recommended question set.
[0276] In one optional embodiment, the knowledge matching module includes:
[0277] The retrieval routing module is used to perform vector retrieval routing based on the first knowledge space coordinates and the knowledge space coordinates of the knowledge points, narrowing the retrieval scope from all documents to documents whose root node's knowledge space coordinates intersect with the first knowledge space coordinates.
[0278] The vector matching module is used to process the target question into a target question vector, and match the target question vector with the knowledge point vector index within the retrieval range to obtain a set of semantically similar knowledge points and the similarity of each knowledge point;
[0279] The intersection calculation module is used to perform intersection calculation between the second knowledge space coordinates corresponding to each knowledge point and the first knowledge space coordinates to obtain a set of knowledge points that satisfy the first knowledge space coordinates.
[0280] The sorting module is used to sort each knowledge point in the knowledge point set according to the similarity of each knowledge point, and select the top N knowledge points as the target knowledge points, where N is an integer greater than 0.
[0281] In an optional embodiment, the apparatus further includes a maintenance module for maintaining the knowledge space coordinates of the current dialogue context, the maintenance module comprising:
[0282] The coordinate overlap module is used to ensure that the knowledge space coordinates of the current dialogue context remain unchanged when the knowledge space coordinates of the current dialogue context overlap with the first knowledge space coordinates;
[0283] The coordinate non-overlapping module is used to use the first knowledge space coordinates as the knowledge space coordinates of the current dialogue context when the knowledge space coordinates of the current dialogue context and the first knowledge space coordinates do not overlap.
[0284] This application also provides an evaluation device for a large-scale clinical auxiliary diagnosis and treatment model, referring to... Figure 8 As shown, Figure 8 This is a schematic diagram of the structure of a large-scale clinical auxiliary diagnosis and treatment model evaluation device provided in an embodiment of this application. The device includes:
[0285] The data construction module 810 is used to construct a test dataset based on the main dimension of the knowledge space, the form of the question paradigm, the value of the question paradigm variable, and the priority order of the synonym expression of the question paradigm.
[0286] The question processing module 820 is used to input the test dataset into the clinical auxiliary diagnosis and treatment professional big model for processing, and obtain the answer content of each test question in the test dataset;
[0287] The response evaluation module 830 is used to evaluate the correctness of the response content by business experts and determine the confidence level of the clinical auxiliary diagnosis and treatment professional big model based on the error rate of the evaluation results.
[0288] In one optional embodiment, the data construction module includes:
[0289] The second determining module is used to determine the number of questions and the set of question paradigms;
[0290] The third determining module is used to determine the target main dimension and the first number of questions for each target main dimension based on the number of questions and the set of question paradigms.
[0291] The fourth determining module is used to determine the target questioning paradigm under each target main dimension and the second questioning quantity for each target questioning paradigm based on the target main dimension and the first questioning quantity.
[0292] The fifth determining module is used to determine the test questioning paradigm and the third number of questions for each test questioning paradigm based on the target questioning paradigm and the second number of questions, wherein the test questioning paradigms are all in the form of intermediate paradigms;
[0293] The sixth determining module is used to determine the variable values of the questioning paradigm based on the test questioning paradigm and the third number of questions, and to generate a test dataset based on the variable values of the questioning paradigm.
[0294] In one optional embodiment, the third determining module includes:
[0295] The first random sampling module is used to randomly sample a main dimension from all main dimensions as a target main dimension when the number of questions is less than the number of main dimensions in the question paradigm set, and to determine the first number of questions for each target main dimension as 1.
[0296] The first uniform distribution module is used to, when the number of questions is not less than the number of principal dimensions in the question paradigm set, take all principal dimensions as target principal dimensions and distribute the number of questions evenly to obtain the first number of questions for each target principal dimension.
[0297] In an optional embodiment, the fourth determining module includes:
[0298] The first selection module is used to select only one representative paradigm from all the question paradigms of each target main dimension, and to form a set of question paradigms to be extracted.
[0299] The second random sampling module is used to randomly sample question paradigms from the set of question paradigms to be sampled when the number of the first question is less than the number of question paradigms in the set of question paradigms to be sampled, and to use them as the target question paradigms under each target main dimension, and to determine the second number of questions for each target question paradigm as 1.
[0300] The second uniform distribution module is used to take all the question paradigms in the question paradigm set to be extracted as target question paradigms and uniformly distribute the first question quantity, when the second number of questions is not less than the number of question paradigms in the question paradigm set to be extracted, so as to obtain the target question paradigm under each main dimension and the second number of questions for each target question paradigm.
[0301] In one optional embodiment, the fifth determining module includes:
[0302] The third uniform distribution module is used to extract intermediate paradigms from the set of intermediate paradigms corresponding to the original paradigm to form a test question paradigm when the test question paradigm is the original paradigm, and uniformly distribute the second question quantity to obtain the third question quantity for each test question paradigm.
[0303] The fourth uniform distribution module is used to, when the test questioning paradigm is the intermediate paradigm, use the target questioning paradigm as the test questioning paradigm, and the number of third questions for each test questioning paradigm is equal to the number of second questions.
[0304] In an optional embodiment, the sixth determining module includes:
[0305] The first substitution module is used to extract the number of variable pairs from the third question to replace the variable names in the test question paradigm when there are variable constraints in the test question paradigm, using the variable value pairs as the set of variable values to be extracted, to obtain the test dataset.
[0306] The second substitution module is used to extract the number of variable values from the third question to replace the variable names in the test question paradigm when there are no variable constraints in the test question paradigm, using the range of variable values as the set of variable values to be extracted, to obtain the test dataset.
[0307] This application also provides an electronic device, see embodiments thereof. Figure 9 , Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 9 As shown, the electronic device 900 includes a memory 910 and a processor 920. The memory 910 and the processor 920 are connected via a bus for communication. The memory 910 stores a computer program that can run on the processor 920 to implement the steps of the method for establishing a large-scale clinical auxiliary diagnosis and treatment model as described in the embodiments of this application, or the steps of the method for evaluating a large-scale clinical auxiliary diagnosis and treatment model as described in the embodiments of this application.
[0308] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0309] This application describes embodiments of methods and apparatus according to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0310] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0311] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0312] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0313] The above provides a detailed description of the method, evaluation method, and apparatus for establishing a large-scale clinical auxiliary diagnosis and treatment model provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for establishing a large-scale clinical auxiliary diagnosis and treatment model, characterized in that, The method includes: Based on the knowledge characteristics of clinical diagnosis and treatment, professional literature is deconstructed into knowledge points, and these knowledge points are mapped to a knowledge space. The knowledge space includes a multi-dimensional, tree-like, multi-level classification tree. These multiple dimensions and their respective multi-level classifications constitute a knowledge space coordinate system. The multiple dimensions include: disease, drug, medical record acquisition, patient indicators, and diagnosis and treatment process. The coordinates of a knowledge point in the knowledge space include primary and secondary coordinates. The primary coordinate represents the classification of the knowledge point's topic within that dimension, while the secondary coordinate represents the classification of the knowledge point's preconditions within that dimension. The relationships between knowledge points in the knowledge space include: overlapping, inclusion, and adjacency. Overlapping relationships refer to knowledge points having identical coordinates. Inclusion relationships refer to identical secondary coordinates, with the primary coordinates classifying them as a primary-child relationship. Adjacency relationships refer to identical secondary coordinates, with the primary coordinates classifying them as sibling nodes under the same primary node. Constructing a questioning paradigm includes: determining the topic of the knowledge points involved in the question as the main dimension of the paradigm, and determining the dimension of the preconditions of the knowledge points involved in the question as the conditional dimension of the paradigm; constructing a questioning paradigm based on the main dimension and the conditional dimension; the questioning paradigm is used to constrain the coordinate mapping relationship between the natural language expression of the question and the knowledge space. Based on the knowledge space and the questioning paradigm, a large-scale professional model for clinical auxiliary diagnosis and treatment is established. This involves deconstructing professional literature into knowledge points based on the knowledge characteristics of clinical diagnosis and treatment, and mapping these knowledge points to a knowledge space, including: Each dimension is used as the root node of a tree-like multi-level classification tree to obtain an initialized knowledge space coordinate system; content extraction and document tree construction are performed on the professional documents to obtain a document tree; Based on the text content of each document node in the document tree, knowledge space coordinates are extracted to form knowledge points in the knowledge space. The extraction of knowledge space coordinates based on the text content of each document node in the document tree includes at least the following steps when the text content is categorized into the multiple dimensions: when the text content involves a new category, a new category name is added to the knowledge space coordinate system as the knowledge space coordinate; when the text content involves an existing category with the same name, the existing category with the same name is used as the knowledge space coordinate; when the text content involves an existing category with the same characteristics but different category names, the existing category name is used as the knowledge space coordinate. The initial knowledge space coordinate system is refined based on the knowledge points, and an index is established for each knowledge point to obtain the knowledge space; The clinical auxiliary diagnosis and treatment professional big model generates the answer to the question according to the following steps: obtaining the target question raised by the user according to the questioning paradigm, and extracting the first knowledge space coordinates from the target question; matching knowledge points from the knowledge space according to the first knowledge space coordinates to obtain the target knowledge points that satisfy the first knowledge space coordinates; obtaining the reference set corresponding to the target knowledge points; generating the answer to the target question according to the target knowledge points, and feeding back the answer and the reference set to the user; Specifically, based on the coordinates of the first knowledge space, knowledge points are matched from the knowledge space to obtain target knowledge points that satisfy the coordinates of the first knowledge space, including: Based on the first knowledge space coordinates and the knowledge space coordinates of the knowledge points, a vector retrieval route is performed to narrow the retrieval scope from all documents to documents whose root node's knowledge space coordinates intersect with the first knowledge space coordinates. The target question is processed into a target question vector, and within the retrieval range, the target question vector is matched with the knowledge point vector index to obtain a set of semantically similar knowledge points and the similarity of each knowledge point; The intersection calculation is performed between the second knowledge space coordinates corresponding to each knowledge point and the first knowledge space coordinates to obtain the set of knowledge points that satisfy the first knowledge space coordinates. Based on the similarity of each knowledge point, each knowledge point in the knowledge point set is sorted, and the top N knowledge points in the sorting are taken as the target knowledge points, where N is an integer greater than 0.
2. The method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to claim 1, characterized in that, Create an index for each knowledge point, including: Generate a text content vector based on the text content of each document node in the document tree; Based on the text content vector, a knowledge point vector index is created for each document node of the document tree in the vector database.
3. The method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to claim 1, characterized in that, Content extraction and document tree construction are performed on the aforementioned professional documents to obtain a document tree, including: Convert the professional documents into the target format file; The text and table content of the target format file are extracted to obtain plain text paragraph content and table text; The natural paragraphs of the professional documents are used as document nodes, and a master-child document node relationship is established according to the chapter affiliation of the professional documents. Add the plain text paragraph content and the table text to the corresponding document nodes to obtain the document tree.
4. The method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to claim 1, characterized in that, The knowledge space coordinates include primary coordinates and secondary coordinates; Based on the text content of each document node in the document tree, extract the knowledge space coordinates, including: The dimension in which the topic of the text content is located is determined as the primary dimension, and the category name of the text content on the primary dimension is determined as the primary coordinate. The dimension containing the preconditions of the text content is defined as the condition dimension, and the category name of the text content on the condition dimension is defined as the coordinate.
5. The method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to claim 1, characterized in that, Based on the text content of each document node in the document tree, extract the knowledge space coordinates, including: The multiple dimensions and one alternative dimension are used as classification options for the text content; When the text content is classified as the candidate dimension, the knowledge space coordinates are determined to be empty coordinates.
6. The method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to claim 1, characterized in that, The questioning paradigms include: the original paradigm containing variables and rule symbols, and the intermediate paradigm containing variables.
7. The method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to claim 6, characterized in that, Based on the main dimension and the conditional dimension, a questioning paradigm is constructed, including: When the question paradigm includes multiple variables, variable constraints are determined based on the values of the multiple variables; Based on the variable constraints, the main dimension, and the conditional dimension, construct the questioning paradigm.
8. The method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to claim 1, characterized in that, The method further includes: A set of recommended questions is generated based on the knowledge space coordinates of the current dialogue context, where the knowledge space coordinates of the current dialogue context represent the historical knowledge space coordinates. The answers to the questions and the set of references will be provided back to the user, including: The answers to the questions, the set of references, and the set of recommended questions will be fed back to the user.
9. The method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to claim 8, characterized in that, The knowledge space coordinates of the current dialogue context are maintained in the following manner: If the knowledge space coordinates of the current dialogue context overlap with the first knowledge space coordinates, the knowledge space coordinates of the current dialogue context remain unchanged. If the knowledge space coordinates of the current dialogue context and the first knowledge space coordinates do not overlap, the first knowledge space coordinates shall be used as the knowledge space coordinates of the current dialogue context.
10. A method for evaluating a large-scale clinical auxiliary diagnosis and treatment model, characterized in that, The large-scale clinical auxiliary diagnosis and treatment model is constructed using the method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to any one of claims 1-9, wherein the method includes: A test dataset is constructed based on the priority order of the main dimensions of the knowledge space, the form of the question paradigm, the values of the question paradigm variables, and the synonymous expression of the question paradigm. The test dataset is input into the clinical auxiliary diagnosis and treatment professional big model for processing to obtain the answer content of each test question in the test dataset; The correctness of the answers is evaluated by business experts, and the confidence level of the clinical auxiliary diagnosis and treatment professional model is determined based on the error rate of the evaluation results. The process involves constructing a test dataset based on the priority order of the main dimension of the knowledge space, the form of the question paradigm, the variable values of the question paradigm, and the synonymous expression of the question paradigm. This includes: determining the number of questions and the set of question paradigms; determining the target main dimension and the first number of questions for each target main dimension based on the number of questions and the set of question paradigms; determining the target question paradigm and the second number of questions for each target question paradigm based on the target main dimension and the first number of questions; determining the test question paradigm and the third number of questions for each test question paradigm based on the target question paradigm and the second number of questions, wherein all test question paradigms are intermediate paradigms; and determining the variable values of the question paradigms based on the test question paradigms and the third number of questions, and generating the test dataset based on the variable values of the question paradigms.
11. The evaluation method for a large-scale clinical auxiliary diagnosis and treatment model according to claim 10, characterized in that, Based on the number of questions and the set of question paradigms, determine the target main dimension and the first number of questions for each target main dimension, including: If the number of questions is less than the number of main dimensions in the question paradigm set, a main dimension is randomly selected from all main dimensions as the target main dimension, and the first number of questions for each target main dimension is determined to be 1. If the number of questions is not less than the number of principal dimensions in the question paradigm set, all principal dimensions are taken as target principal dimensions, and the number of questions is evenly distributed to obtain the first number of questions for each target principal dimension.
12. The evaluation method for a large-scale clinical auxiliary diagnosis and treatment model according to claim 10, characterized in that, Based on the stated primary target dimension and the first number of questions, determine the target questioning paradigm for each primary target dimension and the second number of questions for each target questioning paradigm, including: For all question paradigms of each target main dimension, only one representative paradigm is selected from all synonym paradigms to form a set of question paradigms to be extracted. If the number of questions in the first question is less than the number of question paradigms in the set of question paradigms to be extracted, question paradigms are randomly extracted from the set of question paradigms to be extracted as target question paradigms under each target main dimension, and the number of questions in the second question of each target question paradigm is determined to be 1. If the second number of questions is not less than the number of question paradigms in the set of question paradigms to be extracted, all question paradigms in the set of question paradigms to be extracted are taken as target question paradigms, and the first number of questions is evenly distributed to obtain the target question paradigm under each main dimension and the second number of questions for each target question paradigm.
13. The evaluation method for a large-scale clinical auxiliary diagnosis and treatment model according to claim 10, characterized in that, Based on the target questioning paradigm and the second number of questions, determine the test questioning paradigm and the third number of questions for each test questioning paradigm, including: When the test questioning paradigm is the original paradigm, intermediate paradigms are extracted from the intermediate paradigm set corresponding to the original paradigm to form the test questioning paradigm, and the second questioning quantity is evenly distributed to obtain the third questioning quantity for each test questioning paradigm. When the test questioning paradigm is an intermediate paradigm, the target questioning paradigm is used as the test questioning paradigm, and the number of third questions in each test questioning paradigm is equal to the number of second questions.
14. The evaluation method for a large-scale clinical auxiliary diagnosis and treatment model according to claim 10, characterized in that, Based on the test questioning paradigm and the third number of questions, determine the variable values for the questioning paradigm, and generate a test dataset based on the variable values for the questioning paradigm, including: When there are variable constraints in the test question paradigm, the variable value pairs are used as the set of variable values to be extracted. The third number of variable pairs are extracted and used to replace the variable names in the test question paradigm to obtain the test dataset. In the absence of variable constraints in the test question paradigm, the set of variable values to be extracted is the range of variable values. The third number of variable values are extracted and used to replace the variable names in the test question paradigm to obtain the test dataset.
15. A device for establishing a large-scale clinical auxiliary diagnosis and treatment model, characterized in that, The device includes: The knowledge space mapping module is used to deconstruct professional literature into knowledge points based on the knowledge characteristics of clinical diagnosis and treatment, and map these knowledge points to a knowledge space. The knowledge space includes a multi-dimensional tree-like multi-level classification system. These multiple dimensions and their respective multi-level classifications constitute a knowledge space coordinate system. The multiple dimensions include: disease, drug, medical record acquisition, patient indicators, and diagnosis and treatment process. The coordinates of a knowledge point in the knowledge space include primary coordinates and secondary coordinates. The primary coordinates represent the classification of the knowledge point's topic in the dimension, and the secondary coordinates represent the classification of the knowledge point's preconditions in the dimension. The relationships between knowledge points in the knowledge space include: overlapping relationships, inclusion relationships, and adjacent relationships. Overlapping relationships refer to knowledge points having identical coordinates. Inclusion relationships refer to knowledge points having identical secondary coordinates and a primary coordinate classification of the primary coordinate as a primary-child relationship. Adjacent relationships refer to knowledge points having identical secondary coordinates and a primary coordinate classification of the primary coordinate as sibling nodes under the same primary node. A question paradigm construction module is used to construct a question paradigm, including: determining the topic of the knowledge points involved in the question as the main dimension of the paradigm, and determining the dimension of the preconditions of the knowledge points involved in the question as the conditional dimension of the paradigm; constructing a question paradigm based on the main dimension and the conditional dimension; the question paradigm is used to constrain the coordinate mapping relationship between the natural language expression of the question and the knowledge space. The large model building module is used to build a large professional model for clinical auxiliary diagnosis and treatment based on the knowledge space and the questioning paradigm. The knowledge space mapping module includes: The initialization module is used to take each dimension as the root node of a tree-like multi-level classification tree to obtain the initial knowledge space coordinate system; The document tree construction module is used to extract content from the professional documents and construct a document tree to obtain the document tree. The coordinate extraction module is used to extract knowledge space coordinates based on the text content of each document node in the document tree to form knowledge points in the knowledge space. It includes at least a second knowledge space coordinate module, used to perform the following steps when the text content is classified into the multiple dimensions: when the text content involves a new classification, adding a new classification name as a knowledge space coordinate in the knowledge space coordinate system; when the text content involves an existing classification with the same name, using that classification as the knowledge space coordinate; when the text content involves an existing classification with the same characteristics but different classification names, using the existing classification name as the knowledge space coordinate. The index refinement module is used to refine the initial knowledge space coordinate system according to the knowledge points and establish an index for each knowledge point to obtain the knowledge space. The device further includes a question-and-answer generation module for generating questions and answers in a large-scale clinical auxiliary diagnosis and treatment model. The question-and-answer generation module includes: The question acquisition module is used to acquire the target question raised by the user according to the questioning paradigm, and extract the first knowledge space coordinates from the target question. The questioning paradigm is used to constrain the natural language expression of the question and the coordinate mapping relationship of the knowledge space. The knowledge matching module is used to match knowledge points from the knowledge space based on the coordinates of the first knowledge space to obtain target knowledge points that satisfy the coordinates of the first knowledge space. The reference acquisition module is used to acquire the reference set corresponding to the target knowledge point; The answer generation module is used to generate the answer to the target question based on the target knowledge point, and to provide the answer and the reference set to the user. The knowledge matching module includes: The retrieval routing module is used to perform vector retrieval routing based on the first knowledge space coordinates and the knowledge space coordinates of the knowledge points, narrowing the retrieval scope from all documents to documents whose root node's knowledge space coordinates intersect with the first knowledge space coordinates. The vector matching module is used to process the target question into a target question vector, and match the target question vector with the knowledge point vector index within the retrieval range to obtain a set of semantically similar knowledge points and the similarity of each knowledge point; The intersection calculation module is used to perform intersection calculation between the second knowledge space coordinates corresponding to each knowledge point and the first knowledge space coordinates to obtain a set of knowledge points that satisfy the first knowledge space coordinates. The sorting module is used to sort each knowledge point in the knowledge point set according to the similarity of each knowledge point, and select the top N knowledge points as the target knowledge points, where N is an integer greater than 0.
16. A testing device for a large-scale clinical auxiliary diagnosis and treatment model, characterized in that, The large-scale clinical auxiliary diagnosis and treatment model is constructed using the method for establishing a large-scale clinical auxiliary diagnosis and treatment model according to any one of claims 1-9, and the device includes: The data construction module is used to construct a test dataset based on the main dimensions of the knowledge space, the form of the question paradigm, the values of the question paradigm variables, and the priority order of the synonym expressions of the question paradigm. The question processing module is used to input the test dataset into the clinical auxiliary diagnosis and treatment professional big model for processing, and obtain the answer content of each test question in the test dataset; The response evaluation module is used to assess the correctness of the response content through business experts, and to determine the confidence level of the clinical auxiliary diagnosis and treatment professional big model based on the error rate of the evaluation results. The data construction module includes: The second determining module is used to determine the number of questions and the set of question paradigms; The third determining module is used to determine the target main dimension and the first number of questions for each target main dimension based on the number of questions and the set of question paradigms. The fourth determining module is used to determine the target questioning paradigm under each target main dimension and the second questioning quantity for each target questioning paradigm based on the target main dimension and the first questioning quantity. The fifth determining module is used to determine the test questioning paradigm and the third number of questions for each test questioning paradigm based on the target questioning paradigm and the second number of questions, wherein the test questioning paradigms are all in the form of intermediate paradigms; The sixth determining module is used to determine the variable values of the questioning paradigm based on the test questioning paradigm and the third number of questions, and to generate a test dataset based on the variable values of the questioning paradigm.
17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of establishing a large-scale clinical auxiliary diagnosis and treatment model as described in any one of claims 1-9, or the steps of evaluating a large-scale clinical auxiliary diagnosis and treatment model as described in any one of claims 10-14.