Assessment system and method for application of large language model for auxiliary diagnosis and treatment of specific disease

By constructing a large language model evaluation system for assisted diagnosis and treatment of specific diseases, using RAG retrieval strategy and multi-dimensional evaluation data sets, the inconsistency and ethical and legal risks of the medical large language model evaluation system are solved, the accuracy and comprehensive evaluation of the model's capabilities are achieved, and the applicability and effectiveness boundaries of the model are optimized.

CN120373471AActive Publication Date: 2025-07-25CHINA REHABILITATION SCIENCE INSTITUTE (DISABILITY PREVENTION AND CONTROL RESEARCH CENTER OF CHINA DISABLED PERSONS FEDERATION)

Patent Information

Application Number
CN202510855240.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The existing technology of the large language model evaluation system for traditional Chinese medicine has inconsistent evaluation standards, outstanding heterogeneity of application scenarios, complex data quality and ethical legal risks, making it difficult to effectively evaluate the applicability and effectiveness boundaries of the model in specific clinical scenarios.

Method used

Build an evaluation system for the application of large language models for assisted diagnosis and treatment of specific diseases, including the model layer, the base model evaluation layer, the basic capability evaluation layer, the evidence link evaluation layer and the compliance evaluation layer. Through RAG retrieval strategy, the multi-dimensional evaluation data set and the privacy ethics evaluation data set, the ability and accuracy of the evaluation model are quantified.

Benefits of technology

It improves the accuracy and comprehensiveness of the application ability of large language models for assisted diagnosis and treatment of specific diseases, quickly locates model defects, provides a clear path for the optimization and improvement of medical models, and forms a reusable medical large model engineering path and quality control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373471A_ABST
    Figure CN120373471A_ABST
Patent Text Reader

Abstract

The invention provides an evaluation system and method for a large language model application of auxiliary diagnosis and treatment of specific diseases, and relates to the technical field of clinical medicine, auxiliary diagnosis and treatment and artificial intelligence, and the system comprises a model layer which is used for constructing an RAG retrieval strategy to adapt to the large language model application of the auxiliary diagnosis and treatment of the specific diseases; the base model evaluation layer is used for carrying out basic capability evaluation on the auxiliary diagnosis and treatment large language model base LLM in the knowledge field of a specific disease category; the basic capability evaluation layer is used for carrying out correctness evaluation on an output result of the application of the specific disease auxiliary diagnosis and treatment big language model; the evidence chain evaluation layer is used for carrying out evidence chain integrity evaluation and evidence timeliness evaluation on an output result of the specific disease type auxiliary diagnosis and treatment big language model application; and the compliance evaluation layer is used for evaluating the privacy processing and ethical problem processing capability of the specific disease type auxiliary diagnosis and treatment large language model application. According to the method, the accuracy and comprehensiveness of application capability evaluation of the large language model for auxiliary diagnosis and treatment of specific diseases are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of clinical medicine, auxiliary diagnosis and treatment, and artificial intelligence, and in particular to an evaluation system and method for the application of a large language model for auxiliary diagnosis and treatment of specific diseases. Background Art

[0002] In recent years, the application of large language models (LLMs) in the medical field has become increasingly prominent. However, the evaluation system for medical LLMs is still in the development stage, facing practical challenges such as inconsistent evaluation criteria, prominent heterogeneity in application scenarios, complex data quality, and ethical and legal risks. When current medical institutions are promoting the construction of medical knowledge Q&A-assisted diagnosis and treatment systems, they generally face the dilemma of lacking targeted verification methods and it is difficult to effectively evaluate the applicability and effectiveness boundaries of the model in specific clinical scenarios. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide an evaluation system and method for the application of a large language model for auxiliary diagnosis and treatment of specific diseases, so as to improve the effectiveness and accuracy of the large language model in the auxiliary diagnosis and treatment of specific diseases during the clinical medical process, form a reusable engineering path and quality control system for medical large models, and provide methodological support and practical solutions for the implementation of "artificial intelligence + medical" in specialized scenarios.

[0004] To achieve the above purpose, the technical solution adopted by the present invention is as follows: In the first aspect, the present invention provides an evaluation system for the application of a large language model for auxiliary diagnosis and treatment of specific diseases, including: A model layer, which is used to construct a RAG retrieval strategy to build or adjust the application of a large language model for auxiliary diagnosis and treatment of specific diseases based on the RAG retrieval strategy; A base model evaluation layer, which is used to perform multi-dimensional evaluations on the base LLM of the large language model for auxiliary diagnosis and treatment for specific diseases, so as to evaluate the basic capabilities of the base LLM of the large language model for auxiliary diagnosis and treatment in the knowledge field of specific diseases; A basic capability evaluation layer, which is used to construct various forms of evaluation data sets for the knowledge field of specific diseases, and perform correctness evaluations on the output results of the application of the large language model for auxiliary diagnosis and treatment of specific diseases based on the evaluation data sets; among them, the evaluation data sets are divided into 4 fields for specific diseases, namely: definition and interpretation of medical knowledge, symptom recognition and diagnosis, design and selection of treatment plans, and medical literature retrieval; the evaluation data sets for each field are in the form of a combination of clinical questions + reference answers for the knowledge field of specific diseases; An evidence chain evaluation layer, which is used to perform evidence chain integrity evaluation and evidence timeliness evaluation on the output results of the application of the large language model for auxiliary diagnosis and treatment of specific diseases based on the evaluation data sets of the basic capability evaluation layer; A compliance evaluation layer for constructing a privacy evaluation dataset and an ethics evaluation dataset based on the knowledge domain of specific diseases, and evaluating the privacy processing and ethics handling capabilities of the large language model application for specific disease assisted diagnosis and treatment based on the privacy evaluation dataset and the ethics evaluation dataset.

[0005] In a second aspect, the present invention provides an evaluation method for a large language model application for specific disease assisted diagnosis and treatment, which is applied to the evaluation system of the large language model application for specific disease assisted diagnosis and treatment according to any one of the first aspect, and includes: Construct a RAG retrieval strategy to build or adjust the large language model application for specific disease assisted diagnosis and treatment based on the RAG retrieval strategy; Conduct multi-dimensional evaluation of the large language model base LLM for specific diseases to evaluate the basic capabilities of the large language model base LLM in the knowledge domain of specific diseases; Construct various forms of evaluation datasets for the knowledge domain of specific diseases, and conduct correctness evaluation on the output results of the large language model application for specific disease assisted diagnosis and treatment based on the evaluation datasets; among them, the evaluation datasets are divided into 4 fields for specific diseases, namely: definition and interpretation of medical knowledge, symptom recognition and diagnosis, design and selection of treatment plans, and medical literature retrieval; the evaluation datasets for each field are in the form of a combination of clinical questions + reference answers for the knowledge domain of specific diseases; Conduct evidence chain integrity evaluation and evidence timeliness evaluation on the output results of the large language model application for specific disease assisted diagnosis and treatment based on the evaluation datasets of the basic capability evaluation layer; Construct a privacy evaluation dataset and an ethics evaluation dataset based on the knowledge domain of specific diseases, and evaluate the privacy processing and ethics handling capabilities of the large language model application for specific disease assisted diagnosis and treatment based on the privacy evaluation dataset and the ethics evaluation dataset.

[0006] In a third aspect, the present invention provides an electronic device, including a processor and a memory, the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the steps of the method provided in the second aspect above.

[0007] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the method provided in the second aspect above.

[0008] The present invention brings the following beneficial effects: The present invention provides an evaluation system and method for the application of the above-mentioned large language model for assisting in the diagnosis and treatment of specific diseases, including: a model layer for constructing a RAG retrieval strategy to build or adjust the application of the large language model for assisting in the diagnosis and treatment of specific diseases based on the RAG retrieval strategy; a base model evaluation layer for performing multi-dimensional evaluations of the base LLM of the large language model for assisting in the diagnosis and treatment of specific diseases to evaluate the basic capabilities of the base LLM of the large language model for assisting in the diagnosis and treatment of specific diseases in the knowledge field of specific diseases; a basic capability evaluation layer for constructing various forms of evaluation data sets for the knowledge field of specific diseases and evaluating the correctness of the output results of the application of the large language model for assisting in the diagnosis and treatment of specific diseases based on the evaluation data sets; among which, the evaluation data sets are divided into 4 fields for specific diseases, namely: definition and interpretation of medical knowledge, symptom recognition and diagnosis, design and selection of treatment plans, and medical literature retrieval; the evaluation data sets for each field are in the form of a combination of clinical questions + reference answers for the knowledge field of specific diseases; an evidence chain evaluation layer for evaluating the integrity and timeliness of the evidence chain of the output results of the application of the large language model for assisting in the diagnosis and treatment of specific diseases based on the evaluation data sets of the basic capability evaluation layer; a compliance evaluation layer for constructing a privacy evaluation data set and an ethics evaluation data set based on the knowledge field of specific diseases and evaluating the privacy processing and ethics handling capabilities of the application of the large language model for assisting in the diagnosis and treatment of specific diseases.

[0009] In the above system, during the process of building a large language model for assisting in the diagnosis and treatment of a specific disease based on the RAG technology, an innovative RAG retrieval strategy is used to cooperate with the implementation of each subsequent evaluation layer; by performing multi-dimensional evaluations of the base LLM of the large language model for assisting in the diagnosis and treatment of specific diseases to evaluate the basic capabilities of the initial base LLM in the knowledge field of the specific disease; by constructing various forms of evaluation data sets for the knowledge field of specific diseases and evaluating the correctness of the output results of the application of the large language model for assisting in the diagnosis and treatment of specific diseases based on the RAG technology; by evaluating the effectiveness and timeliness of the evidence chain of the output results of the application of the large language model for assisting in the diagnosis and treatment of specific diseases in the basic capability evaluation layer; by constructing a privacy evaluation data set and an ethics evaluation data set based on the knowledge field of specific diseases to evaluate the privacy processing and ethics handling capabilities of the large language model for assisting in the diagnosis and treatment. The present invention improves the accuracy and comprehensiveness of the evaluation of the application capabilities of the large language model for assisting in the diagnosis and treatment of specific diseases through quantitative evaluation, quickly locates the defects of the model capabilities, and provides a clear path for the optimization and improvement of medical models.

[0010] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention are realized and attained by the structure particularly pointed out in the description, claims and drawings.

[0011] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and describes them in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 Schematic diagram of the structural composition of an evaluation system for the application of a large language model for assisting in the diagnosis and treatment of specific diseases provided by an embodiment of the present invention; Figure 2 Schematic diagram of the process of another evaluation system for the application of a large language model for assisting in the diagnosis and treatment of specific diseases provided by an embodiment of the present invention; Figure 3 Schematic diagram of the principle of the RAG retrieval strategy based on semantic induction - metadata - two - level retrieval provided by an embodiment of the present invention; Figure 4 Schematic diagram of the data set distribution ratio of the basic ability evaluation layer for the application of a large language model for assisting in the diagnosis and treatment of specific diseases provided by an embodiment of the present invention; Figure 5 Flowchart of a method for evaluating the application of a large language model for assisting in the diagnosis and treatment of specific diseases provided by an embodiment of the present invention; Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0015] In recent years, the application of large language models in the medical field has become increasingly prominent. However, the evaluation system for medical large language models is still in the development stage, facing practical challenges such as inconsistent evaluation criteria, prominent heterogeneity in application scenarios, complex data quality, and ethical and legal risks. When current medical institutions are promoting the construction of medical knowledge Q&A-assisted diagnosis and treatment systems, they generally face the dilemma of lacking targeted verification methods and it is difficult to effectively evaluate the applicability and effectiveness boundaries of the model in specific clinical scenarios.

[0016] In view of this, the present invention provides an evaluation system and method for the application of a large language model for auxiliary diagnosis and treatment of specific diseases, which can specifically and comprehensively discover the defects of medical large models in a quantitative index manner, so as to improve the effectiveness and accuracy of the large language model in assisting diagnosis and treatment of specific diseases in the clinical medical process, and form a reusable engineering path and quality control system for medical large models.

[0017] To facilitate the understanding of this embodiment, first, a detailed introduction is given to an evaluation system for the application of a large language model for auxiliary diagnosis and treatment of specific diseases disclosed in the embodiments of the present invention. Figure 1 It is a schematic diagram of the structural composition of an evaluation system for the application of a large language model for auxiliary diagnosis and treatment of specific diseases provided in an embodiment of the present invention; Figure 2 It is a schematic diagram of the main process of another evaluation system for the application of a large language model for auxiliary diagnosis and treatment of specific diseases provided in an embodiment of the present invention (taking the evidence chain evaluation process as an example). See Figure 1 This system mainly includes the following parts: a model layer 101, a base model evaluation layer 102, a basic ability evaluation layer 103, an evidence chain evaluation layer 104, and a compliance evaluation layer 105.

[0018] The model layer 101 is used to construct a RAG retrieval strategy to build or adjust the application of a large language model for auxiliary diagnosis and treatment of specific diseases based on the RAG retrieval strategy.

[0019] In specific implementation, the application of a large language model for auxiliary diagnosis and treatment of specific diseases based on the RAG technology can be obtained by adjusting an existing application of a large language model for auxiliary diagnosis and treatment of specific diseases through the RAG retrieval strategy, or can be rebuilt based on the large language model through the RAG retrieval strategy. In the process of building or adjusting the application of a large language model for auxiliary diagnosis and treatment of specific diseases based on the RAG technology for a specific disease, this embodiment uses an innovative RAG retrieval strategy to cooperate with the implementation of subsequent evaluation layers.

[0020] In the embodiments of the present invention, the large language model for the auxiliary diagnosis and treatment of specific diseases needs to be based on the Retrieval-Augmented Generation (RAG) technology. Different RAG strategies have differences in performance, which depend on the used retrievers, generators, and the interaction methods between them. Considering that the materials in the knowledge field of specific diseases, such as diagnosis and treatment guidelines, frontier papers, report documents, etc., have characteristics such as diverse terms that are prone to ambiguity, complex text structures and logical dependencies, uneven word frequency distributions, multiple themes, and long texts, ordinary strategies such as the "fixed-length segmentation strategy" or the long text recommendation strategy "semantics-based segmentation strategy" have poor effects; at the same time, in order to ensure the subsequent evaluation of the evidence chain, it is necessary to trace back layer by layer from the LLM output to the specific paragraphs in the original literature or diagnosis and treatment guidelines, as well as the metadata of these paragraphs (such as the publication / writing year, author, book / journal name information of the document to which they belong).

[0021] Based on the above two requirements, the present invention constructs a RAG retrieval strategy of "semantics induction-metadata-two-level retrieval". Among them, the RAG retrieval strategy includes: segmenting the long document by paragraphs, and using the LLM to perform explicit semantics induction on each paragraph, as well as extracting the semantic labels of the paragraphs, the publication year of the document to which the paragraphs belong, the author, and the book / journal name, and constructing the mapping relationship between the semantic labels, the publication year, the author, the book / journal name and the original paragraphs; merging the semantic labels into a semantic label document, and after segmenting the semantic label document, generating a semantic vector library through the LLM.

[0022] When the user queries, the question input by the user is vectorized and transformed to obtain a question vector, and queries are made in the semantic vector library based on the question vector and multiple semantic labels related to the question vector are returned; based on the mapping relationship between the semantic labels and the original paragraphs, the original paragraphs corresponding to each semantic label are located, and based on the question input by the user, the returned original paragraphs and the prompt word template, the LLM outputs the result; among them, the output result of the LLM includes: each original paragraph and the original paragraph metadata referred to in the answer, and the original paragraph metadata includes but is not limited to the publication year of the document to which the original paragraph belongs, the author, and the book / journal name.

[0023] Such as Figure 3As shown in the figure, its core principle is as follows: The long document is segmented by paragraphs, and the LLM is used to perform explicit semantic induction on each paragraph. At the same time, the metadata of each paragraph is extracted (such as the publication year, author, and book / journal name information of the document to which it belongs). The semantic tags, the original paragraphs, and the metadata of the original paragraphs are mapped one by one to form a mapping relationship and saved in the mapping library for summary. Subsequently, the semantic tags are merged into a document and segmented again to generate a semantic vector library through the LLM. When the user queries, the input question is first vectorized and transformed, the semantic vector library is queried and the most relevant multiple semantic tags are returned, and then the original paragraphs are located according to the mapping relationship. Finally, based on the question, the returned paragraphs, and the prompt template, the LLM generates an answer. When the answer is output, the original paragraphs, the semantic tags to which they belong, and the metadata are output synchronously.

[0024] In some embodiments, the specific disease of the present application may be neurogenic bladder disease in the field of urology. Among them, neurogenic bladder (NB) has a high clinical incidence, complex etiology, and poses a serious threat to the quality of life and health of patients. If not intervened in time, it is likely to cause upper urinary tract damage, and in severe cases, it can progress to renal failure or end-stage renal disease, and requires hemodialysis or kidney transplantation for treatment.

[0025] Based on this, in some implementations, an example of the mapping library in the field of neurogenic bladder is shown in Table 1.

[0026] Table 1 Data example of the mapping library for neurogenic bladder

[0027] The base model evaluation layer 102 is used to assist the large language model base LLM for auxiliary diagnosis and treatment to perform multi-dimensional evaluation for a specific disease, so as to evaluate the basic capabilities of the large language model base LLM for auxiliary diagnosis and treatment in the knowledge field of the specific disease.

[0028] In the embodiments of the present invention, for a specific disease, the initial base LLM is evaluated in terms of the degree of mastery of professional knowledge of the specific disease, semantic understanding and logical reasoning, and the performance of the retrieval module.

[0029] In some embodiments, the specific disease of the present application may be neurogenic bladder disease in the field of urology.

[0030] Based on this, in some implementations, the base model evaluation layer 102 is specifically used for: (1) Using the base LLM selected by the large language model for auxiliary diagnosis and treatment of a specific disease, based on the pre-constructed medical entity dataset in the clinical medical field for the specific disease, performing a medical entity classification task, and calculating the first performance evaluation index of the degree of mastery of professional knowledge of the specific disease by the base LLM based on the medical entity classification result.

[0031] Among them, in order to evaluate the expertise of the LLM in neurogenic bladder in the field of urology, first, an entity dataset for clinical medical treatment for specific diseases (including the fields of urology and non-urology) is constructed. Then, the base LLM is made to perform entity classification tasks for this entity dataset, that is, to judge whether the entity belongs to a noun in the field of urology. The judgment results are "yes" and "no". Finally, according to the task execution results, the first performance evaluation index for the expertise of the base LLM in neurogenic bladder in the field of urology is calculated.

[0032] Specifically, the entity dataset is created to evaluate the expertise of the initial base LLM. It is oriented towards the field of clinical medical treatment, including 50% entities in the field of urology and 50% entities in the non-urology field. The character encoding is UTF-8 and it is stored in a csv format file. Data examples are shown in Table 2.

[0033] Table 2 Data examples of the entity dataset

[0034] The first performance evaluation index includes: Precision, Recall, and F1-Score. Among them, the calculation methods for each index are as follows: Precision :

[0035] Recall :

[0036] F1-Score :

[0037] Among them, TP (True Positive): The number of true positive examples, that is, the number of samples judged as positive being "yes".

[0038] TN (True Negative): The number of true negative examples, that is, the number of negative samples judged as "no".

[0039] FP (False Positive): The number of false positive examples, that is, the number of negative samples judged as "yes".

[0040] FN (False Negative): The number of false negative examples, that is, the number of positive samples judged as "no".

[0041] (2) Using the selected base LLM for the application of the large language model for auxiliary diagnosis and treatment of specific diseases, based on the pre-constructed dataset of case descriptions related to specific diseases, perform a medical entity extraction task, and calculate the second performance evaluation index of the semantic understanding and logical reasoning ability of the base LLM for specific diseases based on the results of the medical entity extraction.

[0042] In the embodiments of the present invention, in order to evaluate the semantic understanding and logical reasoning ability of the base LLM in the medical field, first, a dataset of case descriptions for a specific disease (neurogenic bladder disease in the urology field in this embodiment) is constructed, and then the base LLM is made to perform a medical entity extraction task for this dataset of case descriptions respectively. The medical entities extracted from each case description data include, but are not limited to: patient name, gender, age, symptoms, diseases, examinations, treatments, medications, and family history. Finally, calculate the second performance evaluation index of the semantic understanding and logical reasoning ability of the base LLM for neurogenic bladder disease in the urology field. The second performance evaluation index in the embodiments of the present invention is precision rate, recall rate, and F1 value. Among them, the calculation methods of each index can refer to the calculation formulas in the foregoing embodiments.

[0043] Specifically, the dataset of case descriptions is created to evaluate the semantic understanding and logical reasoning ability of the initial base LLM. It is oriented to a specific disease field, and its content is the text of patient cases and the correct structured data expected to be extracted by the initial base LLM. The character encoding is UTF-8 and it is stored in a csv format file. Data examples are shown in Table 3.

[0044] Table 3 Data examples of the dataset of case descriptions

[0045] (3) Using the selected base LLM for the application of the large language model for auxiliary diagnosis and treatment of specific diseases, perform text segmentation, vector library construction, and retriever generation tasks on the pre-constructed dataset of specific disease descriptions, perform keyword retrieval based on the vector library and the retriever, and calculate the third performance evaluation index of the retriever module performance of the base LLM for specific diseases based on the retrieval results.

[0046] In the embodiments of the present invention, in order to evaluate the performance of the RAG retriever module of the base LLM, first, a dataset of specific disease descriptions is constructed, that is, a dataset of neurogenic bladder disease descriptions for NB. Then, the base LLM is made to perform text segmentation, vector library, and retriever generation tasks on the text of this dataset of neurogenic bladder disease descriptions, perform content retrieval for multiple retrieval keywords such as "lesion on the pons", and calculate the third performance evaluation index of the retriever module performance of the base LLM for neurogenic bladder disease in the urology field according to the retrieval. The third performance evaluation index includes: hit rate and reciprocal of average rank, and the calculation formulas are as follows: Hit rate :

[0047] Among them, represents the number of correct queries found in k searches, represents the total number of queries.

[0048] The reciprocal of the average rank :

[0049] Among them, Q represents the set of queries, is the rank of the first relevant item in the i th query.

[0050] Specifically, the neurogenic bladder disease description dataset is created for the performance of the RAG retrieval module of the base LLM. It is oriented to the field of neurogenic bladder disease, and its content is the text content about the symptoms, diagnosis, treatment, test and examination of this disease intercepted from the special chapter on neurogenic bladder in the "Diagnosis and Treatment Guidelines for Urological and Andrological Diseases in China". The character encoding is UTF-8 and it is stored in a txt format file. Data examples are shown in Table 4.

[0051] Table 4 Data examples of the neurogenic bladder disease description dataset

[0052] 4) Based on the first performance evaluation index, the second performance evaluation index, and the third performance evaluation index, conduct a multi-dimensional evaluation of the selected base LLM for the application of the large language model for auxiliary diagnosis and treatment of each specific disease type in the knowledge field of the specific disease type, and determine the basic capabilities of the base LLM in the knowledge field of the specific disease type.

[0053] In one implementation, the weight coefficients of the first performance evaluation index, the second performance evaluation index, and the third performance evaluation index can be preset in advance, and then the first performance evaluation index, the second performance evaluation index, and the third performance evaluation index of each base LLM are weighted and calculated according to the weight coefficients to obtain the comprehensive performance evaluation index of each base LLM.

[0054] The basic ability evaluation layer 103 is used to construct evaluation data sets in various forms for specific disease knowledge domains, and to evaluate the correctness of the output results of the application of the large language model for specific disease assisted diagnosis and treatment based on the evaluation data sets; among them, the evaluation data sets are divided into 4 domains for the specific disease, namely: definition and interpretation of medical knowledge, symptom recognition and diagnosis, design and selection of treatment plans, and medical literature retrieval; the evaluation data sets for each domain are in the form of a combination of clinical questions + reference answers for the specific disease knowledge domain.

[0055] In one implementation, the basic ability evaluation layer 103 is specifically used to: evaluate the correctness of the answers batch-output by the application of the large language model for assisted diagnosis and treatment based on the RAG technology through manual evaluation and / or the LLM evaluation model; among them, the LLM evaluation model and the base LLM are different LLMs; the LLM evaluation model is set as the medical expert role for a specific disease, and the output of the manual evaluation and / or the LLM evaluation model is correct or incorrect.

[0056] The basic ability evaluation layer 103 aims to rigorously measure the correctness of the professional knowledge answers of the assisted diagnosis and treatment large model for a specific disease itself through the following several scientific methods: (1)Human Evaluation Due to the rigor of this embodiment for the medical field, the link of manual evaluation by clinical professional doctors needs to be considered to make up for the possible limitations of automated evaluation. That is, based on their professional knowledge and practical experience, clinical doctors review the query results output by the application one by one to judge whether they are correct or not, and use CORRECT and INCORRECT to record the judgment results. The number of CORRECT is counted as the number of correct answers in manual evaluation ( ), and calculate the accuracy rate of manual evaluation: . Among them: is the total number of answers, that is, the total number of query results.

[0057] (2)LLM Evaluation Considering the efficiency of evaluation, the embodiments of the present invention introduce another evaluation method, i.e., LLM evaluation. In some embodiments, to prevent the phenomenon of "self-consistency fallacy", the LLM evaluation model is an LLM other than the optimal base LLM. For example, Llama 3.2 can be used as the LLM evaluation model. In the evaluation process, the LLM evaluation model is set as the role of a medical expert. By providing the LLM evaluation model with problem samples (or questions), reference answers, and answers, the correctness of the output results of the large language model application for assisted diagnosis and treatment based on the RAG technology is determined in batches, and two results, "CORRECT (correct)" or "INCORRECT (incorrect)", are returned. High-efficiency and accurate correctness evaluation is achieved through AI assistance.

[0058] As an example, considering the rigor of the definition and interpretation of medical knowledge, the adjudication criteria are accurate and comprehensive. Therefore, the prompt settings of the LLM evaluation model are as follows: You are a medical expert in urology, specializing in neurogenic bladder, and are now judging whether the answer given by the student is correct. The question is: {query} The correct answer is: {answer} The student's answer is: {result} You will determine whether the student's answer is correct based on the correct answer, and only use two results, CORRECT and INCORRECT, to judge. An incomplete answer is also judged as INCORRECT.

[0059] By counting the number of CORRECT and INCORRECT results, the correctness of the application output answer is measured. The more CORRECT results, the higher the correctness. Counting the number of CORRECT results in the output of the adjudication large language model is the number of correct answers in the LLM evaluation ( )

[0060] The calculation formula for the LLM evaluation - correct rate ( ) is as follows:

[0061] Among them, is the total number of answers.

[0062] In the embodiments of the present invention, the effectiveness of the LLM evaluation results is evaluated using the manual evaluation method as the gold standard, and the evaluation is carried out using precision, recall, and F1 value. Taking the manual evaluation results as the gold standard, the data of the LLM evaluation method are statistically analyzed. The statistical results are shown in Table 5.

[0063] Table 5 Evaluation of the effectiveness of the LLM evaluation method compared with the gold standard

[0064] Compared with traditional manual evaluation methods, the LLM evaluation method performs excellently in terms of precision, while its recall rate is average, and the comprehensive index F1 value shows an above-average level. This indicates that the LLM adopts a more stringent standard when judging the correctness of the output results, and its adjudication results have high credibility and accuracy. In practical applications, this method can be used as an effective auxiliary evaluation method. Especially when the scale of the validation dataset is large, a strategy of combining sampling manual evaluation with batch LLM evaluation can be adopted to improve the evaluation efficiency. In addition, this method has good generality and scalability and is expected to be extended to applications in various specific disease fields.

[0065] Based on this, in the embodiment of the present invention, the basic ability evaluation layer 103 is specifically configured to: first, create an evaluation dataset based on the original database; then, evaluate the application of the specific disease assisted diagnosis and treatment large language model constructed by different hyperparameter combinations based on the evaluation dataset and a preset evaluation system, and obtain correctness quantification indexes such as the correct rate of manual evaluation and the correct rate of LLM evaluation.

[0066] In specific implementation, the evaluation dataset is a standard question-and-answer system covering four core dimensions. The evaluation dataset includes multiple groups of evaluation data for neurogenic bladder disease in the urology field. Each group of evaluation data includes clinical questions for neurogenic bladder disease in the urology field, reference answers, and answers output by the application of the specific disease assisted diagnosis and treatment large language model; among them, the clinical questions for neurogenic bladder disease in the urology field include at least the following four fields: definition and interpretation of medical knowledge, symptom recognition and diagnosis, design and selection of treatment plans, and evidence support and standard citation based on medical literature.

[0067] Considering the application requirements of the model in the embodiments of the present invention mainly for medical students and junior clinicians, reasonable weight distribution is carried out on the number of questions and answers in each dimension of the evaluation dataset in the implementation of the present invention to ensure that both the key fields of medical knowledge can be comprehensively covered and the core content that medical students and junior doctors most need to master in the initial stage of their careers can be effectively focused on. The specific distribution is shown in Figure 4 as shown.

[0068] The character encoding of the evaluation dataset is UTF-8 and is stored in a CSV format file. Data examples are shown in Table 6.

[0069] Table 6 Data Examples of the Evaluation Dataset

[0070] Further, the application of the large language model for auxiliary diagnosis and treatment of specific diseases based on the RAG technology is evaluated according to the evaluation data set, and a quantitative evaluation result of the correct answer rate of the application in the field of neurogenic bladder is obtained. Among them, the preset evaluation system includes: The clinical questions, reference answers, and the output results of the application of the large language model for auxiliary diagnosis and treatment of specific diseases for neurogenic bladder diseases in the field of urology are input into the LLM judgment model to obtain the first accuracy evaluation index value. Specifically, the first accuracy evaluation index value can be the LLM judgment - answer correct rate.

[0071] Next, obtain the second accuracy evaluation index value statistically by users. Specifically, the second accuracy evaluation index value can be the manual judgment - correct rate, and the specific calculation method can be referred to the foregoing embodiments.

[0072] Finally, determine the basic evaluation result of the application of the large language model for auxiliary diagnosis and treatment of specific diseases based on the first accuracy evaluation index value and the second accuracy evaluation index value.

[0073] Specifically, the first accuracy evaluation index value and the second accuracy evaluation index value can be weighted and calculated to obtain the basic evaluation result of the application of the large language model for auxiliary diagnosis and treatment of specific diseases. Or one of the first accuracy evaluation index value or the second accuracy evaluation index value can be selected as the basic evaluation result of the application of the large language model for auxiliary diagnosis and treatment of specific diseases. The specific selection can be made according to the actual situation.

[0074] The evidence chain evaluation layer 104 is used to evaluate the integrity and timeliness of the evidence chain of the output results of the application of the large language model for auxiliary diagnosis and treatment of specific diseases in the basic ability evaluation layer based on the evaluation data set. Specifically, the evidence chain validity and timeliness are evaluated for the output results batch - generated by the basic ability evaluation layer based on the evaluation data set; among them, the evidence chain validity and timeliness evaluation include: through the manual judgment method, each original paragraph referred to by each output result is judged one by one to determine whether the original paragraph is effective in correctly answering the clinical question, and calculate the proportion of valid evidence; according to the publication year of the document to which each original paragraph referred to by each output result belongs and the current year, calculate the evidence timeliness index and the valid evidence timeliness index.

[0075] In some embodiments, the evidence chain evaluation layer aims to measure the validity and timeliness of the evidence chain cited by the application of the large language model for auxiliary diagnosis and treatment of specific diseases when outputting results through the following scientific methods: (1) Proportion of valid evidence (manual evaluation) In the basic ability evaluation layer, for the application of the large language model for auxiliary diagnosis and treatment of specific diseases based on the evaluation dataset, the original paragraphs it references are output along with the output results. Considering the rigor in the medical field, whether the reference original paragraphs are relevant to the input question needs to be considered by means of manual evaluation by clinical professional doctors. Count the proportion of the number ( ) judged as "relevant" by clinical professional doctors in the number (K) of original paragraphs returned, that is = .

[0076] (2) Proportion of valid evidence (similarity evaluation) In the basic ability evaluation layer, for the application of the large language model for auxiliary diagnosis and treatment of specific diseases based on the evaluation dataset, the original paragraphs it references are output along with the output results. Considering the large size of the dataset and the efficiency of batch evaluation. Whether the evidence is relevant to the input question can also be evaluated by calculating the similarity. First, both the input question and each returned reference original paragraph are vectorized. The input question vector is A, and the reference original paragraph vector is , i = [1, k]. The cosine similarity calculation method is as follows: cosine_similarity(A, )=

[0077] Determine the threshold. If cosine_similarity(A, )>0.5, then it is judged that the reference original paragraph is relevant to the input question. Count the proportion of the number ( ) judged as "relevant" by similarity in the number (K) of original paragraphs returned, that is = .

[0078] (3) Evidence timeliness index In the basic ability evaluation layer, for the application of the large language model for auxiliary diagnosis and treatment of specific diseases based on the evaluation dataset, the metadata of the original paragraphs it references are output along with the output results. The metadata includes the publication year Y of the journal / book where the original paragraph is located. Combine with the current year , and calculate the evidence timeliness index , where Δt= - Y, α is a domain-specific attenuation coefficient. For example, for medical papers α = 0.3 (the value drops to 1 / 4 every 3 years). Calculate the average of the evidence timeliness indices of the k reference original paragraphs referenced by the same output result, which is the evidence timeliness index of this output result.

[0079] (4) Effective evidence timeliness index The calculation method of the effective evidence timeliness index is the same as that of the evidence timeliness index. The difference is that only the average value of the evidence timeliness indexes of the n reference original paragraphs marked as "relevant" for the same output result is calculated, which is the effective evidence timeliness index of the output result.

[0080] The proportion of valid evidence index directly reflects the performance of two core links of information retrieval (Retrieval) and semantic alignment (Semantic Alignment) of the large language model for auxiliary diagnosis and treatment of a specific disease. At the same time, combined with the correctness of the output result, further analysis can be carried out: if the proportion of valid evidence is high but the output result is wrong, it indicates that there are problems with the generation logic or the understanding deviation of the LLM; if the proportion of valid evidence is low but the output result is correct, it indicates that the context validity is low, that is, there is redundancy in the reference paragraphs, and the k value can be appropriately reduced to improve the retrieval efficiency.

[0081] The two indexes of evidence timeliness index and effective evidence timeliness index have slightly different focuses when observed. It is recommended to observe the distribution performance of the evidence timeliness index on the entire data set to reflect the timeliness of the current domain database to judge whether the data set needs to be updated. The effective evidence timeliness index focuses on reflecting the timeliness of the output answer in a single instance of question and answer. It is recommended to use the two indexes in combination.

[0082] The compliance evaluation layer 105 is used to construct a privacy evaluation data set and an ethics evaluation data set based on the knowledge domain of a specific disease, and evaluate the privacy processing and ethics processing capabilities of the large language model for auxiliary diagnosis and treatment of a specific disease based on the privacy evaluation data set and the ethics evaluation data set. The privacy evaluation data set includes: a basic privacy evaluation data set and an advanced privacy evaluation data set.

[0083] In some embodiments, the compliance evaluation layer 105 uses the following scientific methods for evaluation: (1) The compliance evaluation layer 105 constructs a basic privacy evaluation data set by randomly inserting simulated user privacy data into the questions of the evaluation data set based on various forms of evaluation data sets for the knowledge domain of a specific disease constructed by the basic ability evaluation layer; among them, the user privacy data includes, but is not limited to: name, ID number, medical treatment ID number.

[0084] In specific implementation, the basic privacy evaluation data set is generated from various forms of evaluation data sets for the knowledge domain of a specific disease constructed by the basic ability evaluation layer. The generation method is to randomly insert simulated user privacy data (such as data related to individuals such as name, ID number, medical treatment ID number, etc.) into the questions on the basis of the evaluation data set of the basic ability evaluation layer.

[0085] For the basic privacy assessment dataset, use the large language model application for assisted diagnosis and treatment of specific diseases to be evaluated to batch output results. Through manual detection and / or LLM detection, check one by one whether there is user privacy information in the output results, and count the number and proportion of the output results with user privacy information.

[0086] In specific implementation, for this basic privacy assessment dataset, use the large language model application for assisted diagnosis and treatment of specific diseases to be evaluated to batch output results. Use manual detection / LLM detection method to check one by one whether there is user privacy information in the output results, and count the number of output results with user privacy information. And proportion . Among them, the calculation formula for the proportion of user privacy output: = , where is the total number of questions. This method evaluates the effectiveness of the pre - privacy filtering mechanism of the large language model application for assisted diagnosis and treatment. The higher the proportion of user privacy output, the worse the privacy filtering performance of the model.

[0087] (2) Compliance assessment layer 105 separately generates an induced question list for specific diseases and medical scenarios, and constructs an advanced privacy assessment dataset; among them, the induced question list includes at least the following four aspects: obtaining user data, processing boundaries of sensitive data, simulating attack scenarios, and inducing the model to expose internal information.

[0088] In specific implementation, the advanced privacy assessment dataset needs to separately generate an induced question list for specific diseases and medical scenarios, including four aspects: obtaining user data, processing boundaries of sensitive data, simulating attack scenarios, and inducing the model to expose internal information.

[0089] For the advanced privacy assessment dataset, use the large language model application for assisted diagnosis and treatment of specific diseases to be evaluated to batch output results. Through manual detection, check one by one whether there is a situation of privacy leakage and blurred data boundaries in the output results, and count the number and proportion of the output results with user privacy information.

[0090] In specific implementation, for this advanced privacy assessment dataset, use the large language model application for assisted diagnosis and treatment of specific diseases to be evaluated to batch output results. Use manual detection to check one by one whether there are situations such as privacy leakage and blurred data boundaries in the output results, and count the number and proportion of the output results with user privacy information. This method evaluates whether the large language model application for assisted diagnosis and treatment of specific diseases has built - in privacy control and access permission control strategies for historical conversations, etc.

[0091] (3) The compliance evaluation layer 105 obtains ethical event cases in the field of specific disease types and constructs an ethical evaluation dataset based on the ethical event cases; among them, the data in the ethical evaluation dataset is in the form of question + reference answer, and the questions in the ethical evaluation dataset include: positive paper dilemma questions or leading questions.

[0092] In specific implementation, ethical evaluation needs to collect the ethical event cases in the field of specific disease types and construct an ethical evaluation dataset, and the ethical evaluation dataset is in the form of question + reference answer. The questions in the ethical evaluation dataset can be positive ethical dilemma questions or leading questions.

[0093] For the ethical evaluation dataset, use the large language model application for assisted diagnosis and treatment of specific disease types to be evaluated to batch output results. By mainly using manual detection and supplemented by LLM detection, check one by one whether the output results are correct compared with the reference answers, and count the number and proportion of correct output results.

[0094] In specific implementation, for this compliance evaluation dataset, use the large language model application for assisted diagnosis and treatment of specific disease types to be evaluated to batch output results. By mainly using manual detection and supplemented by LLM detection, check one by one whether the output results are correct compared with the reference answers. Count the number and proportion of correct output results. This method evaluates the comprehensiveness and coherence of the knowledge of the large language model application for assisted diagnosis and treatment of specific disease types at the levels of morality, ethics, law, etc.

[0095] In some embodiments, the specific disease type is neurogenic bladder disease in the field of urology. An example of the privacy evaluation dataset is shown in Table 7, and an example of the ethical evaluation dataset is shown in Table 8. The evaluation method is mainly manual evaluation and supplemented by LLM evaluation.

[0096] Table 7 Example of the model privacy evaluation dataset in the field of neurogenic bladder

[0097] Table 8 Example of the model ethical evaluation dataset in the field of neurogenic bladder

[0098] In some embodiments, the evaluation system of the large language model application for assisted diagnosis and treatment of specific disease types provided by the embodiments of the present invention further includes a front-end page and a back-end interface.

[0099] The front-end page includes: an import page, an evaluation dataset list page, and a batch evaluation result display page.

[0100] The import page is for users to import evaluation datasets of different categories and different evaluation themes and manage each evaluation dataset.

[0101] The evaluation dataset list page is used to display each detail in the dataset, including the input and the reference answer.

[0102] The batch evaluation result display page shows each detail in the dataset, including the input, the reference answer, the output result of the large language model application for auxiliary diagnosis and treatment of specific diseases, the time consumption, the completion status, the completion time, the manual evaluation result (editable), the LLM evaluation result, the evidence timeliness index, the effective evidence timeliness index, etc. Click on each row to expand and view the original paragraphs cited by the large language model application for auxiliary diagnosis and treatment of specific diseases when outputting the result, display the metadata of each original paragraph (the publication year, author, book / journal name information of the document to which it belongs), and display the similarity between each original paragraph and the input question. It is also possible to manually determine whether each original paragraph is relevant and display the determination result.

[0103] The backend interfaces include the access interface for the large language model application for auxiliary diagnosis and treatment of specific diseases to be evaluated and the evaluation large model interface.

[0104] In specific implementation, the evaluation system for the large language model application for auxiliary diagnosis and treatment of specific diseases is used for: Users import and manage evaluation datasets of various evaluation categories and various topics; Users view the detailed content of the evaluation datasets of various evaluation categories and various topics.

[0105] Users conduct batch evaluation, evaluation result query, evaluation result annotation, statistical index calculation and display for each evaluation dataset.

[0106] The present invention provides the above-mentioned evaluation system for the large language model application for auxiliary diagnosis and treatment of specific diseases. Through the RAG retrieval strategy of "semantic induction-metadata-two-level retrieval" innovatively proposed in the model layer, information is accurately compressed, retrieval noise is reduced; semantic is explicitly induced to enhance the interpretability of model question answering; the two-level retrieval mechanism improves retrieval accuracy and efficiency, taking into account semantic matching and detail accuracy; the extraction and preservation of metadata provide the possibility for the application to achieve multi-dimensional retrieval, and also provide the feasibility for the subsequent traceability and transparency of the evidence chain, and the filtering is especially applicable to scenarios such as logical long documents and multi-topic knowledge bases in the medical field.

[0107] Through the base model evaluation layer, the adaptability of the large language model application for auxiliary diagnosis and treatment of specific diseases to the base LLM model and the RAG architecture can be clarified, and the coverage and performance efficiency of the base LLM in domain knowledge can be quantitatively evaluated. Through quantitative analysis, business requirements are transformed into implementable engineering standards, avoiding resource waste or functional defects.

[0108] Through the basic evaluation layer, the professional capabilities of large language models in the auxiliary diagnosis and treatment of specific diseases can be systematically evaluated, focusing on the core links of clinical decision-making for junior doctors. It includes a four-dimensional domain-specific evaluation data framework, including the definition and interpretation of medical knowledge, symptom recognition and diagnosis, the design and selection of treatment plans, and medical literature retrieval. A dual-modal complementary mechanism is adopted, that is, manual evaluation: clinical experts make accurate judgments based on their own experience and reference answers; LLM evaluation: through an automated evaluation framework, the adjudication LLM model is used to batch and quickly screen the outputs of candidate models. The precise delineation of the professional ability boundaries of LLM in the entire process of clinical diagnosis and treatment is achieved, providing a quantifiable decision for the model optimization direction.

[0109] Through the evidence chain evaluation layer, in-depth analysis and traceability work are carried out on the original text paragraphs cited in the model output results to improve the transparency and credibility of clinical decisions. Specifically, for the key information involved in the model output content, by constructing an explicit semantic evidence chain mechanism, multi-level tracing and precise positioning of the cited evidence are realized. This mechanism traces back layer by layer to the specific paragraphs in the original literature or knowledge base, forming a clear and verifiable evidence context. In this process, the relevance of the traced evidence to the current problem domain is quantified by means of manual judgment or automatic calculation of similarity. At the same time, to ensure the timeliness and reliability of medical knowledge, this study further introduces an evidence timeliness evaluation module. This module quantitatively scores the timeliness of the traced evidence based on the publication time of the literature, providing a double guarantee of timeliness and scientificity for clinical decisions.

[0110] Through the compliance evaluation layer, it is evaluated whether the large language model for auxiliary diagnosis and treatment has a pre-processing mechanism for privacy processing and ethical issues, and the effectiveness of the mechanism is judged. Ensure that the model adheres to the principles of fairness and justice in medical decision-making, and avoid the imbalance of resource allocation and diagnostic discrimination caused by algorithmic bias. At the same time, a strict data protection network is constructed to strictly prevent the leakage of patients' sensitive information, safeguard their information autonomy and privacy rights, and provide a solid guarantee for the safe, compliant and humanized application of medical large models.

[0111] In a feasible implementation, the system can also be monitored in real time to obtain its error information, as shown in Table 9.

[0112] Table 9 Example of System Error Information

[0113] Furthermore, corresponding disaster recovery measures are taken for the monitored error information, including: enabling automatic switching of compatibility models; falling back to conservative tools or manual processes when tool calls fail; setting up multi-node proxies and retry queues at the API layer; realizing fault migration within seconds through containerized orchestration; deploying multi-region DNS load balancing and other strategies.

[0114] In the continuous application of the evaluation system for the Disease-Specific Clinical Large Language Model (DS-CLLM), the present invention proposes a multi-dimensional dynamic evaluation paradigm: (1) Horizontal benchmarking evaluation: By constructing a multi-model benchmark test set (covering ≥5 DS-CLLMs with different architectures), quantitatively compare the performance differences of each model in disease-specific diagnosis and treatment tasks based on a unified evaluation index set; (2) Vertical temporal evaluation: Based on the domain knowledge base version control mechanism, conduct historical retrospective analysis on the diagnosis and treatment outputs of the same DS-CLLM before and after the knowledge base update cycle, and quantitatively evaluate the model knowledge update efficiency; (3) Trend quantitative modeling: Construct a four-dimensional evaluation vector including accuracy, integrity, timeliness, and compliance, and fit the model development trajectory through time series analysis to reveal the dynamic evolution law of the adaptation between technology iteration and clinical needs.

[0115] The evaluation system for the application of the disease-specific clinical large language model provided by the embodiment of the present invention constructs a multi-dimensional quantitative evaluation paradigm anchored by the clinical diagnosis and treatment scenario. It realizes the full-cycle quantitative evaluation of the disease-specific clinical large language model from basic capabilities to clinical implementation. The evaluation results can provide triple optimization evidence for model iteration: ability defect positioning, accurately identifying the knowledge blind spots and logical defects of the model in specific diseases through the analysis of the basic capabilities and professional capabilities of the base; evidence chain strengthening path, guiding the update strategy of the credible data space and the optimization of diagnosis and treatment logic rules in specific fields based on the joint analysis of evidence traceability coverage and timeliness matching degree; compliance risk assessment, establishing a safety threshold verification mechanism before model clinical deployment through the dynamic monitoring of the ethical rule conflict rate and the effectiveness of privacy data desensitization.

[0116] This evaluation system significantly improves the clinical adaptability and credibility of the medical large language model in the field of specific diseases, provides a quantifiable quality assurance standard for promoting the transformation and application of AI technology from the laboratory to the clinic, helps to build a closed-loop iterative ecosystem of "evaluation - improvement - re-evaluation", and promotes the deep integration of precision medicine and artificial intelligence.

[0117] For the evaluation system for the application of the disease-specific clinical large language model provided in the foregoing embodiment, the embodiment of the present invention also provides an evaluation method for the application of the disease-specific clinical large language model. Refer to Figure 5 the flowchart of an evaluation method for the application of a disease-specific clinical large language model shown, which shows that the method mainly includes the following steps S501 to step S505: Step S501: Construct a RAG retrieval strategy to build or adjust the application of the disease-specific clinical large language model based on the RAG technology based on the RAG retrieval strategy.

[0118] In specific implementation, for a knowledge base of specific diseases, an RAG retrieval strategy of "semantic induction-metadata-two-level retrieval" is innovatively designed: the long document is segmented by paragraphs, and the LLM is used to perform explicit semantic induction on each paragraph, extract paragraph semantic tags, the publication year of the document to which the paragraph belongs, the author, and the name of the book / journal, and construct the mapping relationship between the semantic tags, publication year, author, book / journal name and the original paragraph; then the semantic tags are merged into a document and segmented again, and a semantic vector library is generated through the LLM. When the user queries, the input question is first vectorized and processed, the semantic vector library is queried and multiple most relevant semantic tags are returned, and then the original paragraphs are located according to the mapping relationship. Finally, based on the user input question, the returned original paragraphs and the prompt word template, the LLM outputs the result. At the same time, the original paragraphs referred to by the answer, the publication year of the document to which the original paragraph belongs, the author, and the name of the book / journal are also returned.

[0119] Step S502: Perform multi-dimensional evaluation on the large language model base LLM for auxiliary diagnosis and treatment for specific diseases to evaluate the basic capabilities of the large language model base LLM for auxiliary diagnosis and treatment in the knowledge field of specific diseases.

[0120] In specific implementation, using the base LLM selected for the application of the large language model for auxiliary diagnosis and treatment of specific diseases, based on the pre-constructed medical entity data set for the clinical medical field of the specific disease, perform a medical entity classification task, and calculate the first performance evaluation index of the degree of mastery of the specific disease expertise of the base LLM based on the medical entity classification result; Using the base LLM selected for the application of the large language model for auxiliary diagnosis and treatment of specific diseases, based on the pre-constructed data set of case descriptions related to specific diseases, perform a medical entity extraction task, and calculate the second performance evaluation index of the semantic understanding and logical reasoning ability of the base LLM for specific diseases based on the medical entity extraction result; Using the base LLM selected for the application of the large language model for auxiliary diagnosis and treatment of specific diseases, perform text segmentation, vector library construction and retriever generation tasks on the pre-constructed data set of specific disease descriptions, perform keyword retrieval based on the vector library and the retriever, and calculate the third performance evaluation index of the retriever module performance of the base LLM for specific diseases based on the retrieval result; Based on the first performance evaluation index, the second performance evaluation index and the third performance evaluation index, perform multi-dimensional evaluation on the base LLM selected for each application of the large language model for auxiliary diagnosis and treatment of specific diseases in the knowledge field of specific diseases, and determine the basic capabilities of the base LLM in the knowledge field of specific diseases.

[0121] Step S503: Construct evaluation datasets in various forms for the knowledge domain of specific diseases, and evaluate the correctness of the output results of the application of the large language model for specific disease assisted diagnosis and treatment based on the evaluation datasets.

[0122] In specific implementation, the evaluation datasets are divided into 4 domains for specific diseases, namely: definition and interpretation of medical knowledge, symptom recognition and diagnosis, design and selection of treatment plans, and medical literature retrieval; the evaluation datasets for each domain are in the form of a combination of clinical questions + reference answers for the knowledge domain of specific diseases.

[0123] Considering the application requirements of the model in this embodiment of the invention mainly for medical students and junior clinicians, reasonable weight distribution is carried out on the number of questions and answers in each dimension of the evaluation datasets in the implementation of the invention to ensure that both the key domains of medical knowledge can be comprehensively covered and the core content that medical students and junior doctors most need to master in the initial stage of their careers can be effectively focused on. The character encoding of the evaluation datasets is UTF-8 and is stored in a CSV format file.

[0124] Furthermore, evaluate the application of the large language model for specific disease assisted diagnosis and treatment according to the evaluation datasets to obtain a quantitative evaluation result of the correct answer rate for questions and answers in the field of the specific disease of neurogenic bladder. Among them, the preset evaluation system includes: Input the clinical questions, reference answers and the output results of the application of the large language model for specific disease assisted diagnosis and treatment for neurogenic bladder disease in the urology field into the LLM judgment model to obtain the first accuracy evaluation index value. Specifically, the first accuracy evaluation index value can be the LLM judgment - answer correct rate.

[0125] Then, obtain the second accuracy evaluation index value manually counted by users. Specifically, the second accuracy evaluation index value can be the manual judgment - correct rate, and the specific calculation method can be referred to the foregoing embodiments.

[0126] Finally, determine the basic evaluation result of the application of the large language model for specific disease assisted diagnosis and treatment based on the first accuracy evaluation index value and the second accuracy evaluation index value.

[0127] Specifically, the first accuracy evaluation index value and the second accuracy evaluation index value can be weighted and calculated to obtain the basic evaluation result of the application of the large language model for specific disease assisted diagnosis and treatment. Or one of the first accuracy evaluation index value or the second accuracy evaluation index value can be selected as the basic evaluation result of the application of the large language model for specific disease assisted diagnosis and treatment. It can be specifically selected according to the actual situation.

[0128] Step S504: Evaluate the integrity of the evidence chain and the timeliness of the evidence for the output results of the application of the large language model for specific disease assisted diagnosis and treatment based on the evaluation datasets of the basic ability evaluation layer.

[0129] In some embodiments, the evidence chain evaluation layer can measure the effectiveness and timeliness of the evidence chain cited by the auxiliary diagnosis and treatment large model when outputting results through the proportion of valid evidence, the proportion of valid evidence (similarity evaluation), the evidence timeliness index, and the valid evidence timeliness index. For specific details, please refer to the foregoing embodiments.

[0130] Step S505: Construct a privacy evaluation dataset and an ethics evaluation dataset based on the knowledge domain of a specific disease, and evaluate the privacy processing and ethics handling capabilities of the large language model for auxiliary diagnosis and treatment of the specific disease based on the privacy evaluation dataset and the ethics evaluation dataset.

[0131] In some embodiments, the privacy evaluation dataset includes a basic privacy evaluation dataset and an advanced privacy evaluation dataset. For the basic privacy evaluation dataset, use the large language model application for auxiliary diagnosis and treatment of the specific disease to be evaluated to perform batch output results, and use manual detection / LLM detection methods to check one by one whether there is user privacy information in the output results, and count the number of output results containing user privacy information And the proportion . For the advanced privacy evaluation dataset, use the large language model application for auxiliary diagnosis and treatment of the specific disease to be evaluated to perform batch output results, and use manual detection to check one by one whether there are situations such as privacy leakage and blurred data boundaries in the output results, and count the number and proportion of output results containing user privacy information. For the compliance evaluation dataset, use the large language model application for auxiliary diagnosis and treatment of the specific disease to be evaluated to perform batch output results, and mainly use manual detection and supplemented by LLM detection to check one by one whether the output results are correct compared with the reference answers. Count the number and proportion of correct output results.

[0132] It should be noted that the method provided in the embodiments of the present invention has the same implementation principle and technical effects as the foregoing system embodiments. For the sake of brief description, for the parts not mentioned in the method embodiments, reference may be made to the corresponding content in the foregoing system embodiments. The specific values provided in the embodiments of the present invention are only exemplary and are not limited herein.

[0133] The embodiments of the present invention also provide an electronic device. Specifically, the electronic device includes a processor and a storage device; a computer program is stored on the storage device, and when the computer program is run by the processor, it executes the method of any one of the above embodiments.

[0134] Figure 6A schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 includes: a processor 60, a memory 61, a bus 62, and a communication interface 63. The processor 60, the communication interface 63, and the memory 61 are connected through the bus 62. The processor 60 is configured to execute an executable module stored in the memory 61, such as a computer program.

[0135] Among them, the memory 61 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 63 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0136] The bus 62 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0137] Among them, the memory 61 is used to store a program. After receiving an execution instruction, the processor 60 executes the program. The method executed by the device defined by the flow process disclosed in any one of the foregoing embodiments of the present invention can be applied to the processor 60 or implemented by the processor 60.

[0138] The processor 60 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 60 or the instructions in the form of software. The above-mentioned processor 60 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 61, and the processor 60 reads the information in the memory 61 and combines its hardware to complete the steps of the above method.

[0139] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For the specific implementation, reference can be made to the foregoing method embodiments and will not be elaborated here.

[0140] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0141] Finally, it should be noted that the above-mentioned embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An evaluation system for the application of a large language model for assisting in the diagnosis and treatment of specific diseases, characterized in that, Including: A model layer, which is used to build a RAG retrieval strategy to build or adjust a large language model application for assisting in the diagnosis and treatment of specific diseases based on the RAG retrieval strategy; A base model evaluation layer, which is used to perform multi-dimensional evaluations of the large language model base LLM for assisting in the diagnosis and treatment of diseases for specific diseases, so as to evaluate the basic capabilities of the large language model base LLM for assisting in the diagnosis and treatment of diseases in the knowledge field of the specific diseases; A basic capability evaluation layer, which is used to build various forms of evaluation data sets for the knowledge field of specific diseases, and based on the evaluation data sets, evaluate the correctness of the output results of the large language model application for assisting in the diagnosis and treatment of specific diseases; among them, the evaluation data sets are divided into 4 fields for the specific diseases, namely: definition and interpretation of medical knowledge, symptom recognition and diagnosis, design and selection of treatment plans, and medical literature retrieval; the evaluation data sets for each field are in the form of a combination of clinical questions + reference answers for the knowledge field of the specific diseases; An evidence chain evaluation layer, which is used to evaluate the integrity and timeliness of the evidence chain of the output results of the large language model application for assisting in the diagnosis and treatment of specific diseases based on the evaluation data sets of the basic capability evaluation layer; A compliance evaluation layer, which is used to build a privacy evaluation data set and an ethical evaluation data set based on the knowledge field of the specific diseases, and based on the privacy evaluation data set and the ethical evaluation data set, evaluate the capabilities of the large language model application for assisting in the diagnosis and treatment of specific diseases in privacy processing and ethical issue processing.

2. The system according to claim 1, wherein Specifically, the model layer is used for: Based on the knowledge base of specific diseases, build a RAG retrieval strategy based on semantic induction-metadata-two-level retrieval; among them, the RAG retrieval strategy includes: segmenting long documents by paragraphs, and using the LLM to perform explicit semantic induction on each paragraph, and extracting paragraph semantic labels, the publication year of the document to which the paragraph belongs, the author, and the name of the book / journal, and building a mapping relationship between the semantic labels, publication year, author, book / journal name and the original paragraph; Merge the semantic labels into a semantic label document, and after segmenting the semantic label document, generate a semantic vector library through the LLM; When a user queries, convert the question input by the user into a vector to obtain a question vector, query in the semantic vector library based on the question vector, and return multiple semantic labels related to the question vector; Locate the original paragraphs corresponding to each semantic label based on the mapping relationship between the semantic label and the original paragraph, and based on the question input by the user, the returned original paragraph and the prompt word template, output the result by the LLM; among them, the output result of the LLM includes: each original paragraph and original paragraph metadata referred to in the answer, and the original paragraph metadata at least includes the publication year of the document to which the original paragraph belongs, the author, and the name of the book / journal.

3. The system according to claim 1, wherein Specifically, the base model evaluation layer is used for: Using the selected base LLM of the large language model application for auxiliary diagnosis and treatment of specific diseases, based on the pre-constructed medical entity dataset for the clinical medical field of the specific diseases, perform a medical entity classification task, and calculate the first performance evaluation index of the professional knowledge mastery degree of the base LLM for the specific diseases based on the medical entity classification results; Using the selected base LLM of the large language model application for auxiliary diagnosis and treatment of specific diseases, based on the pre-constructed case description dataset related to the specific diseases, perform a medical entity extraction task, and calculate the second performance evaluation index of the semantic understanding and logical reasoning ability of the base LLM for the specific diseases based on the medical entity extraction results; Using the selected base LLM of the large language model application for auxiliary diagnosis and treatment of specific diseases, perform text segmentation, vector library construction and retriever generation tasks on the pre-constructed specific disease description dataset, perform keyword retrieval based on the vector library and the retriever, and calculate the third performance evaluation index of the retrieval module performance of the base LLM for the specific diseases based on the retrieval results; Based on the first performance evaluation index, the second performance evaluation index and the third performance evaluation index, conduct a multi-dimensional evaluation of the selected base LLM of the large language model application for auxiliary diagnosis and treatment of specific diseases in the knowledge field of the specific diseases, and determine the basic capabilities of the base LLM in the knowledge field of the specific diseases.

4. The system according to claim 1, wherein The basic capability evaluation layer is specifically used for: Evaluating the correctness of the answers batch-output by the auxiliary diagnosis and treatment large language model application based on the RAG technology based on the evaluation dataset through manual evaluation and / or an LLM evaluation model; wherein, the LLM evaluation model is a different LLM from the base LLM; the LLM evaluation model is set as a medical expert role for the specific diseases, and the output of the manual evaluation and / or the LLM evaluation model is correct or incorrect.

5. The system according to claim 2, wherein The evidence chain evaluation layer is specifically used for: Conducting evidence chain validity evaluation and evidence timeliness evaluation on the output results batch-generated by the basic capability evaluation layer based on the evaluation dataset; wherein, the evidence chain validity evaluation and evidence timeliness evaluation include: through a manual evaluation method, judging each original paragraph referred to by each output result one by one, judging whether the original paragraph is effective for correctly answering clinical questions, and calculating the proportion of valid evidence; according to the publication year of the documents to which each original paragraph referred to by each output result belongs and the current year, calculating the evidence timeliness index and the valid evidence timeliness index.

6. The system according to claim 1, wherein The privacy evaluation dataset includes: a basic privacy evaluation dataset and an advanced privacy evaluation dataset; The compliance evaluation layer is specifically used for: based on the various forms of evaluation datasets constructed by the basic capability evaluation layer for the knowledge field of specific diseases, randomly inserting simulated user privacy data into the questions of the evaluation dataset to construct the basic privacy evaluation dataset; wherein, the user privacy data at least includes: name, ID number, medical treatment ID number; Generate an inductive question list specifically for the specific disease type and medical scenario, and construct the advanced privacy assessment dataset; wherein, the inductive question list includes at least the following four aspects: obtaining user data, the processing boundary of sensitive data, simulating attack scenarios, and inducing the model to expose internal information; Obtain ethical event cases in the field of the specific disease type, and construct the ethical assessment dataset based on the ethical event cases; wherein, the data in the ethical assessment dataset is in the form of question + reference answer, and the questions in the ethical assessment dataset include: positive ethical dilemma questions or inductive questions.

7. The system according to claim 6, wherein The compliance assessment layer is specifically used for: For the basic privacy assessment dataset, use the large language model application for auxiliary diagnosis and treatment of the specific disease type to be evaluated to batch output results, and check whether there is user privacy information in the output results one by one through manual detection and / or LLM detection, and count the number and proportion of output results with user privacy information; For the advanced privacy assessment dataset, use the large language model application for auxiliary diagnosis and treatment of the specific disease type to be evaluated to batch output results, and check whether there is privacy leakage or blurred data boundary in the output results one by one through manual detection, and count the number and proportion of output results with user privacy information; For the ethical assessment dataset, use the large language model application for auxiliary diagnosis and treatment of the specific disease type to be evaluated to batch output results, and check whether the output results are correct compared with the reference answers one by one mainly through manual detection and supplemented by LLM detection, and count the number and proportion of correct output results; 8. The system according to claim 1, wherein The evaluation system of the large language model application for auxiliary diagnosis and treatment of the specific disease type also includes a front-end page and a back-end interface; The front-end page includes: an import page, an evaluation dataset list page, and a batch evaluation result display page; The import page is used for users to import evaluation datasets of different categories and different evaluation topics, and manage each evaluation dataset; The evaluation dataset list page is used to display each detail in the evaluation dataset, including input and reference answer; The batch evaluation result display page displays each detail in the evaluation result dataset, including input, reference answer, output result of the large language model application for auxiliary diagnosis and treatment of the specific disease type, time consumption, completion status, completion time, manual evaluation result, LLM evaluation result, evidence timeliness index, effective evidence timeliness index; wherein, by clicking on each row of data to expand, view each original paragraph cited by the large language model application for auxiliary diagnosis and treatment of the specific disease type when outputting the result, the metadata of each original paragraph, the similarity between each original paragraph and the input question, and the discrimination result of whether each original paragraph is relevant; The back-end interface includes an access interface for the large language model application for auxiliary diagnosis and treatment of the specific disease type to be evaluated and an evaluation large model interface.

9. An evaluation method for the application of a large language model for assisting diagnosis and treatment of specific diseases, characterized in that, An evaluation system applied to the large language model application for auxiliary diagnosis and treatment of the specific disease type according to any one of claims 1 to 8, includes: Construct a RAG retrieval strategy to build or adjust a large language model application for assisting in the diagnosis and treatment of specific diseases based on the RAG retrieval strategy; Conduct multi-dimensional evaluations of the large language model base LLM for specific diseases to evaluate the basic capabilities of the large language model base LLM in the knowledge domain of the specific diseases; Construct evaluation data sets in various forms for the knowledge domain of specific diseases, and evaluate the correctness of the output results of the large language model application for assisting in the diagnosis and treatment of specific diseases based on the evaluation data sets; among them, the evaluation data sets are divided into 4 fields for the specific diseases, namely: definition and interpretation of medical knowledge, symptom recognition and diagnosis, design and selection of treatment plans, and medical literature retrieval; the evaluation data sets for each field are in the form of a combination of clinical questions + reference answers for the knowledge domain of the specific diseases; Evaluate the integrity of the evidence chain and the timeliness of the evidence of the output results of the large language model application for assisting in the diagnosis and treatment of specific diseases based on the evaluation data sets of the basic capability evaluation layer; Construct a privacy evaluation data set and an ethical evaluation data set based on the knowledge domain of the specific diseases, and evaluate the privacy processing and ethical issue handling capabilities of the large language model application for assisting in the diagnosis and treatment of specific diseases based on the privacy evaluation data set and the ethical evaluation data set.

10. An electronic device, characterized in that, It includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the method described in claim 9.

Citation Information

Patent Citations

  • Medical model evaluation method and device, electronic equipment and storage medium

    CN117407682A

  • Medical auxiliary question and answer method and system based on knowledge calibration and retrieval enhancement

    CN117573843A

  • Breast cancer clinical data analysis diagnosis and treatment platform based on artificial intelligence large language model

    CN118016280A

  • Acupuncture auxiliary diagnosis and treatment and model evaluation method based on large language model

    CN118098506A

  • Large language model selection method based on multi-dimensional evaluation and dynamic weight adjustment

    CN119226448A

Cited By

  • Evidence pollution chain traceability system and method based on data analysis

    CN120613058A

  • Double-domain RAG-driven multi-omics fusion pathology analysis system

    CN121191788A

  • Dual-domain rag-driven multi-omics fusion pathological analysis system

    CN121191788B

  • Medical auxiliary question answering method and device, electronic equipment and storage medium

    CN121233734A