Method, device and equipment for evaluating medical information research capability of large language model
Patent Information
- Application Number
- CN202610730654.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-05-26
AI Technical Summary
[0005]本发明提供一种对大语言模型医疗信息研究能力的评估方法、装置和设备,用以解决现有技术中无法对大语言模型在医疗信息研究能力方面进行准确评估的缺陷,实现对大语言模型在医疗信息研究能力方面进行准确评估的目的
[0018]This invention provides a method, apparatus, and device for evaluating the medical information research capabilities of large language models. The method involves inputting medical research query information into multiple large language models to be tested, obtaining papers to be evaluated from the outputs of each model. For each paper, the method inputs it into a multimodal evaluation model to obtain a quality score. This evaluation model is trained on an initial evaluation model based on positive and negative paper samples and corresponding quality score labels. The quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions. Based on the quality scores of each paper, the medical information research capabilities of the multiple large language models are evaluated, yielding the evaluation results. Because standardized, unified input is used to ensure fairness at the starting point of the evaluation, and a unified, objective multimodal evaluation model trained with domain expertise is used as the evaluation standard, standardized, objective, and quantitative evaluation of papers generated by different large language models can be achieved, thus enabling accurate evaluation of the medical information research capabilities of large language models.
Smart Images

Figure CN122266820B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, and equipment for evaluating the research capabilities of large language models in medical information. Background Technology
[0002] With the rapid development of artificial intelligence technology, Large Language Models (LLMs) have demonstrated enormous application potential in the interdisciplinary field of medical informatics research. Medical informatics research deeply integrates biomedical, data science, and clinical practice knowledge. Its research data patterns are complex and diverse, encompassing multimodal information such as text and images, and the domain knowledge is updated extremely rapidly. Utilizing large language models to drive research agents holds promise for automating or assisting the end-to-end research process, from literature review and data mining to generating complete research papers, thereby significantly improving the efficiency of medical research.
[0003] Currently, research and practice in medical informatics primarily rely on existing generative systems. These systems typically use large language models as their core engine to generate research papers, including text descriptions, data analysis results, and graphical representations, based on user-input research topics or questions.
[0004] However, due to significant differences in the mastery of medical professional knowledge, logical reasoning ability, and the ability to understand and generate multimodal information among different large language models, the quality of papers generated by different large language models varies. Therefore, how to standardize and objectively evaluate the papers generated by different large language models in order to accurately assess the ability of large language models in medical information research is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This invention provides a method, apparatus, and equipment for evaluating the medical information research capabilities of large language models, thereby addressing the shortcomings of existing technologies that cannot accurately evaluate the medical information research capabilities of large language models, and achieving the goal of accurately evaluating the medical information research capabilities of large language models.
[0006] This invention provides a method for evaluating the medical information research capabilities of large language models, comprising: Medical research query information is input into multiple large language models to be tested, and the papers to be evaluated are output by each of the large language models to be tested. For each of the aforementioned papers to be evaluated, the paper to be evaluated is input into a multimodal evaluation model to obtain the quality score of the paper to be evaluated output by the multimodal evaluation model. The multimodal evaluation model is obtained by training an initial evaluation model based on positive paper samples, negative paper samples and corresponding quality score labels. The quality score labels are determined after performing fine-grained quality evaluation on the corresponding paper samples from multiple dimensions. Based on the quality scores of each of the papers to be evaluated, the medical information research capabilities of the multiple large language models to be tested are assessed, and the evaluation results are obtained.
[0007] According to the present invention, a method for evaluating the medical information research capabilities of large language models includes inputting medical research query information into multiple large language models to be tested, and obtaining the papers to be evaluated output by each of the large language models to be tested, comprising: For each of the large language models to be tested, the medical research query information is input into the supervision submodule of the large language model to be tested. The supervision submodule generates literature retrieval tasks, data analysis tasks, coding tasks, and writing tasks. The literature retrieval task is input into the literature review submodule of the large language model to be tested, the data analysis task is input into the data analysis submodule of the large language model to be tested, the coding task is input into the coding submodule of the large language model to be tested, and the writing task is input into the paper generation submodule of the large language model to be tested. The literature review submodule is used to retrieve literature related to the medical research query information. The data analysis submodule processes and analyzes data related to the medical research query information. Generate visual charts through the coding submodule; The paper generation submodule generates the paper to be evaluated based on the retrieved literature, processed data, and generated visualization charts.
[0008] According to the present invention, a method for evaluating the medical information research capabilities of a large language model is provided, wherein the multimodal evaluation model is trained based on the following method: Obtain multiple positive paper samples and multiple negative paper samples respectively; Fine-grained quality annotations were obtained for each of the positive and negative paper samples across multiple dimensions to obtain quality score labels. Each of the positive paper samples and each of the negative paper samples are input into the visual language model to obtain the global feature vector of the paper output by the visual language model. The visual language model is trained based on medical image and text data. The global feature vector of the paper is input into the initial scoring model to obtain the predicted quality score output by the initial scoring model; Loss information is determined based on the predicted quality score and the quality score label; The initial scoring model is iteratively trained based on the loss information to obtain the scoring model; The visual language model and the scoring model are identified as the multimodal evaluation model.
[0009] According to the present invention, a method for evaluating the research capability of a large language model in medical information is provided, wherein the quality score label includes the total quality score of the paper and the score on each dimension; the predicted quality score includes the predicted total quality score and the predicted score on each of the dimensions. The step of determining loss information based on the predicted quality score and the quality score label includes: Based on the total paper quality score and the predicted total quality score, the principal loss is determined; Based on the scores in each dimension and the predicted scores in each dimension, an auxiliary loss is determined; The loss information is determined based on the main loss and the auxiliary loss.
[0010] According to the method for evaluating the research capability of a large language model in medical information according to the present invention, multiple negative paper samples are obtained, including: The sample medical research query information is input into multiple large language models to obtain the paper samples output by each of the large language models; The paper samples output by each of the large language models are determined as the negative paper samples.
[0011] According to the present invention, an evaluation method for medical information research capabilities of a large language model is provided, wherein the multiple dimensions include at least two of the following: accuracy of medical facts, rigor of argumentation logic, consistency between text and graphics, and professionalism of typesetting.
[0012] According to the present invention, a method for evaluating the research capabilities of large language models in medical information includes: Based on the quality scores of each of the papers to be evaluated, target papers to be evaluated whose quality scores are less than a preset score are identified. The target paper to be evaluated is analyzed using interpretability tools to obtain a diagnostic report, which includes the reasons for the low score of the target paper to be evaluated. Output the diagnostic report.
[0013] According to the present invention, a method for evaluating the research capabilities of large language models in medical information includes: Obtain expert scores for each of the papers to be evaluated; For each of the papers to be evaluated, determine the Spearman rank correlation coefficient between the expert scores and quality scores of the papers to be evaluated; The quality score output by the multimodal evaluation model is verified based on the Spearman rank correlation coefficient.
[0014] This invention also provides an evaluation device for the medical information research capabilities of large language models, comprising: The input module is used to input medical research query information into multiple large language models to be tested, and obtain the papers to be evaluated output by each of the large language models to be tested. The input module is further configured to input the paper to be evaluated into a multimodal evaluation model for each of the papers to be evaluated, and obtain the quality score of the paper to be evaluated output by the multimodal evaluation model. The multimodal evaluation model is obtained by training an initial evaluation model based on positive paper samples, negative paper samples and corresponding quality score labels. The quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions. The evaluation module is used to assess the medical information research capabilities of the multiple large language models under test based on the quality scores of each of the papers to be evaluated, and to obtain the evaluation results.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the evaluation method for the research capability of large language model medical information as described above.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the evaluation method for the research capability of large language model medical information as described above.
[0017] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the evaluation method for the research capability of large language model medical information as described above.
[0018] This invention provides a method, apparatus, and device for evaluating the medical information research capabilities of large language models. The method involves inputting medical research query information into multiple large language models to be tested, obtaining papers to be evaluated from the outputs of each model. For each paper, the method inputs it into a multimodal evaluation model to obtain a quality score. This evaluation model is trained on an initial evaluation model based on positive and negative paper samples and corresponding quality score labels. The quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions. Based on the quality scores of each paper, the medical information research capabilities of the multiple large language models are evaluated, yielding the evaluation results. Because standardized, unified input is used to ensure fairness at the starting point of the evaluation, and a unified, objective multimodal evaluation model trained with domain expertise is used as the evaluation standard, standardized, objective, and quantitative evaluation of papers generated by different large language models can be achieved, thus enabling accurate evaluation of the medical information research capabilities of large language models. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the method for evaluating the research capabilities of large language models in medical information, as provided in an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of an evaluation framework for the medical information research capabilities of large language models provided in an embodiment of the present invention.
[0022] Figure 3 This is a schematic diagram of the structure of the device for evaluating the research capabilities of large language models in medical information, provided in an embodiment of the present invention.
[0023] Figure 4 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0025] Medical informatics research is characterized by diverse data patterns, rapid knowledge updates, and the need to integrate biomedical, data science, and clinical practice knowledge. In recent years, research agents based on large language models have demonstrated significant potential in literature reviews, data analysis, and even end-to-end research execution. Currently, research and practice in this field primarily rely on existing generative systems. These systems typically directly utilize mainstream large language models as their core engines, generating research papers that include text descriptions, data analysis results, and even graphical representations based on user-input research topics or questions.
[0026] Because different large language models exhibit significant differences in their mastery of medical expertise, logical reasoning abilities, and capacity for understanding and generating multimodal information, the quality of papers generated by their driving research agents varies considerably. Current technologies for evaluating these papers typically rely on researchers' subjective experience or attempt to use general-domain natural language processing benchmarks, such as text fluency or content consistency, for rough assessment. There is a lack of specific evaluation processes and standards tailored to medical informatics research. Therefore, how to standardize and objectively evaluate papers generated by different large language models to accurately assess their capabilities in medical informatics research is a pressing technical problem that needs to be solved.
[0027] In view of the above-mentioned problems, this invention proposes an evaluation method for the medical information research capabilities of large language models. This method can achieve standardized and objective evaluation of papers generated by different large language models, thereby enabling an accurate assessment of the medical information research capabilities of large language models.
[0028] The following is combined with Figure 1 and Figure 2This invention describes a method for evaluating the research capabilities of large language models in medical information. This invention is applicable to scenarios requiring performance evaluation and capability comparison of medical research agents based on large language models. The executing entity of this method can be an electronic device such as a terminal device, computer, server, server cluster, or a specially designed evaluation device for the research capabilities of large language models in medical information. It can also be an evaluation device for the research capabilities of large language models in medical information, which can be implemented through software, hardware, or a combination of both.
[0029] Figure 1 This is a flowchart illustrating the method for evaluating the medical information research capabilities of a large language model provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes: Step 101: Input the medical research query information into multiple large language models to be tested, and obtain the papers to be evaluated output by each large language model.
[0030] In this step, the medical research query information is standardized. It can be understood as a well-defined medical research task or question, such as "analyzing the correlation between daily steps and systolic blood pressure levels in a hypertensive patient population." By using standardized medical research query information and inputting different large language models to generate papers for evaluation, potential biases caused by different tasks can be eliminated, ensuring the fairness of the entire evaluation process.
[0031] Step 102: For each paper to be evaluated, input the paper to be evaluated into the multimodal evaluation model to obtain the quality score of the paper to be evaluated output by the multimodal evaluation model. The multimodal evaluation model is obtained by training the initial evaluation model based on positive paper samples, negative paper samples and corresponding quality score labels. The quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions.
[0032] In this step, the multimodal evaluation model is a model specifically built for evaluating scientific papers. It has the ability to simultaneously understand and analyze the textual and visual content of the paper. The textual content may include, for example, research background, methodological descriptions, and results analysis, while the visual content may include, for example, statistical charts, flowcharts, and medical image illustrations.
[0033] When training a multimodal evaluation model, the training data includes positive and negative paper samples. Positive paper samples typically come from high-quality papers selected from medical journals or public health databases (PubMed Central, PMC), representing excellent examples. Negative paper samples usually contain various types of defects, such as factual errors, logical inconsistencies, or discrepancies between text and images. These negative paper samples can be papers on the same research topic generated by multiple large language models to be evaluated.
[0034] In addition, the training data includes quality score labels for each paper sample. These quality score labels are determined by domain experts who evaluate each paper sample based on multiple pre-defined fine-grained dimensions. For example, fine-grained dimensions include at least two of the following: accuracy of medical facts, rigor of argumentation logic, consistency between text and images, and professionalism of typesetting. By evaluating the quality of each paper sample from multiple fine-grained dimensions, different dimensions of supervision information can be provided for the training of the multimodal evaluation model, thereby significantly improving the accuracy, robustness, and interpretability of the trained multimodal evaluation model's discriminative ability.
[0035] Furthermore, the initial evaluation model is trained using positive and negative paper samples and their corresponding quality score labels. This allows the model to learn from these finely labeled positive and negative paper samples, thereby gaining the ability to distinguish between the quality of papers.
[0036] For each large language model to be tested, the paper to be evaluated is input into the trained multimodal evaluation model. The multimodal evaluation model will evaluate the paper from different fine-grained dimensions and analyze its overall quality to obtain a quality score. This quality score can be used to measure the quality of the paper to be evaluated, thereby further reflecting the ability of the large language model to be tested that generated the paper to be evaluated in medical information research.
[0037] Step 103: Based on the quality scores of each paper to be evaluated, assess the medical information research capabilities of multiple large language models to be tested, and obtain the evaluation results.
[0038] In this step, the higher the quality score of a paper to be evaluated, the stronger the ability of the large language model to be tested in medical information research. Therefore, the ability of each large language model to be tested in medical information research can be determined based on the quality scores of each paper to be evaluated.
[0039] Alternatively, the quality scores of all papers to be evaluated can be sorted in descending order to obtain a ranking of the capabilities of all large language models to be tested in medical information research, thus achieving a multi-dimensional, interpretable, and horizontally comparable ranking of the performance of the large language models to be tested.
[0040] For example, to further improve the accuracy of the evaluation results and prevent evaluation bias caused by the limitations of a single research topic, medical research query information on different topics can be input into multiple large language models to be tested, resulting in papers to be evaluated corresponding to different topics generated by each large language model. After inputting each paper to be evaluated into a multimodal evaluation model and obtaining the quality score of each paper, the average quality score of all papers generated under different topics can be calculated for each large language model to be tested. Based on this average, the capability of the large language model to be tested in medical information research can be determined.
[0041] The method for evaluating the medical information research capabilities of large language models provided in this invention involves inputting medical research query information into multiple large language models to be tested, obtaining papers to be evaluated from the outputs of each model. For each paper, the method inputs it into a multimodal evaluation model to obtain a quality score. This evaluation model is trained on an initial evaluation model based on positive and negative paper samples and corresponding quality score labels. The quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions. Based on the quality scores of each paper, the medical information research capabilities of the multiple large language models are evaluated to obtain the evaluation results. Because standardized, unified input is used to ensure fairness at the starting point of the evaluation, and a unified, objective multimodal evaluation model trained with domain expertise is used as the evaluation standard, a standardized, objective, and quantitative evaluation of papers generated by different large language models can be achieved, thus enabling an accurate evaluation of the medical information research capabilities of large language models.
[0042] For example, based on the above embodiments, when inputting medical research query information into multiple large language models to be tested and obtaining the papers to be evaluated output by each large language model, this can be achieved in the following way: For each large language model to be tested, medical research query information is input into the supervised submodule of the large language model. The supervised submodule generates literature retrieval tasks, data analysis tasks, coding tasks, and writing tasks. The literature retrieval task is input into the literature review submodule of the large language model, the data analysis task into the data analysis submodule, the coding task into the coding submodule, and the writing task into the paper generation submodule. The literature review submodule retrieves literature related to the medical research query information. The data analysis submodule processes and analyzes the data related to the medical research query information. The coding submodule generates visualization charts. The paper generation submodule generates the paper to be evaluated based on the retrieved literature, the processed data, and the generated visualization charts.
[0043] Specifically, Figure 2 This is a schematic diagram of the evaluation framework for the medical information research capabilities of large language models provided in an embodiment of the present invention, such as... Figure 2 As shown, to ensure the fairness of the evaluation, a standardized testing environment module can be built in each language model to provide a unified and controllable framework for scientific research agents, thereby eliminating evaluation biases caused by different system architectures. This standardized testing environment module receives standardized medical research query information and uses the language model under test as the inference core to execute a predefined, end-to-end research workflow to generate papers for evaluation of the language model under test.
[0044] The standardized testing environment module includes a supervision submodule, a literature review submodule, a data analysis submodule, a coding submodule, and a paper generation submodule. The supervision submodule receives medical research query information and, utilizing the reasoning capabilities of the large language model under test, decomposes the query information into multiple executable structured subtasks. These subtasks include determining which keywords to retrieve from the literature, what statistical analysis to perform on the provided dataset, what types of charts to generate to present the results, and ultimately, how to generate the paper.
[0045] Furthermore, the supervision submodule can input multiple decomposed structured subtasks into different submodules. For example, the literature retrieval task can be input into the literature review submodule of the large language model under test. This submodule automatically retrieves and synthesizes existing knowledge related to medical research query information, providing theoretical background and reference support for the generated paper to be evaluated. The data analysis task can be input into the data analysis submodule of the large language model under test. This submodule cleans, statistically analyzes, and processes the input data related to medical research query information. The coding task can be input into the coding submodule of the large language model under test. This submodule converts the experimental plan into executable code and generates intuitive visualizations, such as scatter plots or bar charts, by running the converted code. The writing task can be input into the paper generation submodule of the large language model under test. This submodule, based on the retrieved literature, processed data, and generated visualizations, integrates these scattered contents into a logically coherent and formatted paper to be evaluated, following the standard structure of academic papers, such as introduction, methods, results, and discussion.
[0046] The literature review, data analysis, and coding submodules mentioned above can be executed in parallel, minimizing idle waiting time between tasks and significantly improving end-to-end efficiency from query input to the generation of the paper to be evaluated. Furthermore, these submodules can also be executed collaboratively and interactively under the unified scheduling of the supervision submodule, based on a pre-defined workflow logic. This allows for dynamic response to and integration of intermediate results generated by other submodules, thus satisfying data and logical dependencies between different stages in complex research tasks. For example, after initial data processing, the data analysis submodule may need to optimize the analysis process based on research methods retrieved by the literature review submodule; or, when generating visualizations, the coding submodule may need to use structured results produced by the data analysis submodule as input. This interactive execution mode not only achieves a balance between efficiency and resource consumption but also allows for multi-threaded, iterative work, thereby improving the quality of the generated paper to be evaluated.
[0047] In the above embodiments, by constructing a standardized testing environment module, a unified and structured scientific research intelligent agent framework can be provided. Medical research query information is decomposed into multiple sub-tasks, and the corresponding sub-tasks are executed through different sub-modules. This ensures that all large language models under test generate papers to be evaluated under the same process and conditions, fundamentally eliminating evaluation bias caused by differences in system architecture, workflow, or task understanding, thereby improving the objectivity and accuracy of the quality score of the papers to be evaluated.
[0048] For example, the multimodal evaluation model mentioned above is trained based on the following method: Multiple positive and negative paper samples are obtained. Fine-grained quality annotations are performed on each positive and negative paper sample across multiple dimensions to obtain quality score labels. Each positive and negative paper sample is input into a visual language model to obtain the global feature vector of the paper output by the visual language model, which is trained based on medical image and text data. The global feature vector of the paper is input into an initial scoring model to obtain the predicted quality score output by the initial scoring model. Loss information is determined based on the predicted quality score and quality score labels. The initial scoring model is iteratively trained based on the loss information to obtain the scoring model. The visual language model and the scoring model are determined as a multimodal evaluation model.
[0049] Specifically, continue to refer to Figure 2 As shown, the medical professional dataset construction unit in the multimodal paper quality assessment benchmark module can be used to construct a medical professional dataset containing both positive and negative paper samples. High-quality full-text papers can be obtained from medical journals or databases and used as positive paper samples. These positive paper samples include papers from several key medical subfields such as oncology, neuroscience, and cardiovascular disease.
[0050] For negative paper samples, the medical research query information of the sample can be input into multiple large language models to obtain the paper samples output by each language model. The paper samples output by each language model can then be identified as negative paper samples.
[0051] Specifically, standardized sample medical research query information on the same research topic can be input into multiple large language models to obtain paper samples of different quality levels output by each language model. These paper samples can be used as negative paper samples. The process of generating paper samples through the large language models is similar to that of generating papers to be evaluated. A standardized testing environment module can be used, with a supervision submodule to decompose the sample medical research query information into multiple sub-tasks. These sub-tasks are executed through literature review, data analysis, coding, and paper generation sub-modules, respectively, to generate paper samples. The negative paper samples contain diverse quality issues, including factual errors, logical inconsistencies, inconsistencies between text and graphics, improper statistical analysis, or non-standard formatting.
[0052] By inputting sample medical research query information into multiple large language models, multiple negative paper samples can be obtained. This enables the construction of a negative paper sample set that covers a wide range of error patterns and closely reflects real-world situations, providing an important data foundation for training a multimodal evaluation model with high generalization.
[0053] After obtaining positive and negative paper samples, for each paper sample, experts in the medical field will conduct fine-grained quality scoring across multiple dimensions based on pre-defined, standardized evaluation criteria, thus obtaining a quality score label for each paper sample. Quality scores can be assigned using a 1-5 point scale or a numerical representation within the range of 1-100. A higher quality score indicates better quality in a particular dimension or overall. Furthermore, the multiple dimensions can include at least two of the following: accuracy of medical facts, rigor of argumentation logic, consistency between text and figures, and professionalism of layout. Therefore, the quality score can include scores for each dimension as well as an overall quality score. These scores for each dimension, together with the comprehensive overall quality score, constitute a complete and structured quality score label for each paper sample.
[0054] Furthermore, it can be achieved through methods such as Figure 2 The domain-adaptive evaluation model unit shown constructs a multimodal evaluation model capable of understanding medical image and text content. This multimodal evaluation model includes a Visual Language Model (VLM) and a scoring model. The VLM prioritizes models fine-tuned on medical image and text data, such as LLaVA-Med, and uses this model as the backbone for feature extraction. Because this VLM has learned from a vast amount of image-text pairs in medical textbooks and research papers during its pre-training and fine-tuning phases, it constructs a deep medical prior knowledge base. This enables it to understand the complex semantic relationships between professional visual elements such as abnormal shadows in X-rays, cell morphology in pathological slides, and clinically significant trends in statistical charts and their corresponding medical text descriptions. Therefore, it can effectively interpret and provide feedback on complex medical visualizations such as medical images, pathological slides, and statistical charts.
[0055] During model training, some weights of the visual language model's backbone are frozen. A lightweight initial scoring model, such as a Multi-Layer Perceptron (MLP), is then fed into the visual language model. Positive and negative paper samples are input into the visual language model, which employs a block-based processing strategy to extract text and image features from long paper samples by chapter or page. These features are then aggregated using a lightweight Transformer to generate a paper-level global feature vector. This global feature vector is then input into the initial scoring model to obtain the predicted quality score in the [0,1] interval output by the initial scoring model.
[0056] By comparing the predicted quality score with the quality score label, the loss information can be determined. Based on the loss information, the model parameters of the initial scoring model are adjusted using the backpropagation algorithm. This process is repeated until the determined loss is minimized or the preset number of iterations is reached. The final model is then determined as the trained scoring model. The visual language model and the trained scoring model are then used as the final multimodal evaluation model.
[0057] By freezing some weights of the visual language model backbone and training the initial scoring model, the complexity of model training can be significantly reduced, the training efficiency can be improved, and the stability of the model can be enhanced, while ensuring the accuracy of the evaluation. This approach has good practicality.
[0058] For example, the quality score labels corresponding to the positive and negative paper samples each include the total paper quality score and the scores on each dimension; the predicted quality score output by the initial scoring model includes the predicted total quality score and the predicted scores on each dimension; therefore, when determining loss information based on the predicted quality score and quality score labels, the main loss can be determined based on the total paper quality score and the predicted total quality score, and the auxiliary loss can be determined based on the scores on each dimension and the predicted scores on each dimension, thereby determining the loss information based on the main loss and the auxiliary loss.
[0059] Specifically, continue to refer to Figure 2 As shown, a composite loss function including main loss and auxiliary loss can be designed through training strategies and optimization units to specifically train the initial scoring model, enabling the resulting multimodal evaluation model to score the papers to be evaluated. The main loss is a regression loss determined based on the labeled total paper quality score. By minimizing the main loss, it can be ensured that the predicted total quality score output by the model is as close as possible to the actual total paper quality score labeled by experts.
[0060] The auxiliary loss is a multi-task regression loss for scores on each fine-grained dimension. It can be understood as the difference between the model's predicted score on each dimension and the expert's score on that dimension. By minimizing each auxiliary loss, the model can be guided to learn the independent evaluation criteria for each fine-grained dimension, enabling the model to build more accurate and robust internal feature representations, thereby significantly improving the accuracy, interpretability, and generalization ability of its overall quality assessment.
[0061] By weighting and combining the main loss and auxiliary loss, we can obtain the final composite loss information. This loss information guides the model training process, enabling the synergistic optimization of overall constraints and micro-level fine-grained discrimination. This allows us to train a robust evaluation model that combines high scoring accuracy, strong generalization ability, and good interpretability.
[0062] Furthermore, LoRA and other parameter-efficient fine-tuning techniques can be used to perform lightweight unfreezing and fine-tuning of the backbone of the visual language model, thereby further improving the multimodal evaluation model's ability to assess the quality of papers in the medical field.
[0063] For example, based on the above embodiments, after obtaining the quality scores of each paper to be evaluated output by the multimodal evaluation model, it is also possible to determine the target paper to be evaluated with a quality score lower than a preset score based on the quality scores of each paper to be evaluated, and to use interpretability tools to analyze the target paper to be evaluated to obtain a diagnostic report. The diagnostic report includes the reasons for the low score of the target paper to be evaluated, and the diagnostic report is output.
[0064] Specifically, the preset score can be set based on actual circumstances or experience, for example, it can be set to 60% of the total score. After obtaining the quality score for each paper to be evaluated, target papers with quality scores lower than the preset score will be selected. For the selected target papers with significant problems, interpretable tools such as Grad-CAM or attention heatmaps will be used to analyze the reasons for the low scores in depth, generating a detailed diagnostic report. The reasons for the low scores may include at least one of the following: factual errors, logical fallacies, discrepancies between text and images, and poor typography.
[0065] In the above approach, for low-scoring target papers to be evaluated, interpretability tools can be used to analyze these target papers to determine the reasons for their low scores. This can generate diagnostic reports with clear guidance, trace and locate specific defects such as factual errors and logical fallacies in various language models, and provide precise directions for optimizing the medical research capabilities of large language models.
[0066] For example, based on the above embodiments, after obtaining the quality scores of each paper to be evaluated output by the multimodal evaluation model, the expert scores of each paper to be evaluated can also be obtained. For each paper to be evaluated, the Spearman rank correlation coefficient between the expert scores and the quality scores of the paper to be evaluated is determined. Based on the Spearman rank correlation coefficient, the quality scores output by the multimodal evaluation model are verified.
[0067] Specifically, it is possible to obtain expert scores from medical experts after blind review of all or part of the papers to be evaluated. These expert scores are obtained by medical experts independently scoring according to the same fine-grained dimensions.
[0068] For each paper to be evaluated, the Spearman rank correlation coefficient between the expert score and the quality score can be calculated. The Spearman rank correlation coefficient measures the consistency between the multimodal evaluation model's score and the expert score in the paper's quality ranking. The Spearman rank correlation coefficient ranges from [-1, 1]. A Spearman rank correlation coefficient of 1 indicates that the quality score ranking output by the multimodal evaluation model and the expert score ranking are consistent for that paper. A Spearman rank correlation coefficient of 0 indicates that the quality score ranking output by the multimodal evaluation model and the expert score ranking are uncorrelated. A Spearman rank correlation coefficient of 1 indicates that the quality score ranking output by the multimodal evaluation model and the expert score ranking are completely opposite.
[0069] In the above embodiments, for each paper to be evaluated, the Spearman rank correlation coefficient between the expert score and the quality score of the paper to be evaluated is determined, and the quality score output by the multimodal evaluation model is verified based on the Spearman rank correlation coefficient, thereby determining the reliability and validity of the quality score output by the multimodal evaluation model.
[0070] The following describes the evaluation device for medical information research capabilities of large language models provided by the present invention. The evaluation device for medical information research capabilities of large language models described below can be referred to in correspondence with the evaluation method for medical information research capabilities of large language models described above.
[0071] Figure 3 This is a schematic diagram of the structure of the device for evaluating the research capabilities of large language models in medical information, as provided in an embodiment of the present invention. Figure 3 As shown, the evaluation device 300 for assessing the medical information research capabilities of large language models includes: Input module 11 is used to input medical research query information into multiple large language models to be tested, and obtain the papers to be evaluated output by each of the large language models to be tested. The input module 11 is further configured to input the paper to be evaluated into a multimodal evaluation model for each of the papers to be evaluated, and obtain the quality score of the paper to be evaluated output by the multimodal evaluation model. The multimodal evaluation model is obtained by training an initial evaluation model based on positive paper samples, negative paper samples and corresponding quality score labels. The quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions. Evaluation module 12 is used to evaluate the medical information research capabilities of the multiple large language models to be tested based on the quality scores of each of the papers to be evaluated, and to obtain evaluation results.
[0072] In one example embodiment, the input module 11 is specifically used for: For each of the large language models to be tested, the medical research query information is input into the supervision submodule of the large language model to be tested. The supervision submodule generates literature retrieval tasks, data analysis tasks, coding tasks, and writing tasks. The literature retrieval task is input into the literature review submodule of the large language model to be tested, the data analysis task is input into the data analysis submodule of the large language model to be tested, the coding task is input into the coding submodule of the large language model to be tested, and the writing task is input into the paper generation submodule of the large language model to be tested. The literature review submodule is used to retrieve literature related to the medical research query information. The data analysis submodule processes and analyzes data related to the medical research query information. Generate visual charts through the coding submodule; The paper generation submodule generates the paper to be evaluated based on the retrieved literature, processed data, and generated visualization charts.
[0073] In one example embodiment, the multimodal evaluation model is trained based on the following method: Obtain multiple positive paper samples and multiple negative paper samples respectively; Fine-grained quality annotations were obtained for each of the positive and negative paper samples across multiple dimensions to obtain quality score labels. Each of the positive paper samples and each of the negative paper samples are input into the visual language model to obtain the global feature vector of the paper output by the visual language model. The visual language model is trained based on medical image and text data. The global feature vector of the paper is input into the initial scoring model to obtain the predicted quality score output by the initial scoring model; Loss information is determined based on the predicted quality score and the quality score label; The initial scoring model is iteratively trained based on the loss information to obtain the scoring model; The visual language model and the scoring model are identified as the multimodal evaluation model.
[0074] In one example embodiment, the quality score label includes a total paper quality score and scores on each dimension; the predicted quality score includes a predicted total quality score and predicted scores on each of the dimensions; the apparatus further includes a determining module, wherein the determining module is specifically used for: Based on the total paper quality score and the predicted total quality score, the principal loss is determined; Based on the scores in each dimension and the predicted scores in each dimension, an auxiliary loss is determined; The loss information is determined based on the main loss and the auxiliary loss.
[0075] In one example embodiment, the apparatus further includes an acquisition module, wherein the acquisition module is specifically used for: The sample medical research query information is input into multiple large language models to obtain the paper samples output by each of the large language models; The paper samples output by each of the large language models are determined as the negative paper samples.
[0076] In one example embodiment, the multiple dimensions include at least two of the following: accuracy of medical facts, rigor of argumentation logic, consistency between text and images, and professionalism of layout.
[0077] In one example embodiment, the device further includes an analysis module and an output module, wherein: The determining module is also used to determine, based on the quality scores of each of the papers to be evaluated, target papers to be evaluated whose quality scores are less than a preset score; The analysis module is used to analyze the target paper to be evaluated using interpretable tools to obtain a diagnostic report, which includes the reasons for the low score of the target paper to be evaluated. The output module is used to output the diagnostic report.
[0078] In one example embodiment, the apparatus further includes an acquisition module and a verification module, wherein: The acquisition module is used to acquire the expert scores for each of the papers to be evaluated. The determination module is used to determine the Spearman rank correlation coefficient between the expert scores and quality scores of each of the papers to be evaluated. The verification module is used to verify the quality score output by the multimodal evaluation model based on the Spearman rank correlation coefficient.
[0079] The apparatus of this embodiment can be used for any of the methods in the side embodiment of the method for evaluating the medical information research capability of large language models. Its specific implementation process and technical effects are similar to those in the side embodiment of the method for evaluating the medical information research capability of large language models. For details, please refer to the detailed description in the side embodiment of the method for evaluating the medical information research capability of large language models, which will not be repeated here.
[0080] Figure 4 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communications bus 440. The processor 410 can call logical instructions in the memory 430 to execute an evaluation method for the medical information research capabilities of large language models. This method includes: inputting medical research query information into multiple large language models to be tested, obtaining papers to be evaluated output by each of the large language models; inputting each paper to be evaluated into a multimodal evaluation model, obtaining a quality score output by the multimodal evaluation model, wherein the multimodal evaluation model is trained on an initial evaluation model based on positive paper samples, negative paper samples, and corresponding quality score labels, and the quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions; and evaluating the medical information research capabilities of the multiple large language models to be tested based on the quality scores of each paper to be evaluated, obtaining an evaluation result.
[0081] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0082] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the evaluation method for the medical information research capability of large language models provided by the above methods. The method includes: inputting medical research query information into multiple large language models to be tested to obtain papers to be evaluated output by each of the large language models to be tested; inputting each paper to be evaluated into a multimodal evaluation model to obtain a quality score of the paper to be evaluated output by the multimodal evaluation model, wherein the multimodal evaluation model is trained on an initial evaluation model based on positive paper samples, negative paper samples, and corresponding quality score labels, and the quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions; and evaluating the medical information research capability of the multiple large language models to be tested based on the quality scores of each paper to be evaluated to obtain an evaluation result.
[0083] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements an evaluation method for the medical information research capabilities of large language models provided by the methods described above. This method includes: inputting medical research query information into multiple large language models to be tested, obtaining papers to be evaluated output by each of the large language models; inputting each paper to be evaluated into a multimodal evaluation model, obtaining a quality score output by the multimodal evaluation model, wherein the multimodal evaluation model is trained on an initial evaluation model based on positive paper samples, negative paper samples, and corresponding quality score labels, and the quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions; and evaluating the medical information research capabilities of the multiple large language models to be tested based on the quality scores of each paper to be evaluated, obtaining an evaluation result.
[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0085] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating the research capabilities of large language models in medical information, characterized in that, include: Medical research query information is input into multiple large language models to be tested, and the papers to be evaluated are output by each of the large language models to be tested. For each of the aforementioned papers to be evaluated, the paper to be evaluated is input into a multimodal evaluation model to obtain the quality score of the paper to be evaluated output by the multimodal evaluation model. The multimodal evaluation model is obtained by training an initial evaluation model based on positive paper samples, negative paper samples and corresponding quality score labels. The quality score labels are determined after performing fine-grained quality evaluation on the corresponding paper samples from multiple dimensions. Based on the quality scores of each of the papers to be evaluated, the medical information research capabilities of the multiple large language models to be tested are evaluated, and the evaluation results are obtained. The multimodal evaluation model was trained using the following method: Obtain multiple positive paper samples and multiple negative paper samples respectively; Fine-grained quality annotations were obtained for each of the positive and negative paper samples across multiple dimensions to obtain quality score labels. Each of the positive and negative paper samples is input into a visual language model. The visual language model uses a block processing strategy to extract image and text features from each of the positive and negative paper samples according to chapters or pages. The features are then aggregated using a lightweight Transformer to generate a global feature vector for the paper. The visual language model is fine-tuned based on medical image and text data to enable the visual language model to understand the complex semantic relationships between visual elements and their corresponding medical text descriptions. The global feature vector of the paper is input into the initial scoring model to obtain the predicted quality score output by the initial scoring model; Loss information is determined based on the predicted quality score and the quality score label; The initial scoring model is iteratively trained based on the loss information to obtain the scoring model; The visual language model and the scoring model are identified as the multimodal evaluation model; Obtain multiple negative paper samples, including: The sample medical research query information is input into multiple large language models to obtain the paper samples output by each of the large language models; The paper samples output by each of the large language models are determined as the negative paper samples.
2. The method for evaluating the research capability of large language models in medical information according to claim 1, characterized in that, The process involves inputting medical research query information into multiple large language models to be tested, resulting in the output of the papers to be evaluated from each of the large language models. This includes: For each of the large language models to be tested, the medical research query information is input into the supervision submodule of the large language model to be tested. The supervision submodule generates literature retrieval tasks, data analysis tasks, coding tasks, and writing tasks. The literature retrieval task is input into the literature review submodule of the large language model to be tested, the data analysis task is input into the data analysis submodule of the large language model to be tested, the coding task is input into the coding submodule of the large language model to be tested, and the writing task is input into the paper generation submodule of the large language model to be tested. The literature review submodule is used to retrieve literature related to the medical research query information. The data analysis submodule processes and analyzes data related to the medical research query information. Generate visual charts through the coding submodule; The paper generation submodule generates the paper to be evaluated based on the retrieved literature, processed data, and generated visualization charts.
3. The method for evaluating the research capability of large language models in medical information according to claim 1, characterized in that, The quality score label includes the total paper quality score and the scores on each dimension; the predicted quality score includes the predicted total quality score and the predicted scores on each of the aforementioned dimensions; The step of determining loss information based on the predicted quality score and the quality score label includes: Based on the total paper quality score and the predicted total quality score, the principal loss is determined; Based on the scores in each dimension and the predicted scores in each dimension, an auxiliary loss is determined; The loss information is determined based on the main loss and the auxiliary loss.
4. The method for evaluating the research capability of large language models in medical information according to any one of claims 1-3, characterized in that, The multiple dimensions include at least two of the following: accuracy of medical facts, rigor of argumentation logic, consistency between text and images, and professionalism of layout.
5. The method for evaluating the research capability of large language models in medical information according to any one of claims 1-3, characterized in that, The method further includes: Based on the quality scores of each of the papers to be evaluated, target papers to be evaluated whose quality scores are less than a preset score are identified. The target paper to be evaluated is analyzed using interpretability tools to obtain a diagnostic report, which includes the reasons for the low score of the target paper to be evaluated. Output the diagnostic report.
6. The method for evaluating the research capability of large language models in medical information according to any one of claims 1-3, characterized in that, The method further includes: Obtain expert scores for each of the papers to be evaluated; For each of the papers to be evaluated, determine the Spearman rank correlation coefficient between the expert scores and quality scores of the papers to be evaluated; The quality score output by the multimodal evaluation model is verified based on the Spearman rank correlation coefficient.
7. An evaluation device for the research capability of large language models in medical information, characterized in that, include: The input module is used to input medical research query information into multiple large language models to be tested, and obtain the papers to be evaluated output by each of the large language models to be tested. The input module is further configured to input the paper to be evaluated into a multimodal evaluation model for each of the papers to be evaluated, and obtain the quality score of the paper to be evaluated output by the multimodal evaluation model. The multimodal evaluation model is obtained by training an initial evaluation model based on positive paper samples, negative paper samples and corresponding quality score labels. The quality score labels are determined after fine-grained quality evaluation of the corresponding paper samples from multiple dimensions. The evaluation module is used to evaluate the medical information research capabilities of the multiple large language models under test based on the quality scores of each of the papers to be evaluated, and to obtain the evaluation results. The multimodal evaluation model was trained using the following method: Obtain multiple positive paper samples and multiple negative paper samples respectively; Fine-grained quality annotations were obtained for each of the positive and negative paper samples across multiple dimensions to obtain quality score labels. Each of the positive and negative paper samples is input into a visual language model. The visual language model uses a block processing strategy to extract image and text features from each of the positive and negative paper samples according to chapters or pages. The features are then aggregated using a lightweight Transformer to generate a global feature vector for the paper. The visual language model is fine-tuned based on medical image and text data to enable the visual language model to understand the complex semantic relationships between visual elements and their corresponding medical text descriptions. The global feature vector of the paper is input into the initial scoring model to obtain the predicted quality score output by the initial scoring model; Loss information is determined based on the predicted quality score and the quality score label; The initial scoring model is iteratively trained based on the loss information to obtain the scoring model; The visual language model and the scoring model are identified as the multimodal evaluation model; Obtain multiple negative paper samples, including: The sample medical research query information is input into multiple large language models to obtain the paper samples output by each of the large language models; The paper samples output by each of the large language models are determined as the negative paper samples.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the evaluation method for the research capability of large language model medical information as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Automatic paper first draft generation method and system and electronic equipment
CN118966162A
Literature quality evaluation method and device, equipment, medium and product
CN122045671A