Evaluation system of medical large model

By constructing an evaluation system based on a large medical model, the problem of lacking a unified evaluation throughout the entire life cycle in existing technologies has been solved. This system achieves fully automated comprehensive assessment, improves evaluation efficiency and accuracy, and supports flexible configuration and weight adjustment of modules.

CN121167232APending Publication Date: 2025-12-19SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511224041.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing medical large-scale model evaluation systems lack a unified evaluation throughout the entire lifecycle. Evaluation indicators are fragmented and cannot effectively cover the model development, deployment, and iteration processes, especially in terms of safety and performance assessment, where an integrated evaluation system is lacking.

Method used

This system provides an evaluation system for large-scale medical models, including modules for compliance evaluation, basic model deployment evaluation, basic model identification, general security evaluation, medical security evaluation, application service deployment evaluation, timeliness evaluation, performance evaluation, and page security evaluation. It conducts a comprehensive evaluation of the entire lifecycle through a serial structure, supports the addition or removal of modules as needed, and generates detailed evaluation reports.

Benefits of technology

It achieves fully automated evaluation of the entire lifecycle of medical large-scale model development, deployment, and iteration, improving evaluation efficiency, reducing evaluation difficulty, and providing more accurate evaluation results. It also supports automatic invocation of evaluation modules and adjustment of module weights for different types of medical large-scale models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167232A_ABST
    Figure CN121167232A_ABST
Patent Text Reader

Abstract

The invention discloses an evaluation system of a medical large model, which comprises an evaluation report generation module, the compliance evaluation module, the basic model deployment evaluation module, the basic model discrimination module, the general safety evaluation module, the medical safety evaluation module, the application service deployment evaluation module, the timeliness evaluation module, the performance evaluation module and the page safety evaluation module are in communication connection with the system. The evaluation report generation module generates an evaluation report based on evaluation results of the compliance evaluation module, the basic model deployment evaluation module, the basic model discrimination module, the general safety evaluation module, the medical safety evaluation module, the application service deployment evaluation module, the timeliness evaluation module, the performance evaluation module and the page safety evaluation module. The evaluation system can be used for evaluating the performance and safety of the whole life cycle of the medical large model, and is simple to operate and relatively high in efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large model evaluation technology, and in particular to an evaluation system for a large medical model. Background Technology

[0002] Medical big data models are complex and highly accurate mathematical models obtained by training medical data through deep learning. Combining deep learning and big data technologies, they can handle complex data relationships and provide more accurate predictions and diagnoses. Medical big data models offer healthcare professionals new tools for disease prediction and risk assessment, aiming to improve the accuracy of disease diagnosis, promote personalized health management plans, and achieve precision in disease prevention strategies. With the explosive growth of medical data and the continuous advancement of computing technology, it has become possible to develop AI models that can efficiently process this data and provide accurate predictions.

[0003] The development and deployment of large-scale medical models often require various evaluations to ensure their safety and performance. However, there is currently no integrated evaluation system that can achieve comprehensive assessments throughout the entire lifecycle of large-scale medical models. Existing evaluation processes are fragmented, with safety and performance assessments typically conducted separately, lacking a unified evaluation system covering model development, deployment, and iteration. Summary of the Invention

[0004] To enable various evaluations of large-scale medical models throughout their entire lifecycle, this invention provides an evaluation system for large-scale medical models, comprising:

[0005] The compliance assessment module is used to conduct compliance assessments based on the information of the medical big data model input by the user, and to confirm the filing information of the medical big data model.

[0006] The basic model deployment and evaluation module is used for the basic model deployment and evaluation of the aforementioned large medical model.

[0007] A schema identification module is used to identify schemas to determine the safety of the source of the medical large model;

[0008] A general security assessment module, which is used for the general security assessment of the aforementioned large medical model;

[0009] A medical safety assessment module, which is used for medical safety assessment of the aforementioned large medical model;

[0010] The application service deployment evaluation module is used for evaluating the application service deployment of the aforementioned medical big data model.

[0011] The timeliness evaluation module is used to evaluate the timeliness of the medical big data model.

[0012] The performance evaluation module is used to evaluate the performance of the large medical model.

[0013] A page security assessment module, used to assess the page security of the aforementioned large medical model; and

[0014] The evaluation report generation module is communicatively connected to the compliance evaluation module, basic model deployment evaluation module, basic model identification module, general security evaluation module, medical security evaluation module, application service deployment evaluation module, timeliness evaluation module, performance evaluation module, and page security evaluation module, and is used to generate evaluation reports based on the evaluation results of the compliance evaluation module, basic model deployment evaluation module, basic model identification module, general security evaluation module, medical security evaluation module, application service deployment evaluation module, timeliness evaluation module, performance evaluation module, and page security evaluation module.

[0015] Furthermore, the evaluation system also includes:

[0016] An interactive page is used to receive user input to obtain information about the medical big data model, wherein the user inputs the information about the medical big data model through text input, and / or drop-down options, and / or check boxes.

[0017] Furthermore, the compliance evaluation module, basic model deployment evaluation module, basic model identification module, general security evaluation module, medical security evaluation module, application service deployment evaluation module, timeliness evaluation module, performance evaluation module, and page security evaluation module are pluggable, and one or more modules can be omitted based on user input information.

[0018] If the information entered by the user in the medical big data model does not include the filing information, the compliance assessment module will be omitted.

[0019] If the aforementioned medical model is intended for internal institutional use or research purposes only, the page security assessment module will be omitted; and

[0020] If the medical big model is a non-local deployment model, then the basic model deployment evaluation module, application service deployment evaluation module, and basic model identification module are omitted.

[0021] Furthermore, the basic model deployment evaluation module generates automatic test scripts based on the medical large model image deployment specifications.

[0022] Furthermore, the schema identification module performs schema identification based on representation in order to identify and trace the source of the medical big model.

[0023] Furthermore, the security evaluation module includes:

[0024] A general security assessment submodule includes seven dimensions of assessment items: fairness and discrimination, commercial violations, infringement of others' legitimate rights and interests, inability to meet the security requirements of specific service types, questions the model should refuse to answer, and questions the model should not refuse to answer, as well as a large-scale adjudicator model. This large-scale adjudicator model is used to perform a general security score on the medical large-scale model based on the seven dimensions of assessment items.

[0025] The medical safety assessment submodule includes assessment items in two dimensions: medical ethics and drug contraindications.

[0026] Furthermore, the application service deployment evaluation module generates automatic test scripts based on the application service image deployment specifications.

[0027] Furthermore, the timeliness evaluation module is used to test the first token generation time, the time required for each output token, the end-to-end time, the input throughput, the output throughput, and the number of requests completed per minute of the medical big data model.

[0028] Furthermore, the performance evaluation module is used to construct a scenario evaluation dataset based on a specific medical scenario question bank, and to evaluate the performance of the large medical model through the scenario evaluation dataset. The performance evaluation metrics include precision, micro-average F1 score, BERT score, and macro-average recall.

[0029] Furthermore, the page security evaluation module is used to train an adversarial example generation model based on historical fault data, automatically synthesize extreme scenario inputs, and use the extreme scenario inputs as inputs to the medical big model for attack and defense, so as to test page security.

[0030] Furthermore, the evaluation report includes the task background and model capability analysis, wherein the model capability analysis includes the comprehensive score of the medical large model application scenario. The comprehensive score of the medical large model application scenario is obtained by weighted averaging of the safety evaluation, performance evaluation and timeliness evaluation results, and the weight of each individual capability is dynamically adjusted using a reinforcement learning strategy.

[0031] This invention provides an evaluation system for large-scale medical models, covering the entire lifecycle of model development, deployment, and iteration. It employs a sequential structure and allows for the addition or removal of evaluation modules as needed. The entire evaluation process is fully automated, effectively improving the efficiency and reducing the difficulty of large-scale model evaluation. The evaluation system can automatically select evaluation modules based on the type of the large-scale medical model and adjust the weights of the evaluation results for each module, resulting in more accurate evaluation results. Attached Figure Description

[0032] To further illustrate the above and other advantages and features of the various embodiments of the present invention, a more specific description of the various embodiments of the present invention will be presented with reference to the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not intended to limit its scope. In the drawings, identical or corresponding parts will be indicated by identical or similar reference numerals for clarity.

[0033] Figure 1 This diagram illustrates a flowchart of a computer-executed medical large-scale model evaluation method according to an embodiment of the present invention; and

[0034] Figure 2 This diagram illustrates the structure of a medical large-scale model evaluation system according to an embodiment of the present invention. Detailed Implementation

[0035] In the following description, the invention is described with reference to various embodiments. However, those skilled in the art will recognize that the embodiments may be practiced without one or more specific details or with other alternatives and / or additional methods, materials, or components. In other instances, well-known structures, materials, or operations are not shown or described in detail so as not to obscure the inventive points of the invention. Similarly, for illustrative purposes, specific quantities, materials, and configurations are set forth to provide a comprehensive understanding of embodiments of the invention. However, the invention is not limited to these specific details. Furthermore, it should be understood that the embodiments shown in the drawings are illustrative representations and are not necessarily drawn to scale.

[0036] In this specification, references to "an embodiment" or "this embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. The phrase "in one embodiment" appearing throughout this specification does not necessarily refer to the same embodiment in all instances.

[0037] It should be noted that the embodiments of the present invention describe the process steps in a specific order; however, this is only for illustrating the specific embodiment and not for limiting the order of the steps. On the contrary, in different embodiments of the present invention, the order of the steps can be adjusted according to the process.

[0038] In order to realize a multi-dimensional and quantifiable medical big model and to evaluate it throughout its entire life cycle, this invention discloses a computer-executed medical big model evaluation method that links compliance, security, performance evaluation, etc., to form a full-chain evaluation index system.

[0039] The technical solution of the present invention will be further described below with reference to the accompanying drawings of the embodiments.

[0040] Figure 1This diagram illustrates a flowchart of a computer-executed medical large-scale model evaluation method according to an embodiment of the present invention. Figure 1 As shown, an evaluation method for a computer-executed large-scale medical model includes:

[0041] First, in step 101, compliance assessment. The user submits an assessment application to initiate the assessment. This application includes the name of the medical large-scale model, registration information, model information, etc. After receiving this information, the computer performs a compliance assessment. In one embodiment of the invention, the compliance assessment verifies whether the information of the medical large-scale model is consistent with the registration information; if they are consistent, the assessment passes. In one embodiment of the invention, the medical large-scale model is a medical large-scale model trained based on a registered base model, or a registered medical large-scale model.

[0042] Next, in step 102, the basic model is deployed and evaluated. The user provides the model image address and model ID based on the medical large-scale model image deployment specification. The model ID can be used to call the model service API to ensure the image runs normally. The computer uses an automated test script to perform automated deployment testing based on the model image address and model ID. In one embodiment of the invention, the medical large-scale model image deployment specification provides users with detailed steps and necessary code when deploying the model, so as to extract intermediate layer representations (activations) of specific layers during the model inference stage. The medical large-scale model image deployment specification can guide users to install the necessary dependency packages for base model detection in the image environment, how to run the code to extract model representations, and provide debugging methods in case of possible errors. The necessary dependency packages for base model detection include Python, PyTorch, transformers, accelerate, pandas, numpy, and other runtime dependency packages. The model to be tested needs to be provided in the image for invocation and access. The model image includes runtime environment dependencies, image files, and inference code to ensure out-of-the-box usability. In one embodiment of the present invention, the model image needs to additionally include model representation extraction code with added hooks for base model identification, wherein the corresponding inference file of the model representation extraction code can be stored in a specified path. In one embodiment of the present invention, a basic model deployment evaluation item is provided through a simulated smoke test generator. The evaluation item can be synchronized in the simulated smoke test generator after local simulation, and then automatically invoked for testing after the user inputs the model image address and model ID;

[0043] Next, in step 103, the baseline model is identified. After the basic model deployment evaluation is passed, baseline model identification is performed to identify and trace the source of the large medical model, and further determine whether its source is safe. In one embodiment of the present invention, the large model under test is traced based on the representation to determine whether it is fine-tuned based on a certain basic model. This does not require pre-implanting watermarks into the large model under test, has no impact on model performance, is robust to subsequent fine-tuning, pruning, and quantization, and has the advantages of low time cost and simplicity.

[0044] Next, in step 104, a general security assessment is conducted. After the baseline model is identified, a general security assessment is performed by computer to improve the ability to identify harmful content and misleading information by utilizing the generalization ability of the large model. The general security assessment includes: extracting assessment items from a preset question bank, wherein the assessment items include seven dimensions: unfair discrimination, commercial violations, infringement of others' legitimate rights and interests, inability to meet the security requirements of specific service types, questions the model should refuse to answer, and questions the model should not refuse to answer. Then, the large medical model is scored based on the seven dimensions of the assessment items by the judge's large model. If the final score is higher than a preset value, it indicates that the general security assessment has been passed. In one embodiment of the present invention, the preset value is 90 points.

[0045] Next, in step 105, a medical safety assessment is conducted. After passing the general safety assessment, a medical safety assessment is performed using a computer to evaluate the ethical and medical safety aspects for medical applications. This assessment mainly includes two dimensions: medical ethics and drug contraindications. Similarly, in one embodiment of the present invention, the medical safety assessment includes extracting assessment items from a preset question bank and scoring the large medical model based on the extracted assessment items. If the final score is higher than a preset value, it indicates that the general safety assessment has been passed. In one embodiment of the present invention, the preset value is 80 points. In one embodiment of the present invention, the algorithm evaluation metric accuracy is used as the score.

[0046] Next, in step 106, the application service is deployed and evaluated. If both general security and medical security evaluations are passed, the medical large-scale model application service can be deployed. At this point, the computer performs automatic deployment testing. In one embodiment of the invention, the user fills in the model image address, model ID, and application model encapsulation description based on the application service image deployment specification. The computer then performs automatic deployment testing based on the user's input. It should be noted that the medical large-scale model application service should share the same basic model foundation as the subsequently deployed model application and the basic model in step 102. In one embodiment of the invention, the application service image deployment specification provides users with detailed steps and necessary code when deploying model application service images (groups), enabling necessary input and output for the corresponding scenario evaluation set during the model inference stage in conjunction with verification and evaluation. To ensure the image meets inference requirements, users must include Python, PyTorch, transformers, accelerate, pandas, numpy, and other runtime dependencies when building the image. If, in addition to the basic model, it also includes RAG, agent, and other project calls, these must be packaged into a single image. The image includes runtime environment dependencies, image files, inference code, etc. The deployed model application service image (group) must conform to the preset calling specifications, and the "model ID" must be filled in on the deployment page to call the model service API to ensure the image runs normally. In one embodiment of the invention, after the image runs, it should provide offline calling functionality for the model interface. The interface calling format follows the preset input and output formats for streaming and non-streaming calls. In one embodiment of the invention, an application service deployment evaluation item is provided through a simulated smoke test generator, wherein the application service deployment evaluation item should meet the requirements of the application service image deployment specifications. The evaluation item can be simulated locally and synchronized in the simulated smoke test generator, and then automatically called for testing after the user inputs the model image address and model ID.

[0047] Next, in step 107, timeliness evaluation is conducted. After the application service deployment evaluation passes, a medical scenario evaluation is performed, including timeliness evaluation based on a hybrid scenario and model performance evaluation based on the selected scenario. In one embodiment of the present invention, the timeliness evaluation includes the first token generation time (TTFT), the time required for each output token (TPOT), end-to-end time (E2E), input throughput, output throughput, and requests per minute (RPM), etc., the purpose of which is to stress test the application service. Here, TTFT refers to the time from when the model starts working to when the first token is output, TPOT refers to the time required for the model to generate each output token, E2E refers to the time required for the entire process from input to output, input throughput refers to the number of input tokens processed per unit time, and output throughput refers to the number of output tokens processed per unit time. In one embodiment of the present invention, the computer first needs to match computing power according to the parameter quantity of the large model. For example, if the parameter quantity is less than or equal to 30B, two A800 chips are automatically matched for automatic evaluation; if the parameter quantity is greater than 30B and less than or equal to 123B, four A800 chips are automatically matched for automatic evaluation; if the parameter quantity is greater than 123B, automatic evaluation is not supported, and an alarm is issued to remind the user to manually perform the evaluation. The parameter quantity is provided by the user in step 101. In one embodiment of the present invention, no fixed-length token is output during the timeliness evaluation process; only the index calculation logic is used.

[0048] Next, in step 108, performance evaluation. A scenario evaluation dataset is constructed based on a specific medical scenario question bank to evaluate the performance of the medical large model in simulated application scenarios. In one embodiment of the present invention, the scenario evaluation dataset is automatically constructed based on the type of the medical large model. In one embodiment of the present invention, the performance evaluation metrics include accuracy, micro-F1 score, BERT score, and macro-recall, where the F1 score is the harmonic mean of precision and recall. The core logic of micro-F1 is to first summarize the total precision (TP), total precision (FP), and total recall (FN) of all categories, then calculate the global precision (Micro-Precision) and global recall (Micro-Recall) based on the total TP, total FP, and total FN, and finally calculate the harmonic mean of the two. BERT Score is an evaluation metric for text generation tasks based on pre-trained language models, such as BERT. Its core logic involves calculating the semantic similarity between the "generated text's token" and the "reference text's token" using BERT's word embeddings, and then calculating the final score using matching strategies such as the Hungarian algorithm. It better captures the semantic consistency of the text. Macro-Recall is a "macro-average" calculation method for recall. Its core logic is to first calculate the recall rate for each category separately, and then take the arithmetic mean of the recall rates for all categories.

[0049] Next, in step 109, page security assessment. Users manually input test questions on the page to test its security. In one embodiment of the invention, a bidirectional detection engineering model is constructed based on a medical adversarial sample library to identify spoofed attacks in user input in real time. Furthermore, a medical intent ambiguity analysis engine is introduced to distinguish between the legitimate use of clinical terminology and malicious commands. In another embodiment of the invention, an adversarial sample generation model is trained based on historical fault data to automatically synthesize extreme scenario inputs, such as embedding drug name confusion words in medical text, thereby improving recognition accuracy; and

[0050] Finally, in step 110, an evaluation report is generated. Based on the evaluation results from steps 101 to 109, an evaluation report is generated. In one embodiment of the present invention, the evaluation report includes two parts: task background and model capability analysis. The task background includes a scenario introduction, evaluation overview, model details, and base model review results. The model details include the application model name, number of parameters, whether it is open source, model release date, and model description. The model capability analysis includes security evaluation, performance evaluation, and timeliness evaluation. As mentioned above, the security evaluation includes general security and medical ethics security. In one embodiment of the present invention, the comprehensive score of the medical large-scale model application scenario is determined based on the results of the security evaluation, performance evaluation, and timeliness evaluation. In one embodiment of the present invention, the final comprehensive score of the model is obtained by weighted averaging of each individual item, wherein the weights of each individual capability are dynamically adjusted using a reinforcement learning strategy.

[0051] In one embodiment of the present invention, if the medical large-scale model requires application for clinical projects, after obtaining the evaluation report, it is necessary to proceed to step 111, ethical review. The evaluation report is then sent to the ethics review committee for ethical review.

[0052] In one embodiment of the present invention, each evaluation is implemented by a separate module, wherein the module can be implemented using software, hardware, firmware, or a combination thereof. When a module is implemented using software, its function can be implemented through computer program flow. For example, the module can be implemented using code segments (such as code segments in languages ​​like C or C++) stored in a storage device (such as a hard disk, memory, etc.), wherein the corresponding function of the module can be implemented when the code segments are executed by a processor. When a module is implemented using hardware, its function can be implemented by setting up a corresponding hardware structure, for example, by hardware programming a programmable device such as a field-programmable gate array (FPGA), or by designing an application-specific integrated circuit (ASIC) that includes multiple electronic devices such as transistors, resistors, and capacitors. When a module is implemented using firmware, the module's function can be written in the form of program code into a read-only memory of the device, such as an EPROM or EEPROM, and the corresponding function of the module can be implemented when the program code is executed by a processor.

[0053] Figure 2 This diagram illustrates the structure of a medical large-scale model evaluation system according to an embodiment of the present invention. Figure 2As shown, an evaluation system for a large medical model includes: a compliance evaluation module 201, a basic model deployment evaluation module 202, a basic model identification module 203, a general security evaluation module 204, a medical security evaluation module 205, an application service deployment evaluation module 206, a timeliness evaluation module 207, a performance evaluation module 208, a page security evaluation module 209, and an evaluation report generation module 210. The compliance evaluation module 201 is used to perform the compliance evaluation described in step 101; the basic model deployment evaluation module 202 is used to perform the basic model deployment evaluation described in step 102; the basic model identification module 203 is used to perform the basic model identification described in step 103; the general security evaluation module 204 is used to perform the general security evaluation described in step 104; the medical security evaluation module 205 is used to perform the medical security evaluation described in step 105; the application service deployment evaluation module 206 is used to perform the application service deployment evaluation described in step 106; the timeliness evaluation module 207 is used to perform the timeliness evaluation described in step 107; the performance evaluation module 208 is used to perform the performance evaluation described in step 108; and the page security evaluation module 209 is used to perform the evaluation described in step 109. The page security evaluation and evaluation report generation module 210 is communicatively connected to the compliance evaluation module 201, the basic model deployment evaluation module 202, the basic model identification module 203, the general security evaluation module 204, the medical security evaluation module 205, the application service deployment evaluation module 206, the timeliness evaluation module 207, the performance evaluation module 208, and the page security evaluation module 209. It is used to generate an evaluation report based on the evaluation results of the compliance evaluation module 201, the basic model deployment evaluation module 202, the basic model identification module 203, the general security evaluation module 204, the medical security evaluation module 205, the application service deployment evaluation module 206, the timeliness evaluation module 207, the performance evaluation module 208, and the page security evaluation module 209.

[0054] In one embodiment of the present invention, the evaluation system further includes an interactive page 211, which is communicatively connected to the compliance evaluation module 201, the basic model deployment evaluation module 202, the basic model identification module 203, the general security evaluation module 204, the medical security evaluation module 205, the application service deployment evaluation module 206, the timeliness evaluation module 207, the performance evaluation module 208, and the page security evaluation module 209. The interactive page 211 is used to receive user input information, wherein the user can input relevant information of the medical big model through text input, drop-down options, check boxes, etc.

[0055] In one embodiment of the present invention, the module can be plugged in or unplugged according to actual needs, thus omitting one or more of steps 101 to 109, such as... Figure 1 The flowchart shown and Figure 2In the structural diagram shown, the steps or modules indicated by dashed boxes represent steps and modules that can be skipped or omitted. In one embodiment of the invention, if the medical big data model is not registered, a compliance assessment is not required; instead, a basic model deployment assessment is performed directly. In practical applications, the compliance assessment can be determined based on the user's input; that is, if the "registration information" field is empty, the compliance assessment is skipped. In one embodiment of the invention, if the medical big data model does not pose a public opinion risk and is only for internal institutional use or research purposes, a page security assessment is not required. In one embodiment of the invention, when the user inputs relevant information about the medical big data model, they can determine whether to check the option "Does not pose a public opinion risk, only for internal institutional use, or only for research purposes" based on the actual situation. If checked, the page security assessment is skipped. In one embodiment of the invention, if a non-localized deployment model (API) is used, a basic model deployment assessment, basic model identification, and application service deployment assessment are not required. In one embodiment of the present invention, when a user inputs information related to a medical big data model, the user can select the deployment method through a drop-down menu or by manual input. If the model is deployed locally, basic model deployment evaluation, basic model identification, and application service deployment evaluation are required. If the model is deployed non-locally, basic model deployment evaluation, basic model identification, and application service deployment evaluation are not required.

[0056] Although various embodiments of the invention have been described above, it should be understood that they are presented by way of example only and not as limitations. It will be apparent to those skilled in the art that various combinations, modifications, and alterations can be made without departing from the spirit and scope of the invention. Therefore, the breadth and scope of the invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined solely by the appended claims and their equivalents.

Claims

1. A system for evaluating a medical large model, characterized by, compliance evaluation module configured to perform compliance evaluation based on user-inputted information of the medical large model, and confirm the filing information of the medical large model; a base model deployment evaluation module configured to perform base model deployment evaluation on the medical large model; a base model discrimination module configured to perform base model discrimination to determine the security of the source of the medical large model; a general security evaluation module configured to perform general security evaluation on the medical large model; a medical safety evaluation module configured to perform medical safety evaluation on the medical large model; an application service deployment evaluation module configured to perform application service deployment evaluation on the medical large model; a timeliness evaluation module configured to evaluate the timeliness of the medical large model; a performance evaluation module configured to evaluate the performance of the medical large model; a page security evaluation module configured to evaluate the page security of the medical large model; and an evaluation report generation module communicatively connected with the compliance evaluation module, the base model deployment evaluation module, the base model discrimination module, the general security evaluation module, the medical safety evaluation module, the application service deployment evaluation module, the timeliness evaluation module, the performance evaluation module, and the page security evaluation module, and configured to generate an evaluation report based on the evaluation results of the compliance evaluation module, the base model deployment evaluation module, the base model discrimination module, the general security evaluation module, the medical safety evaluation module, the application service deployment evaluation module, the timeliness evaluation module, the performance evaluation module, and the page security evaluation module. Further comprising:

2. The evaluation system of claim 1, wherein an interactive page configured to receive user input to obtain information of the medical large model, wherein the user inputs the information of the medical large model through text input, and / or drop-down options, and / or checked options. The compliance evaluation module, the base model deployment evaluation module, the base model discrimination module, the general security evaluation module, the medical safety evaluation module, the application service deployment evaluation module, the timeliness evaluation module, the performance evaluation module, and the page security evaluation module are pluggable structures, and one or more modules can be omitted based on user-inputted information:

3. The evaluation system of claim 1, wherein If the user-inputted information of the medical large model does not include filing information, the compliance evaluation module is omitted; If the medical large model is only used for internal use by an institution or for research purposes, the page security evaluation module is omitted; and If the medical large model is a non-local deployment model, the base model deployment evaluation module, the application service deployment evaluation module, and the base model discrimination module are omitted. The base model deployment evaluation module generates an automatic test script based on a medical large model image deployment specification; and / or 4. The evaluation system of claim 1, wherein The application service deployment evaluation module generates an automatic test script based on an application service image deployment specification. The base model discrimination module performs base model discrimination based on representation to identify and trace the source of the medical large model.

5. The evaluation system of claim 1, wherein, The security evaluation module comprises:

6. The evaluation system of claim 1, wherein, ​ a general safety evaluation submodule including evaluation items of seven dimensions of impartial discrimination, commercial illegal violation, infringement of others' legal rights, inability to meet the safety requirements of specific service types, questions that the model should refuse to answer, and questions that the model should not refuse to answer, and a referee large model configured to perform general safety scoring on the medical large model based on the evaluation items of the seven dimensions; and a medical safety evaluation submodule including evaluation items of two dimensions of medical ethics and drug contraindications.

7. The evaluation system of claim 1, wherein, The timeliness evaluation module is configured to test the first token generation time, the time required for each output token, the end-to-end time, the input throughput, the output throughput, and the number of requests completed per minute of the medical large model.

8. The evaluation system of claim 1, wherein, The performance evaluation module is configured to construct a scenario evaluation dataset based on a specific medical scenario question bank and evaluate the performance of the medical large model through the scenario evaluation dataset, wherein the indicators of the performance evaluation include accuracy, micro-average F1 score, BERT score, and macro-average recall rate.

9. The evaluation system of claim 1, wherein, The page security evaluation module is configured to train an adversarial sample generation model based on historical failure data, automatically synthesize extreme scenario inputs, and perform attack and defense on the medical large model by taking the extreme scenario inputs as inputs to test the page security.

10. The evaluation system of claim 1, wherein, The evaluation report includes task background and model capability analysis, wherein the model capability analysis includes a comprehensive score of the application scenario of the medical large model, and the comprehensive score of the application scenario of the medical large model is obtained by weighted average from the safety evaluation, performance evaluation, and timeliness evaluation results, and the weight of each single capability is dynamically adjusted using reinforcement learning strategy.

Citation Information

Cited By

  • Illusion evaluation method and device for large medical model

    CN121768691A

  • Performance measurement and evaluation system for large model driven distributed system

    CN122220189A