Large-model-oriented evaluation decision-making method, apparatus and device, and medium
By generating evaluation reports and conducting quality gate checks during large model update events, and comparing performance and business metrics in real time, the problem of low accuracy in large model evaluation in existing technologies is solved, enabling dynamic and objective evaluation and decision-making, and improving the security and stability in the financial and medical fields.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-19
AI Technical Summary
Existing large-scale models have low accuracy in evaluation and decision-making, are disconnected from the evaluation and deployment process, rely on manual judgment, and have delayed fault response, making it difficult to meet the security, compliance and stability requirements of the financial and healthcare industries.
In the model update event, a predefined evaluation task is executed in the isolated environment to generate an evaluation report, perform quality access control checks, collect and compare performance and business indicators in real time, and execute traffic scheduling decisions based on the real-time indicator comparison results.
By replacing static and subjective human judgment with dynamic and objective online verification, the accuracy of large model evaluation and decision-making is improved, ensuring model quality and reducing fault response time.
Smart Images

Figure CN122064569A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology and can be applied to the fields of fintech and healthcare. In particular, it relates to an evaluation and decision-making method, apparatus, device, and medium for large models. Background Technology
[0002] With the increasing application of large language models in core businesses such as intelligent investment advisory and risk assessment in the fintech field, and assisted diagnosis and medical record generation in the healthcare field, the need for model iteration is becoming more and more frequent. However, the current mainstream delivery and governance models have significant bottlenecks and are difficult to match the stringent requirements for security, compliance and stability in industries such as finance and healthcare.
[0003] First, there is a severe disconnect between model evaluation and deployment processes. Whether it's compliance testing of financial models or clinical validation of medical AI, evaluations are mostly offline, one-off activities. The results are difficult to automatically translate into decisions for deployment, creating a "quality gap." Second, critical decisions heavily rely on human intervention. In financial scenarios, model deployment depends on expert review; in medical scenarios, the clinical use of AI-assisted tools requires lengthy ethical and efficacy reviews. Both lack real-time data support, resulting in long decision delays and high risks. Third, fault detection and response are lagging. Once an online model experiences performance degradation or compliance / ethical deviations, these issues often only surface after customer complaints in the financial sector, and may only be noticed after adverse events occur in the medical field. Manual investigation and rollback are time-consuming. Therefore, current large-scale model evaluation decisions heavily rely on human intervention, leading to delayed fault response and low accuracy in evaluation decisions. Summary of the Invention
[0004] This invention provides an evaluation and decision-making method, apparatus, computer equipment, and medium for large models, in order to solve the technical problem of low accuracy in existing large model evaluation and decision-making.
[0005] Firstly, it provides an evaluation and decision-making method for large models, including: In response to model update events of large models, predefined evaluation tasks are executed in an isolated environment to generate evaluation reports; Based on the evaluation report, a quality access control check is performed on the target model version corresponding to the model update event; If the target model version passes the quality access control check, initial online traffic is allocated to the target model version, and performance evaluation indicators and business indicators of the target model version and the baseline model version are collected and compared in real time during the service process to obtain real-time indicator comparison results. Traffic scheduling decisions are made based on the real-time indicator comparison results.
[0006] Secondly, an evaluation and decision-making apparatus for large models is provided, including: The execution generation unit is used to respond to model update events of large models and perform predefined evaluation tasks to generate evaluation reports in an isolated environment; The inspection unit is used to perform a quality access control check on the target model version corresponding to the model update event based on the evaluation report. The allocation comparison unit is used to allocate initial online traffic to the target model version if the target model version passes the quality access control check, and to collect and compare the performance evaluation indicators and business indicators of the target model version and the baseline model version in real time during the service process to obtain real-time indicator comparison results. The execution unit is used to make traffic scheduling decisions based on the real-time indicator comparison results.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned evaluation and decision-making method for large models.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned evaluation and decision-making method for large models.
[0009] The aforementioned scheme, implemented by the evaluation and decision-making method, apparatus, computer equipment, and storage medium for large models, can respond to model update events of large models by executing predefined evaluation tasks in an isolated environment to generate evaluation reports. Based on the evaluation reports, a quality gate check is performed on the target model version corresponding to the model update event. If the target model version passes the quality gate check, initial online traffic is allocated to the target model version. Performance evaluation metrics and business metrics of the target model version and the baseline model version are collected and compared in real time during the service process to obtain real-time metric comparison results. Traffic scheduling decisions are then executed based on the real-time metric comparison results. In this invention, before the model goes online, a quality gate check is performed on the target model version based on the generated evaluation report to ensure the basic quality of the model. Then, performance evaluation metrics and business metrics of the target model version and the baseline model version are collected and compared in real time during the service process to obtain real-time metric comparison results. Traffic scheduling decisions are then executed based on the real-time metric comparison results. Dynamic and objective online verification replaces static and subjective manual judgment, thereby improving the accuracy of large model evaluation and decision-making. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating an evaluation and decision-making method for large models in one embodiment of the present invention; Figure 2 yes Figure 1 A schematic diagram of a specific implementation of step S110; Figure 3 yes Figure 1 A schematic diagram of a specific implementation of step S120; Figure 4 yes Figure 1 A schematic diagram of a specific implementation of step S130; Figure 5 This is a schematic block diagram of an evaluation and decision-making device for large models in one embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 7 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] The evaluation and decision-making method for large models provided in this invention can be applied to either the client or server. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. Currently, in the fintech and healthcare fields, the accuracy of existing large model evaluation and decision-making methods is relatively low. To address the above problems, this invention proposes an evaluation and decision-making method for large models. Before the model goes live, this method performs a quality gate check on the target model version based on the generated evaluation report to ensure the basic quality of the model. Then, it collects and compares the performance evaluation indicators and business indicators of the target model version and the baseline model version in real time during the service process to obtain real-time indicator comparison results. Based on the real-time indicator comparison results, it executes traffic scheduling decisions, replacing static and subjective manual judgment with dynamic and objective online verification, thereby improving the accuracy of large model evaluation and decision-making. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 1 As shown, Figure 1 A flowchart illustrating an evaluation and decision-making method for large models provided in an embodiment of the present invention includes the following steps: S110-S140.
[0015] S110. In response to model update events of large models, execute predefined evaluation tasks in an isolated environment to generate evaluation reports.
[0016] Specifically, such as Figure 2As shown, step S110 includes steps S111-S112: S111, when a submission event of the model code or model weights of the large model is detected, the predefined evaluation task is executed in the isolated environment; S112, the evaluation report is generated based on the multidimensional risk results, verification results, and inference results obtained from executing the predefined evaluation task. It should be noted that the predefined evaluation task includes a multidimensional risk task, a verification task, and an inference task; when a model update event related to the model code or model weights of the large model is detected, the predefined evaluation task pipeline is automatically triggered and executed in the isolated environment. This evaluation task system is designed for highly regulated industries and specifically includes: a multidimensional risk task. In the fintech field, this is achieved by running the task on a dedicated test set covering credit, market, operational risk, and compliance ethics to obtain a performance evaluation index vector and thus a multidimensional risk result. In the healthcare field, this is achieved by using a test set covering clinical effectiveness, safety, and ethical compliance to obtain a performance evaluation index vector, thereby quantifying the overall health of the model in the corresponding scenario, i.e., obtaining the multidimensional risk result. The validation task utilizes adversarial test sets (such as malicious language in the financial field and noisy images in the medical field) to evaluate the model's stability under anomalous inputs and obtain validation results. The inference task performs risk-free inference on sampled real business traffic (such as financial transaction requests in the financial field or anonymous medical record data in the medical field) in a shadow environment to verify the model's behavior under real data distribution and obtain inference results. The above multi-dimensional risk results, validation results, and inference results are automatically aggregated and analyzed to generate a unified structured evaluation report.
[0017] It should be noted that, in this embodiment, the isolated environment refers to an independent and controlled computing resource and network space dedicated to performing model evaluation tasks. It ensures complete isolation between the evaluation and production environments, avoiding interference with online services, while also guaranteeing the security of test data, the reproducibility of the evaluation process, and that the evaluation results are not affected by external factors.
[0018] S120. Perform a quality access control check on the target model version corresponding to the model update event based on the evaluation report.
[0019] Specifically, such as Figure 3As shown, step S120 includes steps S121-S122: S121, checking whether the values of each performance evaluation indicator in the evaluation report have reached their corresponding preset indicator thresholds; S122, if the values of each performance evaluation indicator have reached their corresponding preset indicator thresholds, then determining that the target model version corresponding to the model update event has passed the quality gate check. More specifically, firstly, the value of each performance evaluation indicator is parsed from the structured evaluation report. The performance evaluation indicators are defined according to the application field of the model. In the fintech scenario, they are mainly PEM (Performance, Ethics, and Metrics / Compliance) indicators covering performance, ethics, and compliance. In the healthcare scenario, they are CSR (Clinical, Safety, Regulatory) indicators covering clinical, safety, and regulations. Then, a preset indicator threshold rule base is accessed to obtain the threshold of each performance evaluation indicator that strictly corresponds to the current evaluation task and model scenario. Subsequently, the value of each parsed performance evaluation indicator is compared with its corresponding preset indicator threshold. A "pass" result is automatically generated only when all performance evaluation metrics meet or exceed their respective preset threshold requirements, allowing the target model version to proceed to the subsequent deployment process. If any one or more performance evaluation metrics fail to meet the preset threshold, it will be immediately judged as "fail," automatically terminating the subsequent process and triggering an alarm notification, providing feedback to the developers with the specific unmet performance evaluation metric and its value. Understandably, this fully automated access control mechanism constitutes an indispensable quality checkpoint before model deployment, ensuring the basic quality of the model.
[0020] S130. If the target model version passes the quality access control check, initial online traffic is allocated to the target model version, and performance evaluation indicators and business indicators of the target model version and the baseline model version are collected and compared in real time during the service process to obtain real-time indicator comparison results.
[0021] Specifically, once the target model version passes the quality gate check, the canary release process will be automatically triggered. First, a pre-configured small percentage of initial online traffic, such as 1% or 5%, is allocated to the target model version, while the vast majority of traffic is still directed to the stable baseline model version. Then, a real-time monitoring and data collection mechanism is initiated to continuously capture multi-dimensional data on both versions in a real-world service scenario, including performance evaluation metrics (such as response latency, error rate, and resource utilization) and business metrics (such as conversion rate, adoption rate, and manual review rate). This real-time data stream is fed into an aggregation and comparison analysis engine, which performs statistical calculations and significance tests based on a sliding time window, ultimately generating a dynamic real-time metric comparison result. This real-time metric comparison result clearly demonstrates the performance differences and statistical confidence levels of the target model version and the baseline model version across various dimensions, providing immediate and objective data for subsequent automated decision-making.
[0022] S140. Execute traffic scheduling decisions based on the real-time indicator comparison results.
[0023] Specifically, the traffic scheduling decision includes advance traffic decision, maintain traffic decision, and rollback decision, such as... Figure 4 As shown, step S140 includes steps S141-S143: S141. If the real-time indicator comparison result shows that the business indicators of the target model version show a statistically significant positive change compared with the business indicators of the baseline model version, and its performance evaluation indicators do not show a statistically significant deterioration, then the traffic advancement decision is executed to gradually increase the online traffic ratio of the target model version. S142. If the real-time indicator comparison result shows that the business indicators and performance evaluation indicators of the target model version and the baseline model version do not show statistically significant differences, then the traffic maintenance decision is executed to maintain the current traffic allocation ratio and continue monitoring. S143. If the real-time indicator comparison result shows that the absolute value of the business indicator of the target model version is lower than the preset indicator threshold, or the business indicator shows a statistically significant negative change, then the rollback decision is executed to trigger the traffic rollback operation.
[0024] Specifically, the traffic-driven decision-making process is triggered when real-time metric comparisons show that the target model version exhibits statistically significant positive changes in its business metrics compared to the baseline version within a complete observation period (e.g., a click-through rate increase of over 2% in financial marketing scenarios with a p-value (the probability that the observed difference is entirely due to random error in hypothesis testing) less than 0.05, or a significant increase in clinical recommendation adoption rate in medical auxiliary diagnosis scenarios). Simultaneously, all performance evaluation metrics (including but not limited to response latency, error rate, and PEM compliance score in financial scenarios and CSR security score in medical scenarios) show no statistically significant degradation, an automatic advancement decision will be triggered. This decision will automatically and gradually increase the online traffic proportion of the target model version according to a pre-configured incremental strategy, for example, from 5% to 20%, then 50%. After each increase, a new observation period will begin, continuously monitoring metric changes to ensure the promotion process is robust and controllable.
[0025] Maintaining Traffic Decision: When real-time metric comparisons show no statistically significant differences between the target model version and the baseline model version in business metrics (such as user satisfaction and transaction success rate) and all performance evaluation metrics, a maintain decision will be automatically executed. This indicates that the target model version, within the current observation window and traffic scale, has neither demonstrated clear improvement benefits nor shown unacceptable risks. At this point, the current traffic allocation ratio will remain unchanged, and monitoring will continue. This state may mean a longer observation period and a larger traffic sample are needed to obtain statistical confidence, or it may indicate that this iteration is a non-breakthrough optimization. Data will continue to be collected until conditions are triggered to advance or roll back the traffic decision.
[0026] Rollback Decision: This decision serves as a risk circuit breaker and automatic repair mechanism. It is triggered when either of the following conditions is met: First, the absolute value of any business metric or performance evaluation metric is found to be below a preset threshold (e.g., a financial model compliance score below 95%, or a medical model sensitivity score below the safety red line); Second, statistical analysis confirms a statistically significant negative change in the business metric (e.g., a significant increase in the complaint rate or an increase in the failure rate of critical tasks). Once the conditions are met, a rollback decision will be automatically triggered immediately. Within seconds or minutes, the traffic carried by the target model version will be smoothly switched back to the baseline stable version, either entirely or partially according to a preset strategy, and a high-level alert will be issued simultaneously. This design aims to minimize the scope and duration of the failure's impact, ensure the overall stability and security of online services, and provide clear time boundaries and data snapshots for subsequent problem localization and iterative repair.
[0027] It should be noted that the steps for triggering the traffic rollback operation include: performing a partial rollback operation on online traffic based on at least one dimension, such as user profile, transaction type, or service scenario. For example, when an abnormal response is detected in the target model version in the financial advisory scenario for high-net-worth clients, only requests with "high user risk level" and "transaction type financial advisory" can be rolled back, switching their traffic back to the baseline model version. Traffic for other users or business scenarios can continue to be served by the target model version, thereby achieving precise isolation of the fault impact and minimizing the business impact.
[0028] It should also be noted that after step S140, the process further includes: acquiring online monitoring data after the target model version is deployed and running; and performing correlation analysis between the evaluation report and the online monitoring data to obtain analysis results. Specifically, after the target model is deployed and running, online monitoring data of the target model version in the production environment is continuously acquired. This data covers real-time user queries, model responses, performance metrics (such as latency and error rate), and business result feedback. Subsequently, the structured evaluation report before this deployment is deeply correlated and compared with the online monitoring data. Through techniques such as pattern matching and difference attribution, the aim is to identify defects not covered in the evaluation phase, reveal differences between production data distribution and test data, or locate specific scenarios and root causes leading to performance degradation, ultimately generating analysis results containing clear diagnostic conclusions and improvement suggestions.
[0029] To facilitate understanding, the evaluation and decision-making method for large models in this invention will be explained using examples of large-scale model optimization in the fintech field of intelligent investment advisory and iterative AI-assisted diagnostic models in the healthcare field: A securities firm optimized its robo-advisor model to improve the compliance of investment recommendations and user adoption rates. Step S1: The development team submitted the V2.0 code, triggering the evaluation pipeline. Step S2: The evaluation tasks were automatically executed in an isolated environment: 1) Running on a multi-dimensional financial risk test set containing historical transactions and compliance cases to evaluate its PEM metrics (such as the accuracy of portfolio risk level matching and the compliance score of recommendations); 2) Verifying its robustness on an adversarial test set (such as leading and ambiguous user questions); 3) Performing inference on sampled real user queries in a shadow environment. Step S3: A quality gate check ensured all PEM metrics (such as compliance ≥ 98%) met the standards, and V2.0 was approved for gray-scale testing. Step S4: 1% of online user traffic was allocated to V2.0, running it in parallel with V1.0, and the business metrics (such as "recommendation adoption rate") and performance evaluation metrics (such as response latency) of both groups were monitored in real time. Step S5: After 24 hours, the real-time indicator comparison results show that the adoption rate of suggestions in the V2.0 group has significantly improved (p<0.05), and the PEM indicator is stable. The automatic decision is made to increase its traffic proportion to 5%. Step S6: If subsequent monitoring finds that the compliance score of V2.0's suggestions for a certain type of high-risk derivatives has suddenly dropped below the threshold, it will immediately trigger an automatic rollback to V1.0, and correlate the analysis and evaluation report to locate the specific financial product terms not covered by the test set.
[0030] A hospital deployed an AI-assisted diagnostic model for pulmonary nodules using CT images and planned to launch version 2.0, which improved the detection rate of small nodules. Step S1: After the model update is submitted, the evaluation pipeline is automatically triggered. Step S2: Execute the following in an isolated environment: 1) Run on a clinical test set containing various nodule types and rare cases to calculate its CSR metrics (such as sensitivity and specificity); 2) Verify its robustness on an adversarial test set (such as images with added noise and simulated motion artifacts); 3) Perform risk-free inference on desensitized historical images in a shadow environment. Step S3: Quality gate checks confirm that key CSR metrics (such as sensitivity > 97%) meet the standards, and V2.0 passes. Step S4: Deploy V2.0 online and allocate 1% of daily CT image reading traffic for gray-scale testing. Simultaneously, compare the results with the gold standard (expert panel) and V1.0 in real time, monitoring business metrics (such as the rate of doctors accepting AI suggestions) and performance metrics (such as the false negative rate). Step S5: Real-time data shows that the adoption rate of V2.0 has not changed significantly, but the false negative rate for ground-glass nodules shows an upward trend (p<0.05). An automatic decision is made to pause the increased throughput, maintain the 1% rate, and issue an alert. Step S6: After review, a missed diagnosis is confirmed, automatically triggering a full rollback. Correlation analysis reveals that the root cause is insufficient coverage of "atypical ground-glass nodules" in the evaluation set, thus generating a test set enhancement task.
[0031] The evaluation and decision-making method for large models in this invention performs a quality gate check on the target model version based on the generated evaluation report before the model goes live, ensuring the basic quality of the model. Then, it collects and compares the performance evaluation indicators and business indicators of the target model version and the baseline model version in real time during the service process to obtain real-time indicator comparison results. Based on the real-time indicator comparison results, it executes traffic scheduling decisions, replacing static and subjective manual judgment with dynamic and objective online verification, thereby improving the accuracy of large model evaluation and decision-making.
[0032] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0033] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.
[0034] In one embodiment, a large-model-oriented evaluation and decision-making apparatus 200 is provided, which corresponds one-to-one with the large-model-oriented evaluation and decision-making methods described in the above embodiments. For example... Figure 5 As shown, the evaluation and decision-making device 200 for large models includes an execution generation unit 201, an inspection unit 202, an allocation and comparison unit 203, and an execution unit 204. Detailed descriptions of each functional module are as follows: The execution generation unit 201 is used to execute predefined evaluation tasks and generate evaluation reports in an isolated environment in response to model update events of large models. Inspection unit 202 is used to perform quality access control checks on the target model version corresponding to the model update event based on the evaluation report; The allocation comparison unit 203 is used to allocate initial online traffic to the target model version if the target model version passes the quality access control check, and to collect and compare the performance evaluation indicators and business indicators of the target model version and the baseline model version in real time during the service process to obtain real-time indicator comparison results. The execution unit 204 is used to make traffic scheduling decisions based on the real-time indicator comparison results.
[0035] In one embodiment, the generation unit 201 is specifically used for: When a submission event of the model code or model weights of the large model is detected, the predefined evaluation task is executed in the isolated environment, wherein the predefined evaluation task includes a multidimensional risk task, a verification task, and an inference task. The assessment report is generated based on the multidimensional risk results, verification results, and reasoning results obtained from performing the predefined assessment task.
[0036] In one embodiment, the inspection unit 202 is specifically used for: Check whether the values of each performance evaluation indicator in the evaluation report have reached their corresponding preset indicator thresholds; If the values of all the performance evaluation indicators reach their corresponding preset indicator thresholds, then the target model version corresponding to the model update event is determined to have passed the quality gate check.
[0037] In one embodiment, the execution unit 204 is specifically used for: If the real-time indicator comparison result shows that the business indicators of the target model version show a statistically significant positive change compared to the business indicators of the baseline model version, and its performance evaluation indicators do not show a statistically significant deterioration, then the traffic advancement decision is executed to gradually increase the online traffic ratio of the target model version. If the real-time metric comparison results show that there are no statistically significant differences between the business metrics and performance evaluation metrics of the target model version and the baseline model version, then the traffic maintenance decision is executed to maintain the current traffic allocation ratio and monitoring continues. If the real-time indicator comparison result shows that the absolute value of the business indicator of the target model version is lower than the preset indicator threshold, or the business indicator shows a statistically significant negative change, then the rollback decision is executed to trigger a traffic rollback operation.
[0038] In one embodiment, the evaluation and decision-making apparatus 200 for large models further includes: The acquisition unit is used to acquire online monitoring data after the target model version is deployed and run; The analysis unit is used to perform correlation analysis between the evaluation report and the online monitoring data to obtain analysis results.
[0039] The evaluation and decision-making device for large models in this invention performs a quality gate check on the target model version based on the generated evaluation report before the model goes online, ensuring the basic quality of the model. Then, it collects and compares the performance evaluation indicators and business indicators of the target model version and the baseline model version in real time during the service process to obtain real-time indicator comparison results. Based on the real-time indicator comparison results, it executes traffic scheduling decisions, replacing static and subjective manual judgment with dynamic and objective online verification, thereby improving the accuracy of large model evaluation and decision-making.
[0040] Specific limitations regarding the evaluation and decision-making apparatus for large models can be found in the limitations of the evaluation and decision-making methods for large models described above, and will not be repeated here. Each unit in the aforementioned evaluation and decision-making apparatus for large models can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0041] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps of an evaluation and decision-making method for large models.
[0042] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of a large-model-oriented evaluation and decision-making method.
[0043] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described evaluation decision method for large models.
[0044] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described evaluation and decision-making method for large models.
[0045] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0046] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0047] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0048] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for evaluation and decision-making for large models, characterized in that, include: In response to model update events of large models, predefined evaluation tasks are executed in an isolated environment to generate evaluation reports; Based on the evaluation report, a quality access control check is performed on the target model version corresponding to the model update event; If the target model version passes the quality access control check, initial online traffic is allocated to the target model version, and performance evaluation indicators and business indicators of the target model version and the baseline model version are collected and compared in real time during the service process to obtain real-time indicator comparison results. Traffic scheduling decisions are made based on the real-time indicator comparison results.
2. The evaluation and decision-making method for large models as described in claim 1, characterized in that, The steps of generating an evaluation report by performing a predefined evaluation task in an isolated environment in response to a model update event of a large model include: Upon detecting a submission event of the model code or model weights of the large model, the predefined evaluation task is executed in the isolated environment; The assessment report is generated based on the multidimensional risk results, verification results, and reasoning results obtained from performing the predefined assessment task.
3. The evaluation and decision-making method for large models as described in claim 1, characterized in that, The step of performing a quality gate check on the target model version corresponding to the model update event based on the evaluation report includes: Check whether the values of each performance evaluation indicator in the evaluation report have reached their corresponding preset indicator thresholds; If the values of all the performance evaluation indicators reach their corresponding preset indicator thresholds, then the target model version corresponding to the model update event is determined to have passed the quality gate check.
4. The evaluation and decision-making method for large models as described in claim 1, characterized in that, The traffic scheduling decision includes advancing traffic decisions; the step of executing traffic scheduling decisions based on the real-time indicator comparison results includes: If the real-time metric comparison result shows that the business metrics of the target model version show a statistically significant positive change compared to the business metrics of the baseline model version, and its performance evaluation metrics do not show a statistically significant deterioration, then the traffic advancement decision is executed to gradually increase the online traffic ratio of the target model version.
5. The evaluation and decision-making method for large models as described in claim 4, characterized in that, The traffic scheduling decision also includes a traffic maintenance decision; the step of executing the traffic scheduling decision based on the real-time indicator comparison result further includes: If the real-time metric comparison results show that neither the business metrics nor the performance evaluation metrics of the target model version nor the baseline model version show statistically significant differences, then the traffic maintenance decision is executed to maintain the current traffic allocation ratio and monitoring continues.
6. The evaluation and decision-making method for large models as described in claim 4, characterized in that, The traffic scheduling decision also includes a rollback decision; the step of executing the traffic scheduling decision based on the real-time indicator comparison result further includes: If the real-time indicator comparison result shows that the absolute value of the business indicator of the target model version is lower than the preset indicator threshold, or the business indicator shows a statistically significant negative change, then the rollback decision is executed to trigger a traffic rollback operation.
7. The evaluation and decision-making method for large models as described in claim 6, characterized in that, After the step of executing the rollback decision to trigger the traffic rollback operation, the method further includes: Obtain online monitoring data after the target model version is deployed and running; The analysis results are obtained by correlating the assessment report with the online monitoring data.
8. An evaluation and decision-making device for large models, characterized in that, include: The execution generation unit is used to respond to model update events of large models and perform predefined evaluation tasks to generate evaluation reports in an isolated environment; The inspection unit is used to perform a quality access control check on the target model version corresponding to the model update event based on the evaluation report. The allocation comparison unit is used to allocate initial online traffic to the target model version if the target model version passes the quality access control check, and to collect and compare the performance evaluation indicators and business indicators of the target model version and the baseline model version in real time during the service process to obtain real-time indicator comparison results. The execution unit is used to make traffic scheduling decisions based on the real-time indicator comparison results.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the evaluation and decision-making method for large models as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the evaluation and decision-making method for large models as described in any one of claims 1 to 7.