Automatic quantitative screening method and device for large models

By using an automated quantitative screening method, the accuracy parameter is calculated by comparing the processing results of the standard answer set and the candidate large model. This solves the problems of subjectivity and low efficiency in the selection of large models and realizes a scientific and efficient screening process.

CN121808324APending Publication Date: 2026-04-07CHINA ELECTRONICS CORP 6TH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, the selection of large models relies on human experience and judgment, lacking a systematic, repeatable, and statistically significant evaluation mechanism. This results in low screening efficiency, high subjectivity, poor reliability, and poor performance of the selected large models.

Method used

This paper presents an automated quantitative screening method for large models. By obtaining the standard answer set of the target task, calling the candidate large models for processing and comparing the processing results with the standard answer set, calculating the accuracy parameter, and screening out the target large models suitable for processing the target task.

Benefits of technology

It achieves scientific rigor and objectivity in large-scale model screening, reduces labor costs, and improves screening efficiency, especially for the automatic screening process of batch tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808324A_ABST
    Figure CN121808324A_ABST
Patent Text Reader

Abstract

The invention provides a large model automatic quantitative screening method and device. The method comprises the steps of obtaining a standard answer set of a target task; wherein the target task is a batch task, the batch task comprises a plurality of task items, and the standard answer set comprises a standard answer corresponding to each task item; respectively calling each candidate large model in the candidate large model set to process the target task, and obtaining a processing result set of each candidate large model; wherein the processing result set comprises a processing result corresponding to each task item; comparing the processing result set of each candidate large model with a standard answer set to obtain a correct rate parameter of each candidate large model; and comparing the accuracy parameter of each candidate large model, and screening out a target large model suitable for processing the target task from the candidate large model set. Therefore, the target large model suitable for processing the target task can be quantitatively evaluated and automatically screened out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large model technology, and in particular to an automated quantitative screening method and apparatus for large models. Background Technology

[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have demonstrated powerful capabilities in fields such as natural language processing, intelligent question answering, and text generation. Currently, various general-purpose large models are widely used in various industry scenarios.

[0003] Differences in architecture design, training data, and fine-tuning strategies among various large-scale models lead to significant variations in their performance on specific tasks. Therefore, when dealing with a particular task, a crucial issue is how to scientifically and rationally select the most suitable large-scale model from the many available models.

[0004] However, in existing technologies, the selection of large models mostly relies on users' human experience judgment or subjective perception evaluation, lacking a systematic, repeatable and statistically significant evaluation mechanism. This leads to problems such as low efficiency, strong subjectivity, poor reliability, and poor performance of the selected large models in the large model selection process. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method and apparatus for automated quantitative screening of large models, which can quantitatively evaluate and automatically screen out target large models suitable for handling target tasks.

[0006] This application provides an automated quantitative screening method for large models, the method comprising: Obtain the standard answer set for the target task; wherein the target task is a batch task, the batch task includes multiple task items, and the standard answer set includes the standard answer corresponding to each task item; The target task is processed by calling each candidate large model in the candidate large model set respectively, and the processing result set of each candidate large model is obtained; wherein, the processing result set includes the processing result corresponding to each task item; The processing result set of each candidate large model is compared with the standard answer set to obtain the accuracy parameter of each candidate large model; Compare the accuracy parameters of each candidate large model, and select the target large model suitable for processing the target task from the candidate large model set.

[0007] This application embodiment also provides a large-scale automated quantitative screening device, the device comprising: An acquisition module is used to acquire a set of standard answers for a target task; wherein the target task is a batch task, the batch task includes multiple task items, and the set of standard answers includes the standard answer corresponding to each task item; The processing module is used to call each candidate large model in the candidate large model set to process the target task, and obtain the processing result set of each candidate large model; wherein, the processing result set includes the processing result corresponding to each task item; The determination module is used to compare the processing result set of each candidate large model with the standard answer set to obtain the accuracy parameter of each candidate large model; The filtering module is used to compare the accuracy parameters of each candidate large model and filter out the target large model suitable for processing the target task from the set of candidate large models.

[0008] This application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the large model automated quantitative screening method described above are performed.

[0009] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the large model automated quantitative screening method described above.

[0010] This application provides an automated quantitative screening method and apparatus for large-scale models. By comparing the processing result set of each candidate large-scale model for a target task with a standard answer set, the accuracy parameter of each candidate large-scale model is obtained. Then, by comparing the accuracy parameters of each candidate large-scale model, a suitable target large-scale model for processing the target task is selected. This allows for quantitative evaluation and automatic screening of suitable target large-scale models for processing the target task. Quantitative evaluation ensures the scientific rigor and objectivity of the large-scale model screening process. Especially for batch tasks, the automated screening process reduces labor costs and improves screening efficiency.

[0011] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart of an automated quantitative screening method for large models provided in an embodiment of this application is shown; Figure 2 This illustration shows a structural schematic diagram of a large-scale automated quantitative screening device provided in an embodiment of this application; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.

[0015] Research has shown that with the rapid development of artificial intelligence technology, large language models (LLMs) have demonstrated powerful capabilities in fields such as natural language processing, intelligent question answering, and text generation. Currently, various general-purpose large models are widely used in various industry scenarios.

[0016] Differences in architectural design, training data, and fine-tuning strategies among various large-scale models lead to significant variations in their performance on specific tasks. Especially in the Chinese context, domestically developed models typically have advantages in understanding Chinese semantics, following instruction structures, and adapting to local business rules; however, these advantages are difficult to quantify accurately through simple comparative tests. Therefore, when dealing with specific tasks, a crucial issue is how to scientifically and rationally select the most suitable large-scale model from among the many available models.

[0017] However, in existing technologies, the selection of large models mostly relies on users' human experience judgment or subjective perception evaluation, lacking a systematic, repeatable and statistically significant evaluation mechanism. This leads to problems such as low efficiency, strong subjectivity, poor reliability, and poor performance of the selected large models in the large model selection process.

[0018] Based on this, embodiments of this application provide an automated quantitative screening method for large models to quantitatively evaluate and automatically screen out target large models suitable for processing target tasks.

[0019] Please see Figure 1 , Figure 1 This is a flowchart illustrating an automated quantitative screening method for large models provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the method includes: S101. Obtain the standard answer set for the target task.

[0020] The target task is a batch task, which includes multiple task items, meaning that multiple task items need to be processed one by one in the same way. The standard answer set includes the standard answer corresponding to each task item.

[0021] In one example, the objective task is to record the test results according to the evaluation indicators of each item in the preset inspection report, and to fill in the inspection conclusion in the preset inspection report based on the test results.

[0022] The inspection reports require a significant amount of repetitive processing, with approximately 700 reports per year, each 500 pages long. These reports contain various issues requiring manual review, which is time-consuming and labor-intensive. One type of question involves multiple-choice questions, with each report averaging 3000 questions (3000 items). After completion, the answers must be manually verified, and errors corrected to ensure the accuracy meets the standards. Since manual processing is too time-consuming and labor-intensive, a large-scale model is used for processing. Table 1 below provides an example of an inspection report.

[0023] Table 1 Inspection Report

[0024] As shown in Table 1 above, the objective is to record whether the test results S1 meet the requirements according to the evaluation index S2 of each item in the pre-set inspection report, take the degree of compliance as the test result, and fill the test result into the test conclusion S3 in the pre-set inspection report. The options for the degree of compliance may include compliant, partially compliant, non-compliant, and inapplicable.

[0025] In this step, the correct set of standard answers can be obtained from historical inspection reports, or a portion of the current inspection report can be manually filled out to obtain the correct set of standard answers. This set of standard answers serves as a sample for selecting a suitable target model to handle the target task.

[0026] S102. Call each candidate large model in the candidate large model set to process the target task, and obtain the processing result set of each candidate large model.

[0027] In this step, the candidate large model set contains multiple existing large models available for use. Each candidate large model is called to process the target task, and the processing result set output by each candidate large model is obtained. The processing result set includes the processing result corresponding to each task item. Corresponding to the example above, the processing result set includes the test conclusion for each item.

[0028] Corresponding to the example above, step S102 may include: S1021. Construct questioning scripts based on the evaluation indicators and result records in the preset inspection report.

[0029] S1022. Submit the questioning script to each candidate big model. Each candidate big model, for each item in the preset inspection report, understands the requirements of the evaluation indicators and the recorded content of the result record according to the prompting script. By comparing the requirements of the evaluation indicators and the recorded content of the result record, it determines the option that should be filled in the inspection conclusion from multiple options, forming the processing result set of each candidate big model.

[0030] Here, you can fill in the evaluation indicators and results records into the script template to form the question script; for example, the question script is: Help me determine whether the work done in statement 2 meets the requirements of statement 1? There are only four possible answers: meets, partially meets, does not meet, and is not applicable. Just choose one of these answers for me. Statement 1: [S1]. Statement 2: [S2].

[0031] Then, the script is submitted to the large model, which is processed by calling the large model via API to obtain the processing result.

[0032] In one example, the processing procedure for a large model is as follows: Answer format: <thinking>*** <thinking / > conform to <thinking>Okay, let's analyze it step by step.

[0033] 1. Understand the requirements of statement 1. Statement 1 states: The site should be selected in a building with earthquake resistance, wind resistance, and rain protection capabilities.

[0034] This is a site selection or site condition requirement, emphasizing that the building itself possesses these three protective capabilities.

[0035] 2. Understand the content checked in statement 2. Statement 2 provides the following information: The computer room acceptance documents include descriptions of its earthquake resistance, wind resistance, and rain resistance capabilities (with written proof).

[0036] On-site inspection: The roof was not leaking (indicating that the rainproof capability was effective at the time of inspection).

[0037] There are no windows facing outwards (which reduces the possibility of rain and wind entering).

[0038] The accumulation of soil due to wind due to the absence of doors and windows (indicating effective wind protection).

[0039] The walls have no minor damage (which may indirectly indicate that the building structure is intact and is related to its wind and earthquake resistance).

[0040] 3. Comparison and judgment Rainproof: Documentation explains + no leaks on site → Meets requirements.

[0041] Windproof: Documentation provides instructions + No soil accumulation + No windows or windows are properly closed → Meets requirements.

[0042] Earthquake resistance: There are only documented descriptions, without details of on-site structural inspections or seismic assessments. However, if the document is an acceptance document, it can usually prove that the building was constructed according to earthquake-resistant design.

[0043] Statement 1 requires that "the site should be selected in a building with the capability to..." During the verification, based on the acceptance documents (proving the capability) and the absence of any abnormalities on site (indirectly proving the continued validity of the capability), it can be considered that the overall requirements are met.

[0044] Since there is no obvious missing or refutation of evidence, it does not fall under the categories of "partially conforming" or "not conforming".

[0045] 4. Conclusion Documented evidence + no opposing evidence on site → meets the requirements of "having earthquake resistance, wind resistance, and rain resistance".

[0046] <thinking / > Output answer: Match In this manner, each candidate large model processes each task item one by one, summarizes the test results, and forms the processing result set of each candidate large model.

[0047] S103. Compare the processing result set of each candidate large model with the standard answer set to obtain the accuracy parameter of each candidate large model.

[0048] In this step, the processing result set of the candidate large model can be compared with the standard answer set one by one to obtain the accuracy parameters of each candidate large model, such as the accuracy of each item; it can also include accuracy statistics, such as the average accuracy and the standard deviation of accuracy.

[0049] In one possible implementation, step S103 may include: S1031. For each candidate large model, compare the processing result set of the candidate large model with the standard answer set to form the processing statistical sequence of the candidate large model.

[0050] In this process, each element in the statistical sequence corresponds one-to-one with each task item of the target task, and the value of each element is used to characterize whether the candidate large model's processing result for the corresponding task item is correct; for example, 0 indicates incorrect and 1 indicates correct.

[0051] S1032. A sliding window of a preset length is slid across the processing statistical sequence. Each time the sliding window slides once, the accuracy of the candidate large model within the sliding window is determined.

[0052] Suppose a test report has S items, for example, S=3000. We extract the first segment of S, with a length of P items, and calculate the accuracy rate of the large model processing within it as X1. Generally, P=100, so the first segment includes items 1 to 100 in S. Next, we move Q items, extract the second segment, keeping the length the same, and calculate the accuracy rate as X2. It's recommended to use Q=P / 2=50, so the second segment includes items 51 to 150 in S, and so on, continuing to obtain accuracy rates X3 to Xn. Reaching the last item in S, we can return to the first item, meaning the 3001st item is the first item.

[0053] The parameter Q can be used to adjust the number of segments n extracted, as well as the correlation between the accuracy of adjacent segments. For example, if the number of items S is small, decreasing Q ensures that a larger number of segments n are extracted, which is beneficial for calculating the average and standard deviation. If the number of items S is extremely large, increasing Q, even to a value much larger than P, ensures that a relatively small number of segments n are extracted, making the number of segments just right for statistical analysis and saving the overall running time of the large model. When the number of items S is appropriate, increasing Q to a value close to P minimizes the impact of the accuracy of adjacent segments on each other.

[0054] Thus, the embodiments of this application employ a local performance sampling method based on a sliding window, which can detect performance fluctuations of a large model across different data segments, avoid masking local fluctuations with a single global accuracy, and improve evaluation granularity.

[0055] S1033. Based on the determined accuracy rates, calculate the average and standard deviation of the accuracy rates of the candidate large model to obtain the accuracy parameters of the candidate large model.

[0056] In this step, for each candidate large model, based on the accuracy rates determined for that candidate large model, the mean and standard deviation of its accuracy are calculated. The formula is expressed as:

[0057]

[0058] in, The average value representing the accuracy rate; The standard deviation represents the accuracy rate; This indicates the number of samples for which the accuracy is achieved. This reflects the overall accuracy of the large model processing report. The bigger the better. This reflects the stability of the accuracy of the large model's report processing. The smaller the better. , These are the two most important indicators of a model's processing power.

[0059] In addition, the processing time of each item by the candidate large model can be obtained, and the average U and standard deviation V of the processing time for each item can be determined in a similar manner. U reflects the speed at which the large model processes reports, and a smaller U is better. V reflects the stability of the large model's processing speed, and a smaller V is better. The importance of the U indicator depends mainly on the urgency of the processing task's time requirements.

[0060] S104. Compare the accuracy parameters of each candidate large model, and select the target large model suitable for processing the target task from the candidate large model set.

[0061] In this step, after the large model has finished processing the same target task, the processing effects of the two models can be compared based on the accuracy parameter to select the target large model suitable for processing the target task. In this embodiment, the target task is a batch repetitive task, so when the same type of task is generated again, the target large model can be directly called for processing to achieve better processing results.

[0062] In one possible implementation, step S104 may include: S1041. Select two candidate large models from the candidate large model set, compare the processing effects of the two selected candidate large models according to their accuracy parameters, and store each selected candidate large model into the model queue in order according to the processing effect.

[0063] S1042. Select a new candidate large model from the candidate large model set. For the candidate large model, select multiple candidate large models from the model queue and combine them with the candidate large model to form multiple model pairs.

[0064] S1043. For each model pair, compare the processing performance of the two candidate large models based on their accuracy parameters.

[0065] S1044. Based on the processing effect of the two candidate large models in each model pair, determine the position of the candidate large model and store the candidate large model in the model queue according to its position.

[0066] S1045. Repeatedly select a candidate large model from the candidate large model set and store it in the model queue according to its position until all candidate large models in the candidate large model set are stored in the model queue.

[0067] S1046. Select at least one candidate large model with the best processing effect from the model queue as the target large model.

[0068] In one example, two candidate large models are randomly selected from the candidate large model set and their processing effects are compared based on the accuracy parameter. Then, they are stored in the model queue in order of processing effect from best to worst.

[0069] Next, a third candidate large model is selected from the candidate large model set and compared with the two candidate large models in the model queue one by one. The position of the third candidate large model is determined by sorting according to the processing effect and stored in the model queue according to the position.

[0070] Next, each other large model in the candidate large model set is compared with the processing performance of the large models in the model queue in turn. Finally, the first model in the model queue is the best large model for the current target task. Alternatively, the top N large models can be selected. For example, N is 3, which means 1 primary large model and 2 backup large models.

[0071] Furthermore, in S1041 and S1043, the processing performance of the two candidate large models is compared based on their accuracy parameters, including: Step 1: Determine candidate large models and candidate large models The difference between the average accuracy rate and the average accuracy rate.

[0072] Step 2: If the difference is greater than or equal to the preset threshold, the candidate large model with the larger average accuracy is determined to have a better processing effect than the candidate large model with the smaller average accuracy.

[0073] If the difference between the two average values ​​is large, that is In this case, the option with the highest accuracy rate is selected as the best. For example, =0.1.

[0074] Step 3: If the difference is less than a preset threshold, then the candidate large model is determined using the following formula. and candidate large models Reference values ​​for comparing the processing effects between them :

[0075] in, Representing candidate large models The average accuracy rate; Representing candidate large models The standard deviation of the accuracy rate; Representing candidate large models The average accuracy rate; Representing candidate large models The standard deviation of the accuracy rate; This represents the number of candidate large models with high accuracy.

[0076] In this step, if the difference between the two averages is small, based on statistical experience, the accuracy statistics have these characteristics: sample data size. Large enough, average , The difference may be small, standard deviation , The differences may be significant. Taking all factors into consideration, the above formula is used to calculate a comparative reference value for the processing effects of the two large models.

[0077] Step 4, when At that time, candidate large models are determined. The processing effect is better than that of the candidate large model. The processing effect is better; when At that time, candidate large models are determined. The processing effect is better than that of the candidate large model. The processing effect is better; when At that time, candidate large models are determined. The processing effect and candidate large model The processing results are the same; among them, These are empirical parameter values. Based on long-term experiments, a threshold value is generally set. Between 1.0 and 3.0, 2.0 can be taken in the embodiments of this application. = 2.0.

[0078] Alternatively, when this situation rarely occurs; if it does occur, it can be considered that the processing effects of the two large models are approximately the same, and they are considered to be on par or further compare other aspects, such as processing time. That is, according to the processing duration of the two candidate large models for the target task, the processing effects of the two candidate large models are compared.

[0079] Further, after S103, the large model automated quantitative screening method provided by the embodiments of this application further includes: Compare at least one of the number of task items actually processed by each candidate large model, the lowest correct rate within each sliding window, and the standard deviation of the correct rate with a preset qualified standard, and delete the candidate large models that do not meet the preset qualified standard from the candidate large model set.

[0080] Here, through the eligibility test, check whether the processing results of the candidate large models meet the basic requirements. Specifically, there are at least one of the following qualified standards: S0 index: It must satisfy S > S0, which reflects that the number of items in the processed report must be large enough to represent that the large model processes a large number of samples and the results are accurate. When S0 = 2000 is taken, S = 3000 previously, which is obviously met.

[0081] K0 index: It must satisfy Xmin > X * K0, and the value range of K0 is [0, 1], which reflects that the correct rate of the worst case in the large model processing report cannot be too low. Usually, K0 = 0.90 is taken, that is, Xmin must be at least 0.90 times higher than the average value at worst.

[0082] K1 index: It must satisfy Y < K1, which reflects that the fluctuation of the correct rate in the large model processing report cannot be too large. Usually, K1 = 0.05 is taken.

[0083] In addition, when there are time requirements, the U0 index can also be considered: It must satisfy U < U0, which reflects the processing speed of the large model. On average, one item must be processed within U0 time. Set a reasonable U0 according to the specific processing task. Corresponding to the above example, if it is required to process a report in 5 hours, which is equivalent to processing one item in an average of 6 seconds (5 * 60 * 60 / 3000), then U0 = 7 seconds can be set.

[0084] After that, the target large model suitable for processing the target task can be screened from the filtered candidate large model set by comparing the correct rate parameters of each candidate large model in the manner of S104.

[0085] This application provides an automated quantitative screening method for large models, comprising: obtaining a set of standard answers for a target task; wherein the target task is a batch task, the batch task includes multiple task items, and the set of standard answers includes the standard answers corresponding to each task item; processing the target task by calling each candidate large model in the candidate large model set respectively, and obtaining a processing result set for each candidate large model; wherein the processing result set includes the processing result corresponding to each task item; comparing the processing result set of each candidate large model with the set of standard answers to obtain the accuracy parameter of each candidate large model; comparing the accuracy parameters of each candidate large model, and selecting a target large model suitable for processing the target task from the candidate large model set.

[0086] In this way, on the one hand, it is possible to scientifically and rationally screen and rank large models from different companies, series, and models, before and after fine-tuning, based on their performance on specific tasks, thereby selecting the target large models suitable for handling the target tasks. Quantitative evaluation can ensure the scientific and objective nature of the large model selection. On the other hand, the screening method is easily programmed. Fully automated quantitative screening of large models can be achieved through programming. Apart from setting initial parameters, the screening process does not require manual intervention, saving labor costs. Especially for batch tasks, the automated screening process can reduce labor costs and improve screening efficiency.

[0087] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of a large-scale automated quantitative screening device provided in an embodiment of this application. Figure 2 As shown, the large-scale automated quantitative screening device 200 includes: The acquisition module 210 is used to acquire the standard answer set for the target task; wherein the target task is a batch task, the batch task includes multiple task items, and the standard answer set includes the standard answer corresponding to each task item; The processing module 220 is used to call each candidate large model in the candidate large model set to process the target task, and obtain the processing result set of each candidate large model; wherein, the processing result set includes the processing result corresponding to each task item; The determination module 230 is used to compare the processing result set of each candidate large model with the standard answer set to obtain the accuracy parameter of each candidate large model; The filtering module 240 is used to compare the accuracy parameters of each candidate large model and filter out the target large model suitable for processing the target task from the set of candidate large models.

[0088] Furthermore, when the determining module 230 compares the processing result set of each candidate large model with the standard answer set to obtain the accuracy parameter of each candidate large model, the determining module 230 is used to: For each candidate large model, the processing result set of the candidate large model is compared with the standard answer set to form a processing statistical sequence of the candidate large model; wherein, each element in the processing statistical sequence corresponds one-to-one with each task item of the target task, and the value of each element is used to characterize whether the processing result of the candidate large model on the corresponding task item is correct; A sliding window of a preset length is slid over the processing statistics sequence. Each time the sliding window slides once, the accuracy of the candidate large model within that sliding window is determined. Based on the determined accuracy rates, the mean and standard deviation of the accuracy rates of the candidate large model are calculated to obtain the accuracy parameters of the candidate large model.

[0089] Furthermore, when the filtering module 240 compares the accuracy parameters of each candidate large model and filters out the target large model suitable for processing the target task from the candidate large model set, the filtering module 240 is used to: Two candidate large models are selected from the candidate large model set. The processing effects of the two selected candidate large models are compared based on their accuracy parameters. Each selected candidate large model is then stored in the model queue in order based on its processing effect. From the set of candidate large models, a new candidate large model is selected. For the candidate large model, multiple candidate large models are selected from the model queue and combined with the candidate large model to form multiple model pairs. For each model pair, compare the processing performance of the two candidate large models based on their accuracy parameters. Based on the processing effect of the two candidate large models in each model pair, the position of the candidate large model is determined and the candidate large model is stored in the model queue according to its position; Repeatedly select a candidate large model from the candidate large model set and store it into the model queue according to its position until all candidate large models in the candidate large model set are stored into the model queue; Select at least one candidate large model with the best processing performance from the model queue as the target large model.

[0090] Furthermore, the large-scale automated quantitative screening device 200 also includes a filtering module; the filtering module is used for: The number of tasks actually processed by each candidate large model, the lowest accuracy in each sliding window, and the standard deviation of the accuracy are compared with a preset qualification standard. Candidate large models that do not meet the preset qualification standard are deleted from the candidate large model set.

[0091] Furthermore, when the filtering module 240 compares the processing performance of two candidate large models based on their accuracy parameters, the filtering module 240 is used to: Determine candidate large models and candidate large models The difference between the average accuracy rate and the average accuracy rate; If the difference is greater than or equal to the preset threshold, the candidate large model with the larger average accuracy is determined to be better than the candidate large model with the smaller average accuracy. If the difference is less than a preset threshold, the candidate large model is determined by the following formula. and candidate large models Reference values ​​for comparing the processing effects between them :

[0092] in, Representing candidate large models The average accuracy rate; Representing candidate large models The standard deviation of the accuracy rate; Representing candidate large models The average accuracy rate; Representing candidate large models The standard deviation of the accuracy rate; This represents the number of candidate large models with high accuracy. when At that time, candidate large models are determined. The processing effect is better than that of the candidate large model. The processing effect is better; when At that time, candidate large models are determined. The processing effect is better than that of the candidate large model. The processing effect is better; when At that time, candidate large models are determined. The processing effect and candidate large model The processing results are the same; among them, These are empirical parameter values.

[0093] Furthermore, when comparing the processing performance of two candidate large models based on their accuracy parameters, the filtering module 240 is also used to: when At that time, the processing time of the two candidate large models for processing the target task is compared to evaluate the processing performance of the two candidate large models.

[0094] Furthermore, the target task is to record the inspection results according to the evaluation indicators of each item in the preset inspection report, and fill in the inspection conclusion in the preset inspection report based on the inspection results; then, when the processing module 220 is used to call each candidate large model in the candidate large model set to process the target task and obtain the processing result set of each candidate large model, the processing module 220 is used to: Construct questioning scripts based on the evaluation indicators and result records in the preset inspection report; The questioning script is submitted to each candidate big model. Each candidate big model, for each item in the preset inspection report, understands the requirements of the evaluation indicators and the recorded content of the result record according to the prompting script. By comparing the requirements of the evaluation indicators and the recorded content of the result record, it determines the option that should be filled in the inspection conclusion from multiple options, forming the processing result set of each candidate big model.

[0095] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device 300 includes a processor 310, a memory 320, and a bus 330.

[0096] The memory 320 stores machine-readable instructions that can be executed by the processor 310. When the electronic device 300 is running, the processor 310 and the memory 320 communicate via the bus 330. When the machine-readable instructions are executed by the processor 310, the steps of the large model automated quantitative screening method as described in the above method embodiment can be performed. For specific implementation details, please refer to the method embodiment, which will not be repeated here.

[0097] This application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it can execute the steps of the large model automated quantitative screening method as described in the above method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0099] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0100] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0101] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0102] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0103] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.< / thinking> < / thinking>

Claims

1. A large-scale automated quantitative screening method, characterized in that, The method includes: Obtain the standard answer set for the target task; wherein the target task is a batch task, the batch task includes multiple task items, and the standard answer set includes the standard answer corresponding to each task item; The target task is processed by calling each candidate large model in the candidate large model set respectively, and the processing result set of each candidate large model is obtained; wherein, the processing result set includes the processing result corresponding to each task item; The processing result set of each candidate large model is compared with the standard answer set to obtain the accuracy parameter of each candidate large model; Compare the accuracy parameters of each candidate large model, and select the target large model suitable for processing the target task from the candidate large model set.

2. The method according to claim 1, characterized in that, The processing result set of each candidate large model is compared with the standard answer set to obtain the accuracy parameters of each candidate large model, including: For each candidate large model, the processing result set of the candidate large model is compared with the standard answer set to form a processing statistical sequence of the candidate large model; wherein, each element in the processing statistical sequence corresponds one-to-one with each task item of the target task, and the value of each element is used to characterize whether the processing result of the candidate large model on the corresponding task item is correct; A sliding window of a preset length is slid over the processing statistics sequence. Each time the sliding window slides once, the accuracy of the candidate large model within that sliding window is determined. Based on the determined accuracy rates, the mean and standard deviation of the accuracy rates of the candidate large model are calculated to obtain the accuracy parameters of the candidate large model.

3. The method according to claim 2, characterized in that, Compare the accuracy parameters of each candidate large model, and select a target large model suitable for processing the target task from the candidate large model set, including: Two candidate large models are selected from the candidate large model set. The processing effects of the two selected candidate large models are compared based on their accuracy parameters. Each selected candidate large model is then stored in the model queue in order based on its processing effect. From the set of candidate large models, a new candidate large model is selected. For the candidate large model, multiple candidate large models are selected from the model queue and combined with the candidate large model to form multiple model pairs. For each model pair, compare the processing performance of the two candidate large models based on their accuracy parameters. Based on the processing effect of the two candidate large models in each model pair, the position of the candidate large model is determined and the candidate large model is stored in the model queue according to its position; Repeatedly select a candidate large model from the candidate large model set and store it into the model queue according to its position until all candidate large models in the candidate large model set are stored into the model queue; Select at least one candidate large model with the best processing performance from the model queue as the target large model.

4. The method according to claim 3, characterized in that, After comparing the processing result set of each candidate large model with the standard answer set to obtain the accuracy parameter of each candidate large model, the method further includes: The number of tasks actually processed by each candidate large model, the lowest accuracy in each sliding window, and the standard deviation of the accuracy are compared with a preset qualification standard. Candidate large models that do not meet the preset qualification standard are deleted from the candidate large model set.

5. The method according to claim 3, characterized in that, Based on the accuracy parameters of the two candidate large models, compare their processing performance, including: Determine candidate large models and candidate large models The difference between the average accuracy rate and the average accuracy rate; If the difference is greater than or equal to the preset threshold, the candidate large model with the larger average accuracy is determined to be better than the candidate large model with the smaller average accuracy. If the difference is less than a preset threshold, the candidate large model is determined by the following formula. and candidate large models Reference values ​​for comparing the processing effects between them : in, Representing candidate large models The average accuracy rate; Representing candidate large models The standard deviation of the accuracy rate; Representing candidate large models The average accuracy rate; Representing candidate large models The standard deviation of the accuracy rate; This represents the number of candidate large models with high accuracy. when At that time, candidate large models are determined. The processing effect is better than that of the candidate large model The processing effect is better; when At that time, candidate large models are determined. The processing effect is better than that of the candidate large model The processing effect is better; when At that time, candidate large models are determined. The processing effect and candidate large model The processing results are the same; among them, These are empirical parameter values.

6. The method according to claim 5, characterized in that, The comparison of the processing performance of the two candidate large models, based on their accuracy parameters, also includes: when At that time, the processing time of the two candidate large models for processing the target task is compared to evaluate the processing performance of the two candidate large models.

7. The method according to claim 1, characterized in that, The objective task is to record the test results according to the evaluation indicators of each item in the preset inspection report, and to fill in the inspection conclusion in the preset inspection report based on the test results. Then, each candidate large model in the candidate large model set is invoked to process the target task, and the processing result set of each candidate large model is obtained, including: Construct questioning scripts based on the evaluation indicators and result records in the preset inspection report; The questioning script is submitted to each candidate big model. Each candidate big model, for each item in the preset inspection report, understands the requirements of the evaluation indicators and the recorded content of the result record according to the prompting script. By comparing the requirements of the evaluation indicators and the recorded content of the result record, it determines the option that should be filled in the inspection conclusion from multiple options, forming the processing result set of each candidate big model.

8. A large-scale automated quantitative screening device, characterized in that, The device includes: An acquisition module is used to acquire a set of standard answers for a target task; wherein the target task is a batch task, the batch task includes multiple task items, and the set of standard answers includes the standard answer corresponding to each task item; The processing module is used to call each candidate large model in the candidate large model set to process the target task, and obtain the processing result set of each candidate large model; wherein, the processing result set includes the processing result corresponding to each task item; The determination module is used to compare the processing result set of each candidate large model with the standard answer set to obtain the accuracy parameter of each candidate large model; The filtering module is used to compare the accuracy parameters of each candidate large model and filter out the target large model suitable for processing the target task from the set of candidate large models.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the automated quantitative screening method for large models as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the large model automated quantitative screening method as described in any one of claims 1 to 7.