A large model evaluation method and device
By selecting a model and receiving text input through an interactive interface, the system automatically selects target metrics and adjusts the automated evaluation mode, solving the problem of insufficient flexibility in large model evaluation methods. This enables an efficient and flexible evaluation process that can adapt to the personalized needs of different scenarios.
Patent Information
- Application Number
- CN202511557288.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-10-29
AI Technical Summary
Existing large-scale model evaluation methods lack flexibility and cannot adapt to dynamic changes in requirements, resulting in a rigid and inefficient evaluation process that fails to meet dynamic and personalized evaluation needs.
By selecting the model to be evaluated through an interactive interface and receiving text input, the system automatically selects target indicators from the indicator database, determines the automated evaluation mode, generates parallel evaluation tasks, and monitors the tasks in real time to adjust indicators and modes, thereby achieving flexibility and efficiency in the evaluation process.
It achieves flexibility and efficiency in the large-scale model evaluation process, enabling dynamic adjustment of indicators and modes, shortening the evaluation cycle, improving resource utilization, and enhancing the relevance and adaptability of the evaluation.
Smart Images

Figure CN121029623B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a large model evaluation method and device. BACKGROUND
[0002] With the rapid evolution of artificial intelligence technology, large models are increasingly widely used in various fields, and the demand for accurate evaluation of large model performance continues to grow. When evaluating large models in the current industry, a technical framework with a pre-set fixed process is generally used, that is, the evaluation system fixes the evaluation index set in advance (such as only containing general recognition accuracy, inference time, etc.), the user needs to determine all evaluation dimensions before the task starts, and the evaluation task is executed in a serial manner after being generated, that is, one index evaluation needs to be completed before starting the next one, and the automatic processing link mostly uses a single mode (such as full manual review or full system automatic calculation), which is not adapted to the differences in index characteristics.
[0003] The existing evaluation technology cannot meet the dynamic and personalized evaluation needs, and the primary problem is the lack of flexibility. Since the index is fixed in advance, the user cannot pass on personalized needs through natural language input and other convenient methods, nor can they add or adjust the index according to the changes in the scene, resulting in a disconnection between the evaluation scheme and the actual application needs, and the inability to accurately match the special evaluation demands of different fields.
[0004] In addition, the existing technology also has the problems of low efficiency and lack of dynamic adjustment capability. The serial execution of the task mode greatly prolongs the multi-index evaluation period, especially in multi-dimensional complex scenarios, the resource utilization rate is extremely low; at the same time, the single automatic mode either relies too much on manual work for simple indexes, causing cost waste, or over-automates complex indexes, leading to result deviation. More importantly, once the evaluation task is started, it cannot be adjusted in the middle, if the index parameters need to be modified or the evaluation dimensions need to be supplemented, the current task must be terminated, the historical data must be deleted, and then the configuration and start must be restarted, causing a lot of power and time loss, especially in scenarios where sample collection is difficult (such as power equipment defect samples), this fixed process will seriously hinder the evaluation progress. SUMMARY
[0005] Therefore, the present application provides a large model evaluation method and device to solve the problem of fixed evaluation process and lack of flexibility in traditional large model evaluation methods, which cannot adapt to dynamic demand changes.
[0006] Specifically, the present application is realized by the following technical solutions:
[0007] The first aspect of the present application provides a large model evaluation method, the method comprising:
[0008] selecting a model to be evaluated based on an interactive interface, and receiving a text input of evaluation indication content;
[0009] select a plurality of target indicators from an indicator database according to the text input;
[0010] determine an evaluation automation mode corresponding to the target indicators, and generate evaluation tasks according to the evaluation automation mode, wherein the plurality of evaluation tasks are performed in parallel;
[0011] monitor the plurality of evaluation tasks in real time, adjust the target indicators corresponding to any target evaluation task in response to a target indicator information modification instruction of the target evaluation task, calculate the indicator values corresponding to the newly added target indicators based on the evaluation results received before the information modification instruction, update the evaluation automation mode based on the changed target indicators, and the automation mode corresponding to each target indicator is different;
[0012] continue the evaluation of the target evaluation task with the updated automation mode and the adjusted target indicators.
[0013] The second aspect of the present application provides a large model evaluation device, the device comprising a receiving module, a selection module, a determination module, a monitoring module and a processing module;
[0014] The receiving module is configured to select a model to be evaluated based on an interactive interface, and receive a text input of evaluation instruction content;
[0015] The selection module is configured to select a plurality of target indicators from an indicator database according to the text input;
[0016] The determination module is configured to determine an evaluation automation mode corresponding to the target indicators, and generate evaluation tasks according to the evaluation automation mode, wherein the plurality of evaluation tasks are performed in parallel;
[0017] The monitoring module is configured to monitor the plurality of evaluation tasks in real time, adjust the target indicators corresponding to any target evaluation task in response to a target indicator information modification instruction of the target evaluation task, calculate the indicator values corresponding to the newly added target indicators based on the evaluation results received before the information modification instruction, update the evaluation automation mode based on the changed target indicators, and the automation mode corresponding to each target indicator is different;
[0018] The processing module is configured to continue the evaluation of the target evaluation task with the updated automation mode and the adjusted target indicators.
[0019] The large model evaluation method and device provided by the application provide a large model evaluation platform through software interaction, realize intervention, real-time adjustment and configuration and adaptive calculation of the large model evaluation process, and expand the applicable scenarios of the large model evaluation method. Through the complete large model evaluation process of demand input, index matching and subsequent dynamic adjustment of indexes and evaluation modes in the evaluation process, the flexibility, efficiency and adaptability of the evaluation process are improved, and a dynamic and personalized solution is provided for large model evaluation in different scenarios. By monitoring the modification of multiple evaluation task completion indexes and the update of the corresponding evaluation automation mode, on the one hand, the user can flexibly initiate index adjustment during evaluation running without interrupting the current task; on the other hand, the system reuses the evaluation result before receiving the modification instruction to calculate the new index value, avoiding repeated data collection and greatly reducing the computing power and time loss. By selecting multiple target indexes from the index database through text input, when the user inputs personalized requirements in natural language, the system can automatically match the corresponding target indexes from the index database without the user having to find or manually configure one by one, so that the evaluation can accurately meet the specific needs of different fields and different users, and the pertinence and adaptability of the evaluation are improved. In addition, through multi-task parallelism and automatic mode adaptation, on the one hand, multiple target index evaluation tasks are promoted in parallel without waiting for one task to complete before starting another, greatly shortening the overall evaluation period; on the other hand, each index is matched with a dedicated automatic mode, avoiding the efficiency loss caused by excessive reliance on manual or single automatic mode, and further improving the evaluation efficiency and result accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A flowchart of the large model evaluation method provided by the application;
[0021] Figure 2 A structural schematic diagram of the large model evaluation device provided by the application. DETAILED DESCRIPTION
[0022] The exemplary embodiments will be described in detail herein with reference to the drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with the present application.
[0023] The terms used in the present application are merely for the purpose of describing particular embodiments and are not intended to limit the present application. As used in the present application, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or," as used herein, refer to and encompass any or all possible combinations of one or more of the associated listed items.
[0024] It should be understood that, although the terms first, second, third, etc. can be employed in this application to describe various information, the information is not to be limited to these terms. These terms are only used to differentiate one piece of information from another. For example, a first information can also be termed a second information without departing from the scope of the application, and similarly, a second information can also be termed a first information. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining".
[0025] The following specific embodiments are given to introduce the technical solutions of the application in detail.
[0026] Embodiment one
[0027] Figure 1 The flowchart of the large model evaluation method provided by the application is shown in Embodiment one. Please refer to Figure 1 The method provided by the embodiment can include:
[0028] S101, selecting a to-be-evaluated model based on an interactive interface, and receiving a text input of evaluation instruction content.
[0029] It should be noted that the interactive interface is an operation platform for the user to interact with the large model evaluation system, which can be presented through a graphical interface (such as a web interface, a client software interface, etc.), and provides function entrances such as model selection and text input; the evaluation instruction content is text information input by the user to clearly indicate the evaluation requirements of the large model.
[0030] Specifically, the user selects the large model to be evaluated by clicking in the model list displayed on the interactive interface, and then inputs the text content of the evaluation requirements in the text input box. In this process, the large model evaluation system pre-constructs the function logic of the interactive interface, including the interactive logic of model selection and the receiving logic of text input. When the user operates on the interactive interface, the front-end interface of the system converts the user's selection and input operation into specific instructions and transmits them to the back-end processing. After the back-end module receives these information, it performs preliminary analysis and storage for subsequent process.
[0031] In addition, in addition to the above-mentioned method of manually selecting the model and inputting the text through the graphical interactive interface, the configuration file import method can also be used. That is, the user pre-writes the identification information of the to-be-evaluated model and the evaluation instruction content into a configuration file of a specific format, and then uploads the configuration file to the system through the file import function of the interactive interface. The system reads the configuration file content and parses the to-be-evaluated model and the evaluation instruction content. Compared with the former method, this method is suitable for the scene of batch evaluation task and can improve the efficiency of multi-task initiation.
[0032] S102. Select multiple target indicators from the indicator database based on the text input.
[0033] It should be noted that target metrics refer to metrics selected from the metric database that match the evaluation instructions input by the user and are used to conduct large-scale model evaluations. Each metric includes attribute information such as metric name, calculation rules, domain, and applicable scenarios, facilitating system retrieval and matching based on requirements.
[0034] Specifically, the target indicators stored in the indicator database are classified according to the evaluation scenario. In the power equipment inspection scenario, typical target indicators include equipment defect identification accuracy (evaluating the accuracy of defect identification), single sample inference time (evaluating the model running efficiency), defect area positioning accuracy (evaluating the defect location positioning capability), and small defect identification recall rate (evaluating the coverage capability of minor defects).
[0035] It should be noted that the system performs natural language processing on the input evaluation instruction text to extract key information, such as the evaluation objects (e.g., transformers, insulators) and evaluation dimensions (e.g., recognition accuracy, inference speed). Then, based on this key information, it searches and matches in the indicator database, selecting multiple indicators highly relevant to the text content as target indicators. Specifically, natural language processing technology is used to perform semantic analysis on the evaluation instruction text, transforming the natural language descriptions in the text into keywords or semantic vectors that the system can understand. Simultaneously, each indicator in the indicator database has a corresponding keyword tag or semantic feature vector. By calculating the similarity between the text semantics and the indicator semantic features, indicators with similarity higher than a preset threshold are selected as target indicators.
[0036] Optionally, in addition to semantic matching based on natural language processing, template matching can also be used. The system predefines templates for various common assessment needs, and each template corresponds to a fixed set of target indicators. When the assessment instruction text is received, the system matches the text with these templates. If the text matches the characteristics of a certain template (such as containing a specific combination of keywords), the target indicator corresponding to that template is directly selected.
[0037] In this way, based on the text input by the user, the system automatically selects multiple target indicators from the indicator database, eliminating the need for manual searching and filtering of indicators one by one. This reduces the cost of manual operation and the probability of errors, automates and automates indicator selection, improves the efficiency and accuracy of indicator matching, and can quickly respond to different evaluation needs.
[0038] Furthermore, after selecting multiple target indicators from the indicator database based on the text input, the process includes:
[0039] (1) Display the selected multiple target indicators and default attribute parameters in the interactive interface.
[0040] It should be noted that after selecting the target indicators from the indicator database, the interface will display the information in a structured form (such as a list, card), including the name of each target indicator and the default attribute parameters, where these default attribute parameters are the general configurations preset by the indicator database, defining the calculation method of the target indicator and the judgment rule of whether the target indicator meets the standard, covering calculation rules, weight coefficients, and threshold ranges. In specific implementation, the indicator data passed from the backend is converted into visual interface elements through front-end rendering technology, ensuring that users can intuitively understand the basic configuration of the current indicator.
[0041] (2) Receive the adjustment parameters input by the user through the interactive interface.
[0042] It should be noted that for each indicator attribute (such as calculation rules, weights, thresholds) displayed in the interface, the system provides interactive input components (such as text boxes, sliders, drop-down menus), and users can input adjustment values according to actual evaluation needs (such as industry standards, business focus). For example, in the power inspection scenario, if the user believes that the default qualified threshold (300ms) for single-sample inference time is too wide, they can adjust it to 200ms through the slider. Further, the system will perform legality verification (such as format, value range) on the input parameters to ensure that the adjustment is effective and the data is temporarily stored.
[0043] (3) Update the adjusted parameters as new configurations for the corresponding target indicators and store them.
[0044] In this step, after selecting the target indicators, the default attribute parameters are first displayed, and then the user's adjustment parameters are received and updated for storage, realizing flexible pre-configuration of indicator attribute parameters, meeting the user's individual needs in different evaluation scenarios, making the indicator value more suitable for actual evaluation goals, and laying a foundation for subsequent generation of evaluation tasks and accurate calculation.
[0045] S103, determine the evaluation automation mode corresponding to the target indicator, and generate an evaluation task according to the evaluation automation mode, wherein multiple evaluation tasks are performed in parallel.
[0046] It should be noted that the evaluation automation mode refers to the type of automatic process for completing the evaluation of the target indicator, and each mode corresponds to a different human-machine collaboration method, suitable for different complexity and different manual intervention needs of the indicator evaluation scenario; the evaluation task is an independent execution unit generated for a single target indicator according to its matched automation mode, containing complete process logic such as data collection, calculation, and result output. Multiple evaluation tasks can be run simultaneously.
[0047] Specifically, the system matches the most suitable type from multiple modes according to the attribute of each target indicator (such as calculating logical complexity, whether manual review is required), wherein the mode matching is based on a preset rule engine, the system has a built-in mode matching rule library, and the mode allocation is automatically completed by comparing the indicator attribute with the rule library.
[0048] Specifically, multiple evaluation tasks need to be run, and the running time sequence thereof can be scheduled, that is, there are evaluation tasks running simultaneously, and there are also evaluation tasks running in sequence. Preferably, the method further comprises: determining the correlation between each evaluation task; dividing evaluation task groups based on the correlation, each group including multiple evaluation tasks, each group being independently and simultaneously run, and the multiple evaluation tasks in one group being run in sequence; for each evaluation task group, determining the target indicators of each evaluation task group, determining the dependent variables of the target indicators according to the calculation method in the attribute of the target indicators, and generating the time sequence of each evaluation task in the evaluation task group according to the dependent variables. The dependent variable refers to a variable that needs to be known but is calculated by other evaluation tasks in the calculation method of the target indicator. When dividing groups, all dependent variables of each target indicator are determined according to the calculation method in the target indicator of the evaluation task, and evaluation tasks with the same dependent variables or dependent variables having a mutual relationship are divided into a group. After dividing the groups, for a group, a dependent relationship order chain of dependent variables is established according to the calculation method of each target indicator, all variable dependent relationship graphs in the group are obtained by merging each chain, each variable in the dependent relationship graph corresponds to an evaluation task, and one evaluation task can correspond to multiple variables. The time sequence of the corresponding evaluation task is generated according to the variable time sequence in the dependent relationship graph.
[0049] In this way, on the one hand, it is ensured that indicators with data dependencies are calculated in the correct order, ensuring the accuracy of the evaluation results; on the other hand, task groups without dependent relationships are executed in parallel, maximizing the use of system resources, shortening the overall evaluation time, and solving the problems of calculation errors caused by complete parallelism and low efficiency caused by complete serial execution.
[0050] It should be noted that the evaluation automation mode provided by the present application includes a full automation mode, a manual assistance mode, and a manual dominant mode; wherein the full automation mode does not require manual input, automatically collects data from the evaluation data source through a preset algorithm, completes indicator calculation according to the built-in calculation rules, and directly outputs the final indicator value; the manual assistance mode automatically collects data from the evaluation data source and completes preliminary calculation, pushes the preliminary calculation result to the interactive interface, and generates the final indicator value after receiving the manual review confirmation instruction; the manual dominant mode receives the calculation parameters input by the human through the interactive interface, automatically collects matching data from the evaluation data source based on the calculation parameters and completes the calculation, and directly outputs the final indicator value.
[0051] Specifically, the full automation mode is completely without manual operation, relying on system preset logic to complete the whole process from data collection to result output, suitable for index evaluation with clear calculation rules, standardized data sources, and no need for manual judgment. Specifically, the system automatically captures the raw data required for the index (such as single sample reasoning time index needs to collect the processing time data of each sample) from the specified evaluation data source (such as power equipment inspection sample library) through the preset algorithm (such as interface call, batch reading), and then automatically calculates the collected data based on the built-in calculation rules in the index database (such as single sample reasoning time = total processing time / sample total number), without the need for manual input of any parameters; finally, after calculation is completed, the system directly outputs the final index value (such as 280ms / sample), without the need for manual review or confirmation, and the result is directly used for evaluation results.
[0052] The manual assistance mode is that data collection and preliminary calculation are completed by the system, but the final result needs to rely on manual review and confirmation, suitable for index evaluation with automatic calculation process but result accuracy needs to be verified by manual (such as involving sample label accuracy, subjective judgment scene). Specifically, consistent with the full automation mode, the system automatically collects data from the evaluation data source and completes preliminary calculation according to the built-in calculation rules to obtain preliminary results, and then the system pushes the preliminary calculation results and associated support data (such as screenshots of samples with recognition errors, label information) to the interactive interface, prompting the user to review; the user reviews the preliminary results and support data through the interactive interface to confirm whether the results are reasonable (such as judging whether the sample with recognition error is a label annotation error or a model problem), and if the results are correct, clicks to pass the review, and the system generates the final index value; if there is a problem (such as label error), the calculation can be triggered again after correction, and the result is confirmed again.
[0053] The manual assistance mode is that data collection and preliminary calculation are completed by the system, but the final result needs to rely on manual review and confirmation, suitable for index evaluation with automatic calculation process but result accuracy needs to be verified by manual (such as involving sample label accuracy, subjective judgment scene). Specifically, consistent with the full automation mode, the system automatically collects data from the evaluation data source and completes preliminary calculation according to the built-in calculation rules to obtain preliminary results, and then the system pushes the preliminary calculation results and associated support data (such as screenshots of samples with recognition errors, label information) to the interactive interface, prompting the user to review; the user reviews the preliminary results and support data through the interactive interface to confirm whether the results are reasonable (such as judging whether the sample with recognition error is a label annotation error or a model problem), and if the results are correct, clicks to pass the review, and the system generates the final index value; if there is a problem (such as label error), the calculation can be triggered again after correction, and the result is confirmed again.
[0054] In combination with the foregoing description, for example, the single-sample inference time calculation logic is simple, does not require human intervention, and matches the full automation mode; the device defect identification accuracy requires manual confirmation of sample labels, matching the human-assisted mode; and the custom risk score requires user input of weight parameters, matching the human-led mode.
[0055] This step matches a dedicated automation mode for each indicator, generates independent tasks and executes them in parallel, reduces the overall evaluation time through multi-task parallelism, and is particularly suitable for complex multi-indicator scenarios; by matching the mode according to the characteristics of the indicators, the automation efficiency and the need for human intervention are balanced, and the flexibility of the evaluation is increased by independent running of each task.
[0056] S104, real-time monitoring of multiple evaluation tasks, responding to any target evaluation task target indicator information modification instruction, adjusting the target indicator corresponding to the target evaluation task, calculating the indicator value corresponding to the newly added target indicator based on the evaluation results received before the information modification instruction, modifying and updating the evaluation automation mode based on the changed target indicator, and the automation mode corresponding to each target indicator is different.
[0057] It should be noted that the target indicator information modification instruction refers to an adjustment instruction initiated by the user through the interactive interface during the running of the evaluation task, which mainly includes two types: one is to add a target indicator, and the other is to modify an existing target indicator. The evaluation results received before the information modification instruction refer to the evaluation data generated by the target evaluation task before the user initiates the modification instruction, including basic sample data, intermediate results (such as partial sample identification results), and output indicator values.
[0058] Specifically, the system for evaluating large models captures the running logs of each evaluation task (such as data collection progress, computing node status, and result output) in real time, monitors whether the task is running normally, and listens to the target indicator information modification instruction from the interactive interface. Among them, the double mechanism of timing polling and event listening can be used to ensure real-time task progress by polling the running data of each task; at the same time, the instruction triggering event (such as button click, form submission) of the interactive interface is listened to, and once there is a modification instruction, the subsequent adjustment process is triggered immediately to avoid delay. The contents of the listening include task running state, resource occupation, whether an exception is triggered, and whether there is a user-initiated modification instruction, etc.
[0059] When any target evaluation task modification instruction is captured, the instruction is first parsed to determine the modification type and extract the key information, for example, if it is a new target indicator instruction, the indicator name and indicator attribute can be extracted; if it is an existing indicator attribute modification instruction, the indicator ID to be modified and the modification content can be extracted. Through instruction parsing, the object and content of subsequent adjustment are determined.
[0060] Further, different adjustment logics are executed for the two scenarios of adding new indicators and modifying attributes. If it is the scenario of adding new target indicators, the system calls the received information modification instruction to modify the previous evaluation results, analyzes the calculation logic of the added indicators, and judges whether the required data is complete. If the data is complete, the historical data is directly called to calculate the initial indicator value according to the calculation rule of the added indicator; if the data is incomplete, the missing data is supplemented based on the existing intermediate results, and then the initial indicator value is calculated. Further, according to the attribute of the added indicator (such as calculation complexity, whether manual intervention is required), a dedicated mode is matched from the full automation mode, manual assistance mode and manual dominant mode, and finally the added indicator and its initial indicator value and the matched automation mode are added to the indicator set of the current target evaluation task, and are executed in parallel with the original indicators, without affecting the running of other tasks.
[0061] If it is the attribute modification scenario, the historical evaluation results (such as original sample data and intermediate calculation values) are called to recalculate according to the modified attribute (the attribute at least includes calculation rule, weight coefficient and threshold range). If the calculation rule or weight coefficient is modified, the indicator value can be recalculated according to the new rule based on the historical sample data; if the threshold range is modified, the indicator value does not need to be recalculated, and the automatic verification process is directly triggered to compare the original indicator value with the new threshold value to generate the adjusted qualified judgment result.
[0062] In specific implementation, the above two scenarios are further described. The calculation of the indicator value corresponding to the added target indicator based on the evaluation results before the information modification instruction is received, and the modification and update of the evaluation automation mode based on the changed target indicator, include:
[0063] (1) receiving input information of the added target indicator, the input information including an indicator name and an indicator attribute.
[0064] It should be noted that the indicator name is used to uniquely identify the added indicator, and the indicator attribute is used to determine the core configuration of the indicator calculation logic and mode adaptation, including calculation rule, data dimension and manual intervention requirement (here, the indicator attribute is the indicator attribute itself). Specifically, the added indicator information input by the user can be received through an interactive interface, the interactive interface provides a structured input form (such as drop-down selection of data dimension and text box for filling calculation rule), and the system performs legality verification on the input information (such as avoiding duplicate indicator names and ensuring calculation rule format compliance), so as to ensure that the information can be directly used in subsequent steps.
[0065] (2) analyzing the calculation logic of the added target indicator to determine the required basic parameters and intermediate results.
[0066] It should be noted that the system performs semantic analysis on the calculation rules of the newly added indicators, extracts the "basic parameters" (such as defect sample coordinate data and qualified deviation threshold) required for calculation and intermediate results (such as the total number of defect samples calculated in the previous evaluation and the sample ID association table). Based on the pre-set index calculation logic knowledge base, the calculation rules described in natural language (such as the number of qualified samples / total number of samples) can be converted into a parameter-operation logic mapping relationship (such as parameter A = qualified sample number, parameter B = total sample number, operation logic = A / B) that can be recognized by the system, and the data requirement list is clear. Through such steps, the calculation dependency of the newly added indicators can be resolved, and it can be determined whether historical data can be reused to avoid repeated resource collection and waste.
[0067] (3) When the required basic parameters or intermediate results are already included in the evaluation results before receiving the new input, the initial index value of the new target indicator is directly generated by calling the data.
[0068] Compared with the data requirement list determined in step (2), if the basic parameters and intermediate results are already present in the historical data, these data are directly called, automatically operated according to the calculation rules of the newly added indicators, and the initial index value is generated. Here, the system can establish an index of indicators-data association, and the historical evaluation results can be stored according to the index ID and data type, so that the required data can be quickly located through the index when compared, avoiding full-scan database, and the preset index calculation engine is called during the calculation process to perform calculation according to the parsed operation logic.
[0069] (4) When the required basic parameters or intermediate calculation results are not completely included, the missing parameters are calculated based on the existing intermediate results, and the initial index value of the new target indicator is generated; according to the indicator attribute, the target evaluation automation mode corresponding to the new target indicator is matched from the evaluation automation mode, and the new target indicator and the initial index value of the new target indicator are replaced into the current target indicator set.
[0070] Specifically, this process can be executed in two stages. In the first stage, the missing parameters are calculated, and if the historical data only contains part of the basic parameters / intermediate results (such as sample shooting angle data), the system supplements the missing data (such as sample ID associated shooting angle) from the evaluation data source based on the existing intermediate results (such as sample ID), and then generates the initial index value according to the calculation rules. In the second stage, the mode is matched and the task is added, the target mode is matched from the three types of automation modes according to the attribute of the new indicator, and finally the new target indicator set is added to the current target indicator set, sharing computing power with the original indicators and executing in parallel. Among them, when supplementing data, an incremental collection strategy can be used to collect only the missing data rather than full re-collection.
[0071] In a specific implementation, for the attribute modification scenario, after receiving the attribute modification instruction, it includes:
[0072] (1) When the calculation rule or weight coefficient is modified, the calculation logic corresponding to the target index is called, the index value of the target index is recalculated based on the existing evaluation data and intermediate results, and the updated index value is replaced into the current target index set.
[0073] Specifically, after receiving the instruction to modify the calculation rule or weight coefficient, three-step operation is performed, the new calculation logic of the index is called, the existing evaluation data (historical evaluation data) and intermediate results are called, the index value is recalculated according to the new calculation logic, and the old value in the current target index set is replaced by the updated value. Among them, the system retains the historical attribute version of the index, automatically triggers the recalculation of the index value after modification, and preferentially reuses the historical data (avoids repeated collection) during recalculation, while recording the comparison of the index values before and after.
[0074] (2) When the threshold range is modified, the automatic verification process of the target index is retriggered, and the adjusted evaluation result is generated based on the current evaluation data.
[0075] It should be noted that, unlike after the calculation rule or weight coefficient is modified, the verification is performed after the threshold is modified instead of recalculation. At this time, the index value is not recalculated, the system retrigger the automatic verification process of the index, compare the current index value with the new threshold, generate the adjusted evaluation result (such as updating the original unqualified judgment to qualified), and synchronize to the interactive interface and result library.
[0076] (3) After completing the attribute modification, the corresponding evaluation automation mode is re-matched according to the modified index attribute.
[0077] It should be noted that the modification of the index attribute may change the need for manual intervention, and the automation mode needs to be updated synchronously to avoid evaluation process abnormalities.
[0078] In addition, it should also be noted that in addition to adding target indexes and modifying the attributes of existing target indexes, the target index information modification instruction also includes the operation of deleting target indexes. Users can initiate an index deletion instruction for a target evaluation task through the interactive interface during the running of the evaluation task, select the index to be deleted in the current target index set displayed on the interactive interface, click the delete button and confirm, and the system receives the deletion instruction, extracts the to-be-deleted index ID and the target evaluation task identifier to which it belongs. At the same time, the system will perform a legality check on the deletion instruction, for example, checking whether the to-be-deleted index belongs to the currently running target evaluation task and whether there is a dependency relationship (such as the index being the calculation basis for other indexes, then prompting “there is a dependency, cannot be deleted”), to avoid evaluation process abnormalities caused by accidental deletion, and triggering the subsequent deletion process after the verification is passed.
[0079] Further, the system executes the operation of removing the indicators from the target indicator set and terminating the corresponding evaluation process based on the to-be-deleted indicator ID in the deletion instruction. Specifically, in the indicator set of the current target evaluation task, all configuration information of the indicator to be deleted (including the indicator name, attribute, matched automation mode, and generated indicator value) is deleted, and the indicator display list of the interactive interface is updated, so that the deleted indicator is no longer displayed. At the same time, the evaluation task thread corresponding to the indicator is stopped, and system resources such as computing power and memory occupied by the evaluation task thread are released, so that the resources can be allocated to other running indicator tasks. It should be noted that if the indicator has incomplete evaluation operations (such as part of the sample data is being collected, and the preliminary calculation result is to be audited), the system will first terminate these incomplete operations to avoid generating invalid data. At the same time, the deletion time, operator, and historical evaluation result of the indicator are recorded, and the process is completely terminated.
[0080] S105, continue the evaluation of the target evaluation task with the updated automation mode and the adjusted target indicator.
[0081] It should be noted that the specific evaluation process after the update is described above and will not be repeated here. In this step, the evaluation of the target evaluation task with the updated automation mode and the adjusted target indicator continues without interrupting the overall task. The unmodified indicators can continue to run according to the original mode, avoiding the whole process from being stalled due to local adjustment, and greatly improving the evaluation efficiency. Secondly, this step can also ensure the consistency of the evaluation results. The historical incomplete samples are reprocessed according to the new configuration, and the new samples are directly adapted to the new rules, ensuring that all samples follow the same evaluation standard and avoiding the problem of incomparable results caused by configuration differences between old and new samples. In addition, resources can be maximally reused, and there is no need to re-collect historical data or rebuild the task environment. Only incremental adjustment is made to the changed part, which significantly reduces the system computing power and time cost.
[0082] It should also be noted that after the evaluation of the target evaluation task with the updated automation mode and the adjusted target indicator, the following operations are performed: real-time collection of evaluation process data and final result data of each target indicator; generation of a visual chart according to the type and time dimension of the target indicator, marking the time node of the indicator adjustment and the corresponding mode change in the chart; and display of the visual chart on the interactive interface.
[0083] Specifically, the visualization process requires acquiring two types of data for each target indicator: first, assessment process data, covering intermediate results of indicator calculations and records of automation mode switching; and second, final result data, including the final values calculated for each indicator according to the current configuration and the pass / fail judgment results. During data acquisition, information is categorized and stored according to ID, timestamp, and data type. Next, visualization charts are generated based on indicator type and time dimension. Indicators that need to reflect changes in value over time are categorized as numerical indicators, while indicators that need to reflect differences between different categories of data are categorized as categorical indicators. Then, appropriate chart types are generated accordingly. For numerical indicators, a line chart is generated with assessment time as the horizontal axis and indicator value as the vertical axis, clearly showing the real-time trend of indicator changes as the assessment progresses. For categorical indicators, a bar chart is generated with indicator category as the horizontal axis and corresponding indicator value as the vertical axis, intuitively presenting the distribution of indicators under each category.
[0084] It should also be noted that, after continuing the evaluation of the target assessment task with the updated automation mode and adjusted target metrics, the following is also included:
[0085] (1) Based on the real-time evaluation results of each target indicator, automatically identify erroneous samples whose deviation from the preset standard exceeds the threshold.
[0086] It should be noted that, based on the real-time evaluation results of each target indicator, erroneous samples that exceed the preset standard deviation threshold are automatically identified, including:
[0087] (i) Call the preset standard database corresponding to the target indicator to obtain the historical deviation threshold and sample label standard of the target indicator in the same evaluation scenario.
[0088] It should be noted that the system first calls the preset standard database corresponding to the current target indicator. This database stores historical data of each indicator in similar evaluation scenarios, as well as sample label standards. The obtained historical deviation threshold is used as a quantitative standard to judge whether the sample is biased. The sample label standards are used to compare the accuracy of the sample prediction results in the future to ensure that the identification basis complies with the general specifications of the industry or scenario.
[0089] (ii) Compare the predicted results of each evaluation sample with the actual labels and calculate the deviation.
[0090] For each evaluation sample, the system compares the prediction results output by the large model with the actual labels of the sample dimension by dimension, and calculates the deviation between the two using a preset algorithm. For example, for classification indicators (such as defect type identification), the deviation is calculated using category matching degree (100% deviation if there is no match), while for numerical indicators (such as positioning accuracy), the deviation is calculated using absolute error / relative error (e.g., if the positioning deviation is 8 pixels and the preset allowable deviation is 5 pixels, then the deviation is 60%).
[0091] (iii) When the deviation degree exceeds the historical deviation threshold of the target index, mark as a suspected error sample.
[0092] The calculated sample deviation degree is compared with the historical deviation threshold obtained in step (i). If the deviation degree exceeds the threshold (e.g., a sample deviation degree of 8%, exceeding the threshold of "±5%"), the sample is marked as a suspected error sample. This step can preliminarily exclude normal samples within a reasonable range of deviation, focus on abnormal samples that need further verification, and reduce the workload of subsequent verification.
[0093] (iii) Perform secondary verification on the suspected error sample by searching the sample feature library.
[0094] To avoid misjudgment caused by label annotation errors, sample data abnormalities, and other non-model factors, the system searches the sample feature library (which stores typical features of various samples) to perform secondary verification on the suspected error sample. If it is confirmed that the model misjudged, rather than a label error, the sample passes the secondary verification; if the search finds that it is a label annotation error, the sample is excluded from the error sample range.
[0095] Further, the identification, target index, deviation degree, and feature parameters of the suspected error sample that passes the secondary verification are recorded, which can provide a basis for subsequent analysis of deviation causes and parameter problem positioning.
[0096] (2) For the identified error sample, analyze the calculation logic and parameter configuration of the associated target index to locate the fine-tuning parameter item that caused the deviation.
[0097] It should be noted that the system first obtains the complete configuration of the target index to which the error sample belongs, covering the calculation rules, parameter configuration (weight coefficient, threshold range, sample screening condition, etc.), and automatic mode associated parameters, to ensure coverage of key nodes in the whole evaluation process. For example, the configuration of the device defect recognition accuracy index includes the calculation rule (number of correctly recognized samples / total number of effective samples), sample screening condition (defect area pixel ratio ≥ 5%), and sample proportion for manual assistance mode ≥ 30%. For the scenario-based deviation of error samples, compare the index configuration one by one to find the directly related parameters. If the error sample set is a small pixel defect sample (such as defect area pixel ratio 3%-4%), compare the sample screening condition to find that the threshold of defect area pixel ratio ≥ 5% in the index configuration causes this type of sample to be excluded from the effective calculation, and the model's recognition result is not included in the accuracy statistics, thus producing a deviation. At this time, the defect area pixel ratio screening threshold is the fine-tuning parameter item that causes the deviation. Preferably, as an optional embodiment, for the identified error sample, obtain the model parameter combination of the error sample recognition, determine the calculation logic and parameter configuration of the error sample associated target index, and determine the parameter combination with the highest impact in the model parameter combination according to the calculation logic. The parameter combination is used as the fine-tuning parameter item.
[0098] (3) Generate a parameter correction suggestion, and recalculate the target index value corresponding to the error sample based on the corrected parameters, and update the evaluation result.
[0099] Specifically, generating a parameter correction suggestion requires judgment based on actual conditions, which can be referred to the description of related technologies and will not be repeated here.
[0100] The method provided in the embodiment realizes flexible adjustment of indexes and automatic modes in the evaluation process by monitoring the evaluation task and the response index modification instruction in real time, that is, the user can add target indexes, modify the attributes of existing indexes, or delete unnecessary indexes during the task running, and the system can reuse the historical evaluation results to calculate the added index values and re-match the automatic mode based on the modified attributes, without terminating and restarting the whole process, so that the evaluation can dynamically adapt to the changing requirements in the middle of the process, and the flexibility and scene adaptation ability of the evaluation are significantly improved. Through the design of automatic index selection by text input and multi-task parallel execution, the user only needs to input the natural language evaluation requirements through the interactive interface, and the system can automatically match the target indexes from the index database based on the text without manual selection and configuration one by one. At the same time, the evaluation tasks corresponding to multiple target indexes are pushed forward in parallel according to the exclusive automatic mode, avoiding the time-consuming problem of traditional serial evaluation, especially in the multi-index complex scene, which can significantly shorten the overall evaluation period and improve the resource utilization. In addition, through the design of automatic process instead of manual operation and reuse of historical data, the core links such as index selection, task generation, and mode matching are automatically completed by the system, reducing the operation cost and error probability of manual configuration; when the indexes are adjusted, the evaluation results before receiving the modification instruction are reused, which can avoid repeated data collection and task environment reconstruction, greatly reducing the system computing power and time loss, so that the evaluation can be efficiently promoted while realizing the optimal utilization of resources.
[0101] Embodiment two
[0102] Corresponding to the foregoing embodiment of the large model evaluation method, the present application also provides an embodiment of a large model evaluation device.
[0103] Figure 2 The structure diagram of the second embodiment of the large model evaluation device provided in the present application is shown in FIG. 2. Please refer to Figure 2 The device provided in the embodiment includes a receiving module 210, a selection module 220, a determination module 230, a monitoring module 240, and a processing module 250.
[0104] The receiving module 210 is configured to select a model to be evaluated based on an interactive interface, and receive a text input of evaluation instruction content.
[0105] The selection module 220 is configured to select multiple target indexes from an index database according to the text input.
[0106] The determination module 230 is configured to determine an evaluation automatic mode corresponding to the target indexes, and generate evaluation tasks according to the evaluation automatic mode, wherein multiple evaluation tasks are executed in parallel.
[0107] The monitoring module 240 is configured to monitor a plurality of evaluation tasks in real time, adjust a target index corresponding to a target evaluation task in response to a target evaluation task target index information modification instruction, calculate an index value corresponding to the newly added target index based on an evaluation result received before the information modification instruction, update an evaluation automation mode based on the changed target index, and the automation mode corresponding to each target index is different;
[0108] The processing module 250 is configured to continue the evaluation of the target evaluation task in the updated automation mode and the adjusted target index.
[0109] The device of the embodiment can be used to perform Figure 1 The steps of the method embodiment are similar in specific implementation principles and implementation processes, and will not be described here.
[0110] The implementation processes of the functions and roles of the units in the device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be described here.
[0111] For the device embodiment, since it basically corresponds to the method embodiment, the related parts can be referred to the part of the method embodiment. The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the scheme of the present application. Those skilled in the art can understand and implement without creative labor.
[0112] The above is only the preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for evaluating large models, characterized in that, The method includes: Select the model to be evaluated based on the interactive interface and receive text input of the evaluation instructions; Select multiple target metrics from the metric database based on the text input; Determine the automated assessment mode corresponding to the target indicator, and generate assessment tasks according to the automated assessment mode, wherein multiple assessment tasks are performed in parallel; The system monitors multiple assessment tasks in real time and responds to any target assessment task's target indicator information modification instruction. It adjusts the target indicator corresponding to the target assessment task, calculates the indicator value corresponding to the newly added target indicator based on the assessment results prior to receiving the information modification instruction, and updates the assessment automation mode based on the changed target indicator. Each target indicator corresponds to a different automation mode. The process of calculating the indicator value corresponding to the newly added target indicator based on the assessment results prior to receiving the information modification instruction, and updating the assessment automation mode based on the changed target indicator, includes: receiving input information for the newly added target indicator, including the indicator name and indicator attributes; parsing the calculation logic of the newly added target indicator to determine its required basic parameters and intermediate results; when the required basic parameters or intermediate results are already included in the assessment results before receiving the new input, directly calling the data to generate the initial indicator value of the newly added target indicator; when the required basic parameters or intermediate results are not fully included, supplementing the calculation of missing parameters based on existing intermediate results, and generating the initial indicator value of the newly added target indicator; and matching the target assessment automation mode corresponding to the newly added target indicator from the assessment automation modes according to the indicator attributes, and replacing the newly added target indicator and its initial indicator value into the current target indicator set. The evaluation of the target assessment task will continue with the updated automation mode and adjusted target metrics.
2. The large model evaluation method according to claim 1, characterized in that, The automated assessment modes include fully automated mode, human-assisted mode, and human-led mode; The fully automated mode requires no manual input, automatically collects data from the evaluation data source through a preset algorithm, completes the indicator calculation according to the built-in calculation rules, and directly outputs the final indicator value. The manual assistance mode automatically collects data from the assessment data source and completes preliminary calculations. The preliminary calculation results are pushed to the interactive interface, and the final indicator value is generated after receiving the manual review and confirmation instruction. The human-led mode receives calculation parameters input manually through an interactive interface, automatically collects matching data from the evaluation data source based on the calculation parameters, completes the calculation, and directly outputs the final indicator value.
3. The large model evaluation method according to claim 1, characterized in that, The adjustment of the target indicators corresponding to the target evaluation task includes attribute modification, and the attributes include at least calculation rules, weight coefficients and threshold ranges. Among them, after receiving the attribute modification instruction, when the calculation rules or weight coefficients are modified, the calculation logic corresponding to the target indicator is called, the indicator value of the target indicator is recalculated based on the existing evaluation data and intermediate results, and the updated indicator value is replaced in the current target indicator set. When the threshold range is modified, the automated verification process of the target indicator is retried, and the adjusted evaluation result is generated based on the current evaluation data. After the attribute modification is completed, the corresponding automated assessment mode is re-matched based on the modified indicator attributes.
4. The large model evaluation method according to claim 1, characterized in that, After selecting multiple target indicators from the indicator database based on the text input, the process includes: The interactive interface displays the selected target metrics and their default attribute parameters. Receive adjustment parameters input by the user through the interactive interface; Update the adjusted parameters to the new configuration corresponding to the target indicators and save them.
5. The large model evaluation method according to claim 1, characterized in that, After continuing the evaluation of the target assessment task with the updated automation mode and adjusted target metrics, including: Real-time collection of evaluation process data and final result data for each target indicator; Generate visualization charts based on the type and time dimension of the target indicators, and mark the time nodes for indicator adjustments and the corresponding pattern changes in the charts; The visualization chart is displayed in the interactive interface.
6. The large model evaluation method according to claim 1, characterized in that, After continuing the evaluation of the target assessment task with the updated automation mode and adjusted target metrics, it also includes: Based on the real-time evaluation results of each target indicator, erroneous samples that exceed the preset standard deviation threshold are automatically identified. For the identified erroneous samples, analyze the calculation logic and parameter configuration of the target indicators associated with the erroneous samples, and locate the fine-tuning parameter items that cause the deviation; Generate parameter correction suggestions, recalculate the target index values corresponding to the erroneous samples based on the corrected parameters, and update the evaluation results.
7. The large model evaluation method according to claim 6, characterized in that, The real-time evaluation results based on each target indicator automatically identify erroneous samples whose deviation from the preset standard deviation exceeds a threshold, including: Call the preset standard database corresponding to the target indicator to obtain the historical deviation threshold and sample label standard of the target indicator in the same evaluation scenario; The predicted results of each evaluation sample are compared with the actual labels, and the deviation is calculated. When the deviation exceeds the historical deviation threshold of the target indicator, it is marked as a suspected erroneous sample; The suspected erroneous samples are then verified by searching the sample feature database.
8. The large model evaluation method according to claim 1, characterized in that, The generation of assessment tasks according to the automated assessment mode includes: Determine the correlation between the various assessment tasks; Based on the aforementioned correlation, the assessment task groups are divided, with each group containing multiple assessment tasks. Each group runs independently and simultaneously, and the multiple assessment tasks in a group run in sequence. For each assessment task group, the target indicators for each assessment task group are determined, the dependency variables of the target indicators are determined according to the calculation method in the attribute of the target indicators, and the time sequence of each assessment task in the assessment task group is generated according to the dependency variables.
9. A large model evaluation device, characterized in that, The device includes a receiving module, a selection module, a determination module, a monitoring module, and a processing module; The receiving module is used to select the model to be evaluated based on the interactive interface and receive text input of the evaluation instructions. The selection module is used to select multiple target indicators from the indicator database based on the text input; The determining module is used to determine the evaluation automation mode corresponding to the target indicator, and generate evaluation tasks according to the evaluation automation mode, wherein multiple evaluation tasks are performed in parallel. The monitoring module is used to monitor multiple assessment tasks in real time, respond to any target assessment task's target indicator information modification instruction, adjust the target indicator corresponding to the target assessment task, calculate the indicator value corresponding to the newly added target indicator based on the assessment results prior to receiving the information modification instruction, and modify and update the assessment automation mode based on the changed target indicator. The automation mode corresponding to each target indicator is different. The step of calculating the indicator value corresponding to the newly added target indicator based on the assessment results prior to receiving the information modification instruction, and modifying and updating the assessment automation mode based on the changed target indicator, includes: receiving input information for the newly added target indicator, the input information including the indicator name and... The system analyzes the calculation logic of the newly added target indicator to determine its required basic parameters and intermediate results. When the required basic parameters or intermediate results are already included in the evaluation results before receiving the new input, the system directly calls the data to generate the initial indicator value of the new target indicator. When the required basic parameters or intermediate results are not fully included, the system supplements the calculation of missing parameters based on the existing intermediate results and generates the initial indicator value of the new target indicator. According to the indicator attributes, the system matches the target evaluation automation mode corresponding to the new target indicator from the evaluation automation mode and replaces the new target indicator and its initial indicator value into the current target indicator set. The processing module is used to continue the evaluation of the target evaluation task in an updated automated mode and with adjusted target indicators.
Citation Information
Patent Citations
Model evaluation method and device, electronic equipment, storage medium and program product
CN119621503A
Model data processing method and device, equipment and medium
CN119884682A