Model evaluation system, data processing method and apparatus, and device and medium
By using a multi-dimensional evaluation system and user feedback optimization logic, the problem of single evaluation of model data processing results in existing technologies has been solved, and accurate evaluation and optimization of model performance has been achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INSPUR SUZHOU INTELLIGENT TECH CO LTD
- Filing Date
- 2025-11-11
- Publication Date
- 2026-05-21
AI Technical Summary
Existing technologies use a single method to evaluate the results of model data processing, which cannot effectively assess the actual performance of each model.
A multi-dimensional evaluation system is adopted, which evaluates the data processing results of multiple first models through a second model, including dimensions such as format, thematic expression, fluency, repetition, rhythm and rhetorical skills, and optimizes the evaluation logic by combining user feedback.
It enables multi-dimensional and accurate evaluation of model data processing results, accurately compares the performance differences of different models, selects the optimal data processing results, and continuously optimizes the evaluation system based on user feedback.
Smart Images

Figure CN2025134160_21052026_PF_FP_ABST
Abstract
Description
A model evaluation system, data processing method, apparatus, equipment, and medium
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411612753.8, filed on November 12, 2024, entitled "A Model Evaluation System, Data Processing Method, Apparatus, Device and Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of data processing technology, and more specifically, to a model evaluation system, data processing method, apparatus, device, and medium. Background Technology
[0004] Currently, numerous artificial intelligence models have emerged for performing data processing tasks. These models can generate corresponding data processing results based on user-defined requirements. The accuracy and quality of these model-generated results need to be evaluated. In related technologies, the evaluation of model data processing results is typically conducted through manual assessment or by using simple rules.
[0005] In related technologies, the methods used to evaluate the data processing results of models are relatively simple and cannot effectively evaluate the actual performance of each model. Summary of the Invention
[0006] This application provides a model evaluation system, data processing method, apparatus, device, and medium, which are intended to effectively evaluate the data processing results of a model.
[0007] The first aspect of this application provides a model evaluation system, the model evaluation system comprising:
[0008] A data processing unit is used to call at least one first model to perform a data processing task and obtain at least one data processing result, wherein the at least one first model is a different model;
[0009] The result evaluation unit is used to evaluate at least one data processing result through the second model, obtain at least one evaluation result, output the evaluation result to the user page, and use the evaluation result to compare the differences between the data processing results of different first models.
[0010] The user feedback acquisition unit is used to acquire user feedback information on the evaluation results, so that the second model can optimize the evaluation logic based on the feedback information.
[0011] In some embodiments of this application, the result evaluation unit includes:
[0012] The dimensional evaluation subunit is used to evaluate the data processing results from at least one preset dimension to obtain the evaluation results.
[0013] In some embodiments of this application, the user feedback acquisition unit includes at least one of the following:
[0014] The first user feedback acquisition subunit is used to acquire the user's evaluation selection for each evaluation result. The evaluation selection includes any of the following: evaluation too high, evaluation normal, evaluation too low.
[0015] The second user feedback acquisition subunit is used to acquire the user's evaluation content for each evaluation result. The evaluation content includes the score for each evaluation result and evaluation information.
[0016] In some embodiments of this application, the model evaluation system further includes:
[0017] The model loading unit is used to load the corresponding first model based on the model information input by the user.
[0018] The status setting unit is used to set the usage status of the first model, which includes an enabled status and a deprecated status.
[0019] The model information display unit is used to display the model information of the first model;
[0020] The parameter configuration unit is used to configure the parameters of the first model according to the model configuration parameters.
[0021] In some embodiments of this application, the model evaluation system further includes:
[0022] The file receiving unit is used to receive a task file through the first interface. The task file includes task data corresponding to the data processing task.
[0023] The result file acquisition unit is used to output the result file corresponding to the data processing task through the second interface. The result file includes the data processing results of the first model and the evaluation results of the second model.
[0024] In some embodiments of this application, the model evaluation system further includes:
[0025] The evaluation results display unit is used to display the evaluation score corresponding to each first model, as well as the evaluation reason corresponding to the evaluation score;
[0026] The image generation unit is used to calculate the average score of the evaluation scores of all the first models after all the first models have completed the data processing tasks, and to generate the corresponding score distribution map, and to mark the average score in the score distribution map.
[0027] The evaluation history display unit is used to display the data processing results corresponding to each data processing task executed by the first model, as well as the evaluation records corresponding to the data processing results. The evaluation records include the task content of the data processing task, the evaluation results corresponding to the data processing results, and the evaluation reasons.
[0028] A second aspect of this application provides a data processing method applied to a model evaluation system. The data processing method includes:
[0029] Receive task data corresponding to the data processing task;
[0030] At least one data processing result is obtained by performing a data processing task using at least one first model.
[0031] Input at least one data processing result into the second model;
[0032] The data processing results are evaluated from at least one preset dimension using a second model to obtain at least one evaluation result;
[0033] Based on the evaluation results, the target data processing result is determined from at least one data processing result.
[0034] In some embodiments of this application, before performing a data processing task through at least one first model to obtain at least one data processing result, the data processing method further includes:
[0035] Receive the model information of the first model to be loaded;
[0036] If the model address included in the model information is a local path, the first model is obtained from the storage location corresponding to the local path;
[0037] If the model address included in the model information is the program interface address, the first model is downloaded through the program interface address;
[0038] Based on the model parameter types included in the model information, input the corresponding model parameter values into the first model;
[0039] Set the first model to enabled.
[0040] In some embodiments of this application, the data processing method further includes:
[0041] The data processing results, data processing results, and corresponding evaluation reasons corresponding to the first model are stored in the evaluation record corresponding to the first model;
[0042] Upon receiving a request to view the evaluation records corresponding to the first model, the evaluation records are displayed.
[0043] In some embodiments of this application, the data processing method further includes:
[0044] Upon receiving a re-evaluation instruction, at least one data processing result is re-evaluated using a second model.
[0045] In some embodiments of this application, the data processing results are evaluated from at least one preset dimension using a second model to obtain at least one evaluation result, including:
[0046] Using preset format matching rules, the data processing results are scored from the format dimension to obtain the format score result;
[0047] By calculating the consistency score between the content and the topic of the data processing results, the data processing results are scored from the perspective of topic expression, and the topic expression score results are obtained.
[0048] The data processing results are sorted by confusion level, and the data processing results are scored from the fluency dimension to obtain the fluency score results;
[0049] Using preset repetition evaluation rules, the data processing results are scored from the repetition dimension to obtain the repetition score results;
[0050] Using pre-set rhyme judgment rules, the data processing results are scored from the prosody dimension to obtain the prosody score result;
[0051] Using a pre-set rhetoric judgment method, the data processing results are scored from the perspective of rhetoric skills to obtain rhetoric score results;
[0052] The format score, theme expression score, fluency score, repetition score, rhythm score, and rhetoric score are weighted according to a preset weight ratio to obtain the data processing result.
[0053] A third aspect of this application provides a data processing apparatus, which is applied to a model evaluation system. The data processing apparatus includes:
[0054] The task data receiving module is used to receive task data corresponding to data processing tasks.
[0055] The data processing result acquisition module is used to execute a data processing task through at least one first model and obtain at least one data processing result.
[0056] The data processing result input module is used to input at least one data processing result into the second model;
[0057] The data processing result evaluation module is used to evaluate the data processing results from at least one preset dimension using a second model, and obtain at least one evaluation result.
[0058] The data processing result determination module is used to determine the target data processing result from at least one data processing result based on the evaluation result.
[0059] In some embodiments of this application, the data processing apparatus further includes:
[0060] The model information receiving module is used to receive model information of at least one first model to be loaded;
[0061] The first model acquisition submodule is used to retrieve the first model from the storage location corresponding to the local path when the model address is a local path.
[0062] The second model acquisition submodule is used to download the first model through the program interface address when the model address is the program interface address.
[0063] The parameter input submodule is used to input the corresponding model parameter values into the first model according to the model parameter type;
[0064] The model activation module is used to set the first model to the enabled state after the first model has been loaded.
[0065] In some embodiments of this application, the data processing apparatus further includes:
[0066] The evaluation record module is used to store the data processing results, data processing results and corresponding evaluation reasons corresponding to the first model in the evaluation record corresponding to the first model;
[0067] The evaluation record display module is used to display the evaluation records when a request is received to view the evaluation records corresponding to the first model.
[0068] In some embodiments of this application, the data processing apparatus further includes:
[0069] A re-evaluation module is used to re-evaluate at least one data processing result using a second model upon receiving a re-evaluation instruction.
[0070] In some embodiments of this application, the data processing result evaluation module includes:
[0071] The first evaluation submodule is used to score the data processing results from the format dimension using preset format matching rules, and obtain the format score result;
[0072] The second evaluation submodule is used to score the data processing results from the perspective of topic expression by calculating the consistency score between the content and topic of the data processing results;
[0073] The third evaluation submodule is used to sort the data processing results by confusion level and score the data processing results from the fluency dimension to obtain the fluency score result.
[0074] The fourth evaluation submodule is used to score the data processing results from the repeatability dimension using preset repeatability evaluation rules, and obtain the repeatability score results.
[0075] The fifth evaluation submodule is used to score the data processing results from the prosody dimension using pre-set rhyme judgment rules, and obtain the prosody score result;
[0076] The sixth evaluation submodule is used to score the data processing results from the perspective of rhetoric skills using a pre-set rhetoric judgment method, and obtain the rhetoric score result.
[0077] The evaluation result acquisition submodule is used to perform weighted calculations on the format score, topic expression score, fluency score, repetition score, rhythm score, and rhetoric score according to preset weight ratios to obtain the data processing results.
[0078] In some embodiments of this application, the data processing apparatus further includes:
[0079] The first weight ratio adjustment module is used to adjust the weight ratio based on the data processing results;
[0080] The second weight ratio adjustment module is used to adjust the weight ratio based on user feedback.
[0081] The third weight ratio adjustment module is used to adjust the weight ratio by iteratively optimizing different combinations of weight ratios through a preset optimization algorithm.
[0082] A fourth aspect of this application provides a non-volatile readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method as described in the first aspect of this application.
[0083] A fifth aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method of the first aspect of this application.
[0084] Using the data processing method provided in this application, the method involves receiving task data corresponding to a data processing task; executing the data processing task through at least one first model to obtain at least one data processing result; inputting the at least one data processing result into a second model; evaluating the data processing result from at least one preset dimension through the second model to obtain at least one evaluation result; and determining the target data processing result from the at least one data processing result based on the evaluation result.
[0085] In this application, at least one first model is preloaded in the model evaluation system. Task data is input into the first model to obtain the corresponding data processing results. Then, the data processing results are evaluated in multiple dimensions by a second model to obtain the final data processing result. This method can freely load the first model that needs to be evaluated, evaluate multiple first models at the same time, and evaluate the data processing results from multiple dimensions. In this way, the performance differences of different models can be accurately grasped through the evaluation results, and the optimal data processing result can be selected. Attached Figure Description
[0086] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0087] Figure 1 is a schematic diagram of the structure of a model evaluation system proposed in an embodiment of this application;
[0088] Figure 2 is an architecture diagram of a model evaluation system proposed in an embodiment of this application;
[0089] Figure 3 is a flowchart of the model evaluation system proposed in an embodiment of this application;
[0090] Figure 4 is a system evaluation flowchart proposed in an embodiment of this application;
[0091] Figure 5 is a flowchart of a data processing method proposed in an embodiment of this application;
[0092] Figure 6 is a flowchart of model parameter setting according to an embodiment of this application;
[0093] Figure 7 is a schematic diagram of a data processing apparatus according to an embodiment of this application;
[0094] Figure 8 is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0095] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0096] Referring to Figure 1, which is a schematic diagram of the model evaluation system structure proposed in an embodiment of this application, the model evaluation system includes:
[0097] The data processing unit is used to call at least one first model to perform data processing tasks and obtain at least one data processing result. The at least one first model can be different models. The result evaluation unit is used to evaluate the at least one data processing result through the second model and obtain at least one evaluation result. The evaluation result is output to the user page and is used to compare the differences between the data processing results of different first models. The user feedback acquisition unit is used to acquire user feedback information on the evaluation results so that the second model can optimize the evaluation logic based on the feedback information.
[0098] In some embodiments of this application, the first model is an artificial intelligence model that performs data processing tasks. It can execute corresponding data processing tasks based on the user's task requirements. For example, if a user needs to generate a poem using the model, the data processing task to be performed is a poem generation task. The user inputs the corresponding task data, such as the main body and format of the poem, and the first model can generate the corresponding poem. This system can simultaneously use multiple different first models to perform data processing tasks, thereby obtaining multiple data processing results.
[0099] In some embodiments of this application, the second model is a text content evaluation model trained using Transformer (a deep learning framework) based on Natural Language Processing (NLP). Taking poetry recognition as an example, when training the second model, processed poems conforming to standard formats are used as training material and input into the second model to be trained. The parameters of the second model are adjusted to obtain a trained second model, which has the performance to evaluate and score elements such as the format, theme, and rhyme of poems, thereby making a comprehensive evaluation of a poem. This evaluation score can reflect the quality of the data processing results, and thus reflect the performance of the first model. By comparing multiple evaluation scores, the differences between the data processing results of the first model can be determined.
[0100] The results evaluation unit includes:
[0101] The dimensional evaluation subunit is used to evaluate the data processing results from at least one preset dimension to obtain the evaluation results.
[0102] In some embodiments of this application, the second model evaluates the data processing results through at least one preset dimension. Each dimension of the evaluation is equipped with a corresponding algorithm. For example, when evaluating poetry, the content of the poetry is evaluated from multiple dimensions, such as format, theme expression, fluency, repetition, rhythm, and rhetorical techniques.
[0103] The user feedback acquisition unit includes at least one of the following: a first user feedback acquisition subunit, used to acquire the user's evaluation selection for each evaluation result, the evaluation selection including any of the following: evaluation too high, evaluation normal, evaluation too low; and a second user feedback acquisition subunit, used to acquire the user's evaluation content for each evaluation result, the evaluation content including the score for each evaluation result and evaluation information.
[0104] In some embodiments of this application, users can evaluate each evaluation result on the result display page of the model evaluation system. Options include evaluation too high, evaluation normal, evaluation too low, or users can directly score the evaluation result and specify the corresponding evaluation information, including the evaluation reason, etc. The user feedback acquisition unit obtains the user's evaluation selection or evaluation information for optimizing the parameters of the second model.
[0105] The model evaluation system also includes:
[0106] The model loading unit is used to load the corresponding first model based on the model information input by the user.
[0107] The status setting unit is used to set the usage status of the first model, including the enabled status and the deactivated status; the model information display unit is used to display the model information of the first model; and the parameter configuration unit is used to configure the parameters of the first model according to the model configuration parameters.
[0108] In some embodiments of this application, users can load a first model in the model evaluation system, specifying the model name, loading path, and other information. The system can then load the corresponding first model according to the loading path. Users can also set the usage status of each model, and the model information of each model will be displayed in the model list. Users can also configure the parameters of the model independently by inputting the corresponding parameters.
[0109] The model evaluation system also includes: a file receiving unit, used to receive task files through a first interface, the task files including task data corresponding to the data processing task; and a result file acquisition unit, used to output the result file corresponding to the data processing task through a second interface, the result file including the data processing results of the first model and the evaluation results of the second model.
[0110] In some embodiments of this application, the model evaluation system can receive a task file containing task data through a preset first interface, and after performing the data processing task, it can output a result file containing the corresponding result data through a second interface. Users can provide task files and receive corresponding result files.
[0111] The model evaluation system also includes:
[0112] The evaluation result display unit displays the evaluation score for each first model and the evaluation reason for that score. The image generation unit calculates the average score of all first models after all first models have completed the data processing task, generates a corresponding score distribution map, and marks the average score on the score distribution map. The evaluation history display unit displays the data processing results for each data processing task performed by the first model, as well as the evaluation record corresponding to the data processing results. The evaluation record includes the task content of the data processing task, the evaluation result corresponding to the data processing result, and the evaluation reason.
[0113] In some embodiments of this application, the evaluation results can be recorded and statistically analyzed. Users can view the output results of each first model and second model recorded after each data processing task is executed on the corresponding page, as well as the detailed information of the evaluation results corresponding to each data processing task, including evaluation content, evaluation reasons, evaluation scores, etc. The average score of multiple evaluation results can also be calculated, and a score distribution chart can be generated. The average score is marked on the score distribution chart, so that users can intuitively see the score of the data processing results output by each model.
[0114] Referring to Figure 2, which is an architecture diagram of a model evaluation system proposed in an embodiment of this application, the front end of the model evaluation system includes a homepage, a model management page, an evaluation history page, and an API (Application Programming Interface). The homepage includes functions for displaying multiple models, model inference, model evaluation, and visualizing model evaluation scores. The model management page includes functions for model download, model configuration, model parameter settings, and model on / off. The evaluation history page includes the scores and related data for each evaluation of each first model, presented in a list format. The details page displays detailed data for each inference iteration of the first model. The inference evaluation interface in the program interface can import user-defined content to be evaluated, and the export interface outputs the corresponding evaluation scores. The model evaluation system relies on multiple first models for multi-model inference, on second models (i.e., the source model) for scoring, and on database-provided functions.
[0115] Referring to Figure 3, which is a flowchart of the model evaluation system proposed in an embodiment of this application, when a user uses the model evaluation system, the homepage will be opened by default. The homepage mainly contains functions such as inputting inference content and setting and displaying the second model evaluation content. Users can switch to the model settings page to perform model-related operations, such as downloading models, setting parameters, enabling / disabling models, etc. Users can view the evaluated content and corresponding details on the historical evaluation page.
[0116] When a user enters the homepage, the system displays the models the user has set up for inference. The system will then use these models for inference. The user can also set the evaluation method, such as automatic or manual evaluation. In automatic evaluation, the system will automatically evaluate the data processing results using the second model after all the first models have completed inference. In manual evaluation, the user clicks the evaluation button on the homepage. For self-evaluation, after all models have completed inference, the user fills in the evaluation information (score) below each model and submits it to the second model for the larger model's reference.
[0117] After entering the model management page, a list of models is displayed, including downloaded models, local models, and online models. Online models will display status information such as pending download, downloading, and download complete, as well as the download path and parameter configuration information. The local model list displays relevant information such as the local path, online path, and parameter configuration. Online models will display the online address and parameter configuration information. Both downloaded and local models require model parameter configuration.
[0118] After entering the historical evaluation page, the system will display the evaluation records of each first model in the evaluation list according to the evaluation time order. This includes model information of the models that have been evaluated. You can click on the details to enter the evaluation details page, which can display the reasoning content, data processing results, and user self-evaluation information.
[0119] The model evaluation system also supports user feedback, allowing users to rate the evaluation score of the second model. Based on user feedback, the second model optimizes its evaluation logic to better meet user requirements. The second model determines whether its performance evaluation score is high, low, or normal based on user ratings, and it can further optimize and adjust its parameters according to user feedback scores to ensure the evaluation score better meets user needs.
[0120] Referring to Figure 4, which is a system evaluation flowchart proposed in an embodiment of this application, when a user operates the system, they first add the model to be used on the model configuration page, enter the name of the configured model, the local path (required for local model calls), the API (Application Programming Interface) address (required for online model calls), the model parameters and corresponding values, and save the configuration. At this time, the system will enable the model and make interface calls using the configured address and parameters.
[0121] Next, select the model you want to use, enable or disable the model on the model configuration page, and after configuration, click the Start Inference button to perform inference and display the data processing results on the homepage.
[0122] The homepage is the main page for user-input inference evaluation. It includes controls for user input of inference content, multi-model data processing result display, large-model evaluation methods, large-model evaluation scoring results, large-model evaluation reasons, and user self-evaluation. Users input the content to be inferred in the inference content control, select the evaluation method, and click the "Start Inference" control. The system then begins inference. Multiple models perform inference simultaneously, displaying the results in real-time in the multi-model data processing result display until all models have completed inference.
[0123] After all the first models have completed inference, the system automatically evaluates the data processing results based on the instructions received from the large model evaluation method control, or waits for the user to click the manual evaluation button. After the evaluation is completed, the user can click the re-evaluation button to re-evaluate the data processing results. The system's autonomous evaluation control can receive custom content sent by the user, call the source large model to evaluate the custom content, and obtain the corresponding evaluation score.
[0124] After evaluation, the visualization interface will display the evaluation score and corresponding evaluation reasons below the data processing results of each model. Finally, an average evaluation score and a score distribution chart are generated to intuitively display the score rankings. Users can go to the evaluation history page to view all evaluation records, which will display information such as inference content, average inference score, and score for each model. Users can also use a local API to automatically perform inference and evaluation on local files. The local file content contains all the content the user wants to infer. Then, through the API, the final multi-model data processing results and evaluation scores are obtained.
[0125] Referring to Figure 5, which is a flowchart of a data processing method according to an embodiment of this application, the data processing method is applied to a model evaluation system. As shown in Figure 5, the data processing method includes the following steps:
[0126] S11: Receive task data corresponding to the data processing task.
[0127] In some embodiments of this application, at least one first model is pre-loaded in the model evaluation system. These first models may be obtained locally or downloaded from a model library on the network via a program interface. Task data corresponding to the data processing task is received through the pre-loaded at least one first model.
[0128] For example, the first model is a language model that can generate verses, and the task data is "generate a five-character quatrain about the moon".
[0129] S12: Perform a data processing task through at least one first model to obtain at least one data processing result.
[0130] In some embodiments of this application, the data processing result is the output obtained by the first model based on the input task data.
[0131] In some embodiments of this application, multiple first models simultaneously receive task data and generate corresponding data processing results based on the task data. These first models are all pre-trained models that process the input task data based on parameters learned during training to obtain the corresponding data processing results.
[0132] For example, the first model will generate a five-character quatrain containing the moon or the element of the moon based on the input task data "generate a five-character quatrain about the moon".
[0133] S13: Input at least one data processing result into the second model.
[0134] In some embodiments of this application, when multiple data processing results are obtained from multiple first models, the multiple data processing results are respectively input into the second model.
[0135] S14: Evaluate the data processing results from at least one preset dimension using a second model to obtain at least one evaluation result.
[0136] In some embodiments of this application, the evaluation result is an evaluation score obtained after evaluating the data processing results of the first model.
[0137] In some embodiments of this application, the input data processing results are evaluated in multiple dimensions using a second model to obtain evaluation results. Multiple aspects need to be considered during the evaluation, and thus multiple dimensions are set to evaluate the data processing results.
[0138] In some embodiments of this application, taking poetry as an example, the second model is mainly based on Natural Language Processing (NLP) technology. It uses Transformer (a deep learning framework) technology to train the model, making it capable of understanding the language structure and style of poetry.
[0139] In some embodiments of this application, taking poetry evaluation as an example, most current poetry evaluation methods employ BLEU (bilingual evaluation understudy) with certain rules. However, rule-based evaluation alone can only assess format and cannot achieve other evaluation effects. While BLEU is fast, language-insensitive, and has better applicability, it only considers the similarity of N-grams and lacks semantic-level evaluation.
[0140] In some embodiments of this application, the content of poems is evaluated from multiple dimensions, including format, thematic expression, fluency, repetition, rhythm, and rhetorical techniques.
[0141] In some embodiments of this application, the specific steps for obtaining data processing results by scoring the data processing results using at least one preset evaluation dimension include:
[0142] S14-1: Using preset format matching rules, score the data processing results from the format dimension to obtain the format score result.
[0143] In some embodiments of this application, different inference content has different format requirements. Matching rules can be set according to the actual situation. The preset format matching rules are used to score the data processing results from the format dimension to obtain the format score results. When building the second model, the set format matching rules are added to the model. During the training process, the second model will adjust the corresponding parameters of the format matching rules, thereby continuously improving the format matching rules.
[0144] In some embodiments of this application, taking poetry as an example, the second model supports the recognition of formats including five-character quatrains, seven-character quatrains, five-character regulated poems, and seven-character regulated poems. It also supports various ci (lyric poetry) titles such as *Zhegutian*, *Huanxisha*, *Manjianghong*, *Pusa Man*, *Dielianhua*, *Linjiangxian*, *Shuidiaogetou*, *Qingpingyue*, and *Niannujiao*. Furthermore, it supports different character counts for the same ci title, enabling better format verification and evaluation of different types of poetry. The model uses rule matching, where the content is segmented based on punctuation marks. It then determines whether each segment conforms to the corresponding format and character count. If it does, the data processing result is 100 points; otherwise, it is 0 points. For example, if each line of a five-character quatrain has five characters, and each line in the data processing result has five characters, it is considered a perfect score.
[0145] The corresponding evaluation formula is:
[0146] Where format_score represents the format score, f(x) represents the final score, and weight represents the weight. As can be seen from the formula, the score result is divided into two parts. If format_score is 0, the overall score is 0. If format_score is not 0, the final score is the sum of the scores of each dimension multiplied by the scores of each dimension.
[0147] S14-2: By calculating the consistency score between the content and the topic of the data processing results, the data processing results are scored from the topic expression dimension to obtain the topic expression score result.
[0148] In some embodiments of this application, the consistency score between the content and the topic of the data processing result is also called the answer fidelity score. It is an evaluation index obtained using LLM (Large Language Model) to evaluate whether the topic of a content meets the proposed requirements. The second model integrates the language recognition function of the large language model and continuously adjusts the parameters for content recognition during the training process to accurately obtain the answer fidelity score for each inference content.
[0149] In some embodiments of this application, the data processing results are scored from the topic expression dimension by calculating the answer fidelity score of the data processing results, thereby obtaining the corresponding topic expression score results.
[0150] In some embodiments of this application, a second model is used to decompose the poetry composition result into multiple statements, and the consistency of each statement with the theme is examined. A "fidelity score" is calculated based on the ratio of the number of supported statements to the total number of statements. The score is awarded in an additive form, with 5 points added for each supported statement, up to a maximum of 100 points.
[0151] Here, `score` represents the final score, `w` represents the weight, `k` represents the number of sums, and `y` represents the score added for a match, where `y` is always equal to 5. The weighted sum is compared to 100, and the minimum is taken. Finally, the score is obtained by multiplying by the weight `w`.
[0152] In some embodiments of this application, the second model fine-tunes the model to achieve a deeper fit to the theme. By using model fine-tuning, the content of the poem is compared with the user's questions and answers to evaluate the degree of relevance between the content and the theme. This includes the theme words and words related to the theme words. For example, if the theme is "Write a five-character quatrain with the Mid-Autumn Festival as the theme", then the theme word is "Mid-Autumn Festival". However, if the title contains "moon" or "Chang'e", it will also be evaluated as theme relevance.
[0153] S14-3: Sort the data processing results by perplexity and score the data processing results from the fluency dimension to obtain the fluency score results.
[0154] In some embodiments of this application, the basic idea of perplexity ranking (PPL) is to calculate the probability of a sentence appearing. The higher the score, the lower the probability of its appearance. That is, the higher the score, the lower the probability of the sequence appearing, and the more difficult it is to produce fluent text. The lower the score, the easier it is to produce fluent text. Based on this idea, the fluency of the poem is calculated.
[0155] In some embodiments of this application, the probability of a sentence appearing is calculated and raised to the power of 1 / N, as shown in the following formula:
[0156] Where w1 w2…wN represents N tokens in a sentence (tokens correspond to the dictionary in the language model; for Chinese, they may be a character or a word), and P is the probability of generating each word.
[0157] As can be seen from the above formula, when using PPL sorting analysis to analyze a poem, the lower the score, the more likely the poem is to be a fluent text and the higher its fluency.
[0158] The scoring formula is as follows:
[0159] score = w * PPL (4)
[0160] Where score is the final score, w represents the weight, and PPL is the result calculated by the above method. Finally, multiply by the weight w to obtain the final score.
[0161] In some embodiments of the present application, fluency mainly uses a formula form to sort the final results by perplexity, and judges the fluency of the data processing results according to the evaluation criteria of the formula, greatly improving the evaluation efficiency.
[0162] S14-4: Use a preset repetition evaluation rule to score the data processing results from the repetition dimension to obtain a repetition score result.
[0163] In some embodiments of the present application, a preset repetition evaluation rule is used to score the data processing results from the repetition dimension to obtain a repetition score result. When constructing the second model, the repetition evaluation rule is added, and the relevant parameters of the repetition evaluation rule can be iteratively adjusted during the training of the model, so as to better process and analyze the poems through this rule.
[0164] In some embodiments of the present application, the repetition degree is evaluated by a rule. The poem content is split character by character, and then the repetition degree is judged. The higher the repetition degree, the lower the score. This evaluation dimension is in the form of deduction. The starting score is 100 points, and 5 points are deducted for each repetition. The highest score is 100 points, and the lowest score is 0 points. The formula is:
[0165] Where score is the final score, w represents the weight, k represents the number to be summed, y represents the score of the repeated characters, where y = 5 is a constant value. Compare 100 points minus the qualified score with 0, and take the larger one as the final result, and multiply by the weight w to obtain the final score.
[0166] In some embodiments of the present application, the repetition degree is a measurement of the repetition rate of the poem content of the final result. For a poem, if there is too much repeated content, the result may be rather rigid. Of course, it does not rule out that some good poems use repeated content, such as "唧唧复唧唧". Therefore, the method of subtracting the repeated number from the total score is adopted, and the repeated content is subtracted to obtain the final result, which not only judges the repeated number, but also optimizes the result in a form from high to low, avoiding the situation of overly weakening the repeated words.
[0167] S14-5: Use a preset rhyme judgment rule to score the data processing results from the rhyme dimension to obtain a rhyme score result.
[0168] In some embodiments of this application, a pre-set rhyme judgment rule is used to score the data processing results from the prosody dimension to obtain a prosody score result. The second model incorporates the rhyme judgment rule during construction, and can continuously optimize the relevant parameters as the model is trained, so as to more accurately perceive and evaluate the prosody and rhythm of a text content.
[0169] In some embodiments of this application, a Chinese character-to-pinyin conversion method is used to convert the last character of each sentence into pinyin and determine whether it rhymes. A score is assigned based on the degree of rhyme. The higher the proportion of rhyming words and phrases, the higher the score. The formula is as follows:
[0170] Where score is the final score, w represents the weight, n represents the total number of rhymes, t is the maximum number of rhymes for this poetry type, and multiplying by 100 is to calculate the ratio and take values within 100 to maintain a consistent unit with other dimensions for statistical purposes. The final score is obtained by multiplying the score by the weight w.
[0171] In some embodiments of this application, the second model uses a pinyin recognition method to compare the pinyin of the last character to determine whether it rhymes, which can better evaluate the rhythm.
[0172] S14-6: Using a pre-set rhetoric judgment method, score the data processing results from the perspective of rhetoric skills to obtain the rhetoric score results.
[0173] In some embodiments of this application, when generating articles or poems, it is necessary to determine the rhetorical devices used. A pre-set rhetorical judgment method is used to score the data processing results from the perspective of rhetorical skills, and a rhetorical score result is obtained. During the training process, the second model will continuously adjust the parameters corresponding to the rhetorical judgment method so that the model can more accurately determine the rhetorical devices used in the text content.
[0174] In some embodiments of this application, the second model translates the input poem into modern Chinese, and then judges the rhetorical skills of the poem by determining whether the modern Chinese contains rhetorical devices such as metaphor and personification, thereby assigning a score. The more rhetorical devices used, the higher the score. Each rhetorical device adds 5 points, up to a maximum of 100 points. The formula is as follows:
[0175] Where score is the final score, w represents the weight, k represents the number of sums, and y represents the score added for rhetorical devices, where y = 5. The result is multiplied by the weight w to obtain the final score.
[0176] In some embodiments of this application, the second model uses the method of translating ancient poems into modern text to determine the rhetorical devices. This method simplifies the judgment logic and increases the accuracy of the judgment. The large model is trained to judge different rhetorical devices through fine-tuning. It can judge various devices such as metaphor, parallelism, and personification. By judging and evaluating different devices, the dimensions of the poem evaluation are increased.
[0177] S14-7: The format score, theme expression score, fluency score, repetition score, rhythm score, and rhetoric score are weighted according to a preset weight ratio to obtain the data processing result.
[0178] In some embodiments of this application, the format score, topic expression score, fluency score, repetition score, rhythm score, and rhetoric score are weighted and calculated according to preset weight ratios to obtain the data processing result. The data processing result is a comprehensive score.
[0179] In some embodiments of this application, the second model supports specifying the weight of each evaluation dimension in the form of parameters. For example, if the theme expression is the most important, it can be set to the highest weight, with 100 points being the highest score and 0 points being the lowest. Currently, the sum of all parameter weights is 100. The system will first select the dimension with the highest weight, and then select the dimension with the next highest weight. If the sum of the preceding parameter settings reaches 100, the subsequent parameters will be automatically set to 0. If there are dimensions with the same weight, they will be selected according to the system's preset order. For example, for fluency and repetition, the system sets fluency to have a higher priority than repetition. If these two weights are set the same, the system will prioritize fluency for weight summation. If there are remaining weights, then repetition will be allocated.
[0180] In some embodiments of this application, the initial weighting ratio is based on expert experience and opinions. Initially, the expert opinions are: topic expression: 40%, fluency: 25%, repetition: 15%, rhythm: 10%, and rhetorical skills: 10%.
[0181] In some embodiments of this application, the method further includes:
[0182] S14-8: Adjust the weight ratio based on the data processing results.
[0183] In some embodiments of this application, a feedback loop is established. After obtaining the data processing results, the model performance model adjusts the proportion of the scoring criteria by automatically feeding back the evaluation score. The feedback loop can be implemented by methods such as reinforcement learning or genetic algorithms. Through these algorithms, the model can gradually optimize its internal parameters based on the evaluation score.
[0184] S14-9: Adjust the weight ratio based on user feedback.
[0185] In some embodiments of this application, user feedback information refers to the user's autonomous rating of the data processing results of each first model, and also includes the user's rating of the second model.
[0186] In some embodiments of this application, user feedback information is introduced to further adjust the proportion of the scoring criteria. Users can score the data processing results or provide other forms of feedback, and the user feedback information guides the parameter optimization of the second model.
[0187] S14-10: By using a preset optimization algorithm, iteratively optimize different combinations of weight ratios and adjust the weight ratios.
[0188] In some embodiments of this application, hyperparameter optimization techniques can be used to automatically adjust the weight ratios to maximize the overall score. For example, a Bayesian model can be used for optimization, and a corresponding acquisition function can be designed to iteratively optimize the model. Different ratio combinations can be automatically tried within a given search space to find the optimal solution.
[0189] In some embodiments of this application, the initial weighting of each dimension is based on the expert opinion rating ratio. Then, hyperparameter optimization techniques (such as Bayesian optimization) are used to find an optimized set of weighting ratios. These optimized weighting ratios are then input into a second model, and further fine-tuned through a feedback loop (such as reinforcement learning). The feedback loop can automatically adjust the ratios based on the performance of the poems generated by the model in real or simulated environments to achieve better results.
[0190] In addition to the automatic feedback loop, user feedback is introduced to guide the optimization process. Users can rate the poems generated by the model or provide other forms of feedback. This user feedback serves as additional input data to adjust the optimization objectives or weights in the feedback loop. For example, if most users feel that the rhetoric of a poem is not good enough, the system can automatically increase the weight of the rhetoric score and adjust the weights of other parts accordingly.
[0191] For example, after continuous optimization, the resulting weight ratios are: topic expression: 38.3%, fluency: 30.5%, repetition: 10.6%, rhythm: 11.2%, and rhetorical skills: 9.4%.
[0192] S15: Based on the evaluation results, determine the target data processing result from at least one data processing result.
[0193] In some embodiments of this application, after obtaining the evaluation result corresponding to each data processing result, the evaluation result with the highest score is determined from multiple data processing results and used as the target data processing result.
[0194] For example, when the data processing task is to generate a poem, the poem with the highest score is selected as the poem to be used, and the first model that generates the poem is determined as the optimal model.
[0195] In some embodiments of this application, a comprehensive optimization framework is constructed that integrates multiple optimization strategies such as hyperparameter optimization, feedback loops, and user feedback. Different optimization stages and priorities are set within this framework. By continuously optimizing the weights of each dimension of the evaluation text of the second model through multiple strategies, the second model makes a more accurate evaluation of the data processing results, effectively improving the efficiency and accuracy of the second model's evaluation.
[0196] In another embodiment of this application, before receiving task data via at least one pre-loaded first model, the method further includes:
[0197] S21: Receive the model information of the first model to be loaded.
[0198] In some embodiments of this application, the model information includes the model name, local path, program interface address, model parameters, and model parameter values.
[0199] In some embodiments of this application, the system first receives model information of at least one first model to be loaded.
[0200] In some embodiments of this application, the model information includes the model name, model address, model parameter type, and model parameter value.
[0201] S22: If the model address included in the model information is a local path, retrieve the first model from the storage location corresponding to the local path.
[0202] In some embodiments of this application, if the model address is a local path, it means that the model is stored locally, so the first model can be obtained from the storage location corresponding to the local path.
[0203] S23: If the model address included in the model information is the program interface address, download the first model through the program interface address.
[0204] In some embodiments of this application, when the model address is the program interface address, the first model needs to be downloaded from the program interface address.
[0205] S24: Based on the model parameter types included in the model information, input the corresponding model parameter values into the first model.
[0206] In some embodiments of this application, the model parameter type specifies the specific type of the model parameters, and the model parameter value is the specific numerical value of the parameter.
[0207] In some embodiments of this application, the model information includes the type of model parameters and the corresponding model parameter values. Based on the type of model parameters, the corresponding model parameter values are loaded into the first model, and the first model is loaded into the model evaluation system.
[0208] S25: Set the first model to enabled.
[0209] In some embodiments of this application, the first model is set to the startup state after the first model has been loaded.
[0210] In some embodiments of this application, it is also necessary to set model parameters. Referring to Figure 6, which is a flowchart of model parameter setting according to an embodiment of this application, as shown in Figure 6, the user selects the model whose parameters need to be set on the page, fills in the request header information, the specific parameters of the request, and the URL (network address) request parameters, and saves the information. The system verifies the parameters, and if the parameters are correct, stores the filled parameters in the database. When using the model, the stored parameters are called and loaded into the model. For model-related parameters, the following parameter data tables are designed. The fields of the main parameters are shown in Table 1 and Table 2. Table 1 is the model parameter type table, and Table 2 is the model parameter value table.
[0211] Table 1
[0212] Table 2
[0213] In some embodiments of this application, users can load the corresponding model into the system according to their needs and set the corresponding parameters accordingly, which can achieve flexible expansion of the first model and meet the user's needs.
[0214] In another embodiment of this application, after obtaining the performance evaluation score, the average score of all the data processing results is calculated.
[0215] For example, if the performance evaluation scores of the three data processing results are 95, 96, and 97 respectively, then the average score is 96.
[0216] In some embodiments of this application, after calculating the average score, a score distribution chart of these data processing results is generated. The average score of these performance evaluation scores is marked on the score distribution chart, which can intuitively show whether the performance of the first model is higher than the average.
[0217] In some embodiments of this application, by generating average scores and score distribution charts, the scores of each model can be seen intuitively, which helps users to grasp the specific performance of each first model.
[0218] In another embodiment of this application, the data processing method further includes:
[0219] S31: Store the data processing results, data processing results and corresponding evaluation reasons corresponding to the first model in the evaluation record corresponding to the first model.
[0220] In some embodiments of this application, for each first model, the data processing results of the model, the data processing results, and the corresponding evaluation reasons are stored in the evaluation record corresponding to the first model, so that users can view them at any time.
[0221] S32: Upon receiving a request to view the evaluation record corresponding to the first model, display the evaluation record.
[0222] In some embodiments of this application, the evaluation records are displayed when the system receives a request to view the evaluation records of the first model.
[0223] In some embodiments of this application, users can view the evaluation records at any time to determine the overall performance of each first model and gain a global understanding of the model's performance.
[0224] In another embodiment of this application, the data processing method further includes:
[0225] S41: Upon receiving a re-evaluation instruction, re-evaluate at least one data processing result using a second model.
[0226] In some embodiments of this application, when the system receives a re-evaluation instruction, the model evaluation system re-evaluates the data processing results using a second model. Since parameters can be adjusted through feedback loops during the evaluation process, the scores may differ each time, and multiple evaluations have greater reference value.
[0227] In some embodiments of this application, the user can issue a re-evaluation command at any time to re-evaluate the evaluation score. The user can grasp the overall performance status of each first model based on the results of multiple evaluations.
[0228] In another embodiment of this application, the system statistically analyzes the evaluation scores of the second models. If the score of a first model is lower than a preset score threshold in multiple evaluations, the first model is marked as a model not recommended for use, to remind the user that the first model has poor performance and cannot provide good service to the user. Alternatively, the system can sort the scores of each first model, and if a first model has a low ranking in multiple rankings, it is marked as a model not recommended. In this way, the system can help users filter out high-performance first models and eliminate low-performance first models, allowing users to quickly find high-performance models and improve the user experience.
[0229] Based on the same inventive concept, one embodiment of this application provides a data processing apparatus applied to a model evaluation system. Referring to FIG7, FIG7 is a schematic diagram of a data processing apparatus 700 according to an embodiment of this application. As shown in FIG7, the apparatus includes:
[0230] The task data receiving module 701 is used to receive task data corresponding to the data processing task.
[0231] The data processing result acquisition module 702 is used to perform a data processing task through at least one first model and obtain at least one data processing result.
[0232] The data processing result input module 703 is used to input at least one data processing result into the second model;
[0233] The data processing result evaluation module 704 is used to evaluate the data processing result from at least one preset dimension through a second model, and obtain at least one evaluation result;
[0234] The data processing result determination module 705 is used to determine the target data processing result from at least one data processing result based on the evaluation result.
[0235] In some embodiments of this application, the apparatus further includes:
[0236] The model information receiving module is used to receive model information of at least one first model to be loaded;
[0237] The first model acquisition submodule is used to retrieve the first model from the storage location corresponding to the local path when the model address is a local path.
[0238] The second model acquisition submodule is used to download the first model through the program interface address when the model address is the program interface address.
[0239] The parameter input submodule is used to input the corresponding model parameter values into the first model according to the model parameter type;
[0240] The model activation module is used to set the first model to the enabled state after the first model has been loaded.
[0241] In some embodiments of this application, the apparatus further includes:
[0242] The evaluation record module is used to store the data processing results, data processing results and corresponding evaluation reasons corresponding to the first model in the evaluation record corresponding to the first model;
[0243] The evaluation record display module is used to display the evaluation records when a request is received to view the evaluation records corresponding to the first model.
[0244] In some embodiments of this application, the apparatus further includes:
[0245] A re-evaluation module is used to re-evaluate at least one data processing result using a second model upon receiving a re-evaluation instruction.
[0246] In some embodiments of this application, the data processing result evaluation module includes:
[0247] The first evaluation submodule is used to score the data processing results from the format dimension using preset format matching rules, and obtain the format score result;
[0248] The second evaluation submodule is used to score the data processing results from the perspective of topic expression by calculating the consistency score between the content and topic of the data processing results;
[0249] The third evaluation submodule is used to sort the data processing results by confusion level and score the data processing results from the fluency dimension to obtain the fluency score result.
[0250] The fourth evaluation submodule is used to score the data processing results from the repeatability dimension using preset repeatability evaluation rules, and obtain the repeatability score results.
[0251] The fifth evaluation submodule is used to score the data processing results from the prosody dimension using pre-set rhyme judgment rules, and obtain the prosody score result;
[0252] The sixth evaluation submodule is used to score the data processing results from the perspective of rhetoric skills using a pre-set rhetoric judgment method, and obtain the rhetoric score result.
[0253] The evaluation result acquisition submodule is used to perform weighted calculations on the format score, topic expression score, fluency score, repetition score, rhythm score, and rhetoric score according to preset weight ratios to obtain the data processing results.
[0254] In some embodiments of this application, the apparatus further includes:
[0255] The first weight ratio adjustment module is used to adjust the weight ratio based on the data processing results;
[0256] The second weight ratio adjustment module is used to adjust the weight ratio based on user feedback.
[0257] The third weight ratio adjustment module is used to adjust the weight ratio by iteratively optimizing different combinations of weight ratios through a preset optimization algorithm.
[0258] Based on the same inventive concept, another embodiment of this application provides a non-volatile readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the data processing method of any of the above embodiments of this application.
[0259] Based on the same inventive concept, another embodiment of this application provides an electronic device. FIG8 is a schematic diagram of an electronic device 800 proposed in an embodiment of this application, including a memory 802, a processor 801 and a computer program stored in the memory and executable on the processor. When executed by the processor, the steps in the data processing method of any of the above embodiments of this application are implemented.
[0260] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0261] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0262] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0263] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0264] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0265] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0266] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0267] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.
[0268] The above provides a detailed description of the model evaluation system, data processing method, apparatus, device, and medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A model evaluation system, characterized by, The system includes: A data processing unit is used to call at least one first model to perform a data processing task and obtain at least one data processing result, wherein the at least one first model is a different model; The result evaluation unit is used to evaluate the at least one data processing result using the second model, obtain at least one evaluation result, and output the evaluation result to the user page. The evaluation result is used to compare the differences between the data processing results of different first models. The user feedback acquisition unit is used to acquire user feedback information on the evaluation results, so that the second model can optimize the evaluation logic based on the feedback information.
2. The model evaluation system of claim 1, wherein, The result evaluation unit includes: The dimension evaluation subunit is used to evaluate the data processing results from at least one preset dimension to obtain the evaluation results.
3. The model evaluation system of claim 1, wherein, The user feedback acquisition unit includes at least one of the following: The first user feedback acquisition subunit is used to acquire the user's evaluation selection for each of the evaluation results, and the evaluation selection includes any of the following: evaluation too high, evaluation normal, evaluation too low; The second user feedback acquisition subunit is used to acquire the user's evaluation content for each of the evaluation results, the evaluation content including the score and evaluation information for each evaluation result.
4. The model evaluation system of claim 1, wherein, The system also includes: The model loading unit is used to load the corresponding first model based on the model information input by the user. A status setting unit is used to set the usage status of the first model, the usage status including an enabled status and a deprecated status; A model information display unit is used to display the model information of the first model; The parameter configuration unit is used to configure the parameters of the first model according to the model configuration parameters.
5. The model evaluation system of claim 4, wherein, The model configuration parameters include: request header parameters, request body parameters, and address request parameters.
6. The model evaluation system of claim 1, wherein, The system also includes: The file receiving unit is used to receive a task file through a first interface, wherein the task file includes task data corresponding to the data processing task. The result file acquisition unit is used to output the result file corresponding to the data processing task through the second interface. The result file includes the data processing result of the first model and the evaluation result of the second model.
7. The model evaluation system of claim 1, wherein, The system also includes: The evaluation result display unit is used to display the evaluation score corresponding to each of the first models, and the evaluation reason corresponding to the evaluation score; An image generation unit is configured to calculate the average score of the evaluation scores corresponding to all the first models after all the first models have completed the data processing task, and generate a corresponding score distribution map, and mark the average score in the score distribution map. The evaluation history display unit is used to display the data processing results corresponding to each data processing task executed by the first model, and the evaluation record corresponding to the data processing results. The evaluation record includes the task content of the data processing task, the evaluation result corresponding to the data processing result, and the evaluation reason.
8. A data processing method, characterized by, The method is based on the model evaluation system according to any one of claims 1 to 7, and includes: Receive task data corresponding to the data processing task; The data processing task is executed by at least one first model to obtain at least one data processing result; Input the at least one data processing result into the second model; The data processing results are evaluated from at least one preset dimension using the second model to obtain at least one evaluation result; Based on the evaluation results, a target data processing result is determined from the at least one data processing result.
9. The data processing method according to claim 8, characterized in that, Before obtaining at least one data processing result by performing the data processing task through at least one first model, the method further includes: Receive the model information of the first model to be loaded; If the model address included in the model information is a local path, the first model is obtained from the storage location corresponding to the local path; If the model address included in the model information is a program interface address, the first model is downloaded through the program interface address; According to the model parameter types included in the model information, the corresponding model parameter values are input into the first model; Set the first model to enabled.
10. The data processing method according to claim 8, characterized in that, The method further includes: The data processing results corresponding to the first model, the data processing results, and the corresponding evaluation reasons are stored in the evaluation record corresponding to the first model. Upon receiving a request to view the evaluation record corresponding to the first model, the evaluation record is displayed.
11. The data processing method of claim 8, wherein, The method further includes: Upon receiving a re-evaluation instruction, the at least one data processing result is re-evaluated using the second model.
12. The data processing method of claim 8, wherein, The evaluation result is an evaluation score obtained after evaluating the data processing results of the first model; determining the target data processing result from the at least one data processing result based on the evaluation result includes: The highest evaluation score is determined from the at least one evaluation score, and the data processing result corresponding to the highest evaluation score is taken as the target data processing result.
13. The data processing method of claim 8, wherein, The step of evaluating the data processing results from at least one preset dimension using the second model to obtain at least one evaluation result includes: Using preset format matching rules, the data processing results are scored from the format dimension to obtain a format score result; By calculating the consistency score between the content and the topic of the data processing results, the data processing results are scored from the perspective of topic expression, and a topic expression score result is obtained. The data processing results are sorted by confusion level, and the data processing results are scored from the fluency dimension to obtain a fluency score result; Using preset repetition evaluation rules, the data processing results are scored from the repetition dimension to obtain repetition score results; Using pre-set rhyme judgment rules, the data processing results are scored from the prosody dimension to obtain a prosody score result; Using a pre-set rhetoric judgment method, the data processing results are scored from the perspective of rhetoric skills to obtain rhetoric score results; The format score, the theme expression score, the fluency score, the repetition score, the rhythm score, and the rhetoric score are weighted and calculated according to a preset weight ratio to obtain the data processing result.
14. The data processing method according to claim 13, characterized in that, The process involves calculating a consistency score between the content and the topic of the data processing results, and then scoring the data processing results from the perspective of topic expression to obtain a topic expression score, including: The data processing results are decomposed into multiple statements using a second model; The consistency between each statement and the topic corresponding to the data processing result is examined to obtain the number of statements that are consistent with the topic corresponding to the data processing result. The topic expression score is calculated based on the ratio between the number of statements that correspond to the topic and the total number of statements.
15. The data processing method according to claim 13, characterized in that, The method of using a pre-set rhetoric judgment to score the data processing results from the perspective of rhetoric skills yields a rhetoric score result, including: The data processing results are translated into modern text using a second model; Determine the number of rhetorical devices used in the aforementioned modern text; The rhetorical score is determined based on the number of times the rhetorical techniques are used.
16. The data processing method of claim 13, wherein, The method further includes: Based on the evaluation results, the weighting ratios are adjusted. The weight ratios are adjusted based on user feedback. The weight ratio is adjusted by iteratively optimizing different combinations of the weight ratios using a preset optimization algorithm.
17. The data processing method according to claim 16, characterized in that, The user feedback information includes the user's self-rated score for the data processing results of each of the first models, and the user's score for the second model.
18. A data processing apparatus, characterized by The device includes: The task data receiving module is used to receive task data corresponding to data processing tasks. A data processing result acquisition module is used to perform a data processing task through the at least one first model to obtain at least one data processing result; A data processing result input module is used to input the at least one data processing result into the second model; The data processing result evaluation module is used to evaluate the data processing result from at least one preset dimension using the second model, and obtain at least one evaluation result; A data processing result determination module is used to determine a target data processing result from the at least one data processing result based on the evaluation result.
19. A computer non-volatile readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 8 to 17.
20. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 8 to 17.