A model evaluation system, data processing method, apparatus, equipment, and medium

By using a multi-dimensional evaluation system and user feedback optimization logic, the problem of single evaluation of model data processing results in existing technologies has been solved, enabling accurate evaluation of model performance and selection of the optimal result.

CN119719705BActive Publication Date: 2026-04-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies use a single method to evaluate the results of model data processing, which cannot effectively assess the actual performance of each model.

Method used

A multi-dimensional evaluation system is adopted, which evaluates the data processing results of multiple first models through a second model, including evaluation of dimensions such as format, thematic expression, fluency, repetition, rhythm and rhetorical skills, and optimizes the evaluation logic by combining user feedback.

Benefits of technology

It enables multi-dimensional and accurate evaluation of model data processing results, accurately compares the performance differences of different models, and selects the optimal data processing result.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719705B_ABST
    Figure CN119719705B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing technology, specifically to a model evaluation system, data processing method, apparatus, device, and medium, aimed at effectively evaluating the performance of a model. The system includes: a data processing unit, used to invoke at least one first model to execute a data processing task and obtain at least one data processing result, wherein the at least one first model is a different model; a result evaluation unit, used to evaluate the at least one data processing result using a second model, obtaining at least one evaluation result, and outputting the evaluation result to a user page, wherein the evaluation result is used to compare the differences between the data processing results of different first models; and a user feedback acquisition unit, used to acquire user feedback information on the evaluation results, so that the second model can optimize the evaluation logic based on the feedback information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a model evaluation system, data processing method, apparatus, device, and medium. Background Technology

[0002] Currently, numerous artificial intelligence models have emerged for performing data processing tasks. These models can generate corresponding data processing results based on user-defined requirements. The accuracy and quality of these model-generated results need to be evaluated. In related technologies, the evaluation of model data processing results is typically conducted through manual assessment or by using simple rules.

[0003] In related technologies, the methods used to evaluate the data processing results of models are relatively simple and cannot effectively evaluate the actual performance of each model. Summary of the Invention

[0004] This application provides a model evaluation system, data processing method, apparatus, device, and medium, which are intended to effectively evaluate the data processing results of a model.

[0005] A first aspect of this application provides a model evaluation system, the system comprising:

[0006] A data processing unit is used to call at least one first model to perform a data processing task and obtain at least one data processing result, wherein the at least one first model is a different model;

[0007] The result evaluation unit is used to evaluate the at least one data processing result using the second model, obtain at least one evaluation result, and output the evaluation result to the user page. The evaluation result is used to compare the differences between the data processing results of different first models.

[0008] The user feedback acquisition unit is used to acquire user feedback information on the evaluation results, so that the second model can optimize the evaluation logic based on the feedback information.

[0009] Optionally, the result evaluation unit includes:

[0010] The dimension evaluation subunit is used to evaluate the data processing results from at least one preset dimension to obtain the evaluation results.

[0011] Optionally, the user feedback acquisition unit includes at least one of the following:

[0012] The first user feedback acquisition subunit is used to acquire the user's evaluation selection for each of the evaluation results, and the evaluation selection includes any of the following: evaluation too high, evaluation normal, evaluation too low;

[0013] The second user feedback acquisition subunit is used to acquire the user's evaluation content for each of the evaluation results, the evaluation content including the score and evaluation information for each evaluation result.

[0014] Optionally, the system further includes:

[0015] The model loading unit is used to load the corresponding first model based on the model information input by the user.

[0016] A status setting unit is used to set the usage status of the first model, the usage status including an enabled status and a deprecated status;

[0017] A model information display unit is used to display the model information of the first model;

[0018] The parameter configuration unit is used to configure the parameters of the first model according to the model configuration parameters.

[0019] Optionally, the system further includes:

[0020] The file receiving unit is used to receive a task file through a first interface, wherein the task file includes task data corresponding to the data processing task.

[0021] The result file acquisition unit is used to output the result file corresponding to the data processing task through the second interface. The result file includes the data processing result of the first model and the evaluation result of the second model.

[0022] Optionally, the system further includes:

[0023] The evaluation result display unit is used to display the evaluation score corresponding to each of the first models, and the evaluation reason corresponding to the evaluation score;

[0024] An image generation unit is configured to calculate the average score of the evaluation scores corresponding to all the first models after all the first models have completed the data processing task, and generate a corresponding score distribution map, and mark the average score in the score distribution map.

[0025] The evaluation history display unit is used to display the data processing results corresponding to each data processing task executed by the first model, and the evaluation record corresponding to the data processing results. The evaluation record includes the task content of the data processing task, the evaluation result corresponding to the data processing result, and the evaluation reason.

[0026] A second aspect of this application provides a data processing method, which is applied to a model evaluation system, comprising:

[0027] Receive task data corresponding to the data processing task;

[0028] The data processing task is executed by at least one first model to obtain at least one data processing result;

[0029] Input the at least one data processing result into the second model;

[0030] The data processing results are evaluated from at least one preset dimension using the second model to obtain at least one evaluation result;

[0031] Based on the evaluation results, a target data processing result is determined from the at least one data processing result.

[0032] Optionally, before performing the data processing task through at least one first model to obtain at least one data processing result, the method further includes:

[0033] Receive the model information of the first model to be loaded;

[0034] If the model address included in the model information is a local path, the first model is obtained from the storage location corresponding to the local path;

[0035] If the model address included in the model information is a program interface address, the first model is downloaded through the program interface address;

[0036] According to the model parameter types included in the model information, the corresponding model parameter values ​​are input into the first model;

[0037] Set the first model to enabled.

[0038] Optionally, the method further includes:

[0039] The data processing results corresponding to the first model, the data processing results, and the corresponding evaluation reasons are stored in the evaluation record corresponding to the first model.

[0040] Upon receiving a request to view the evaluation record corresponding to the first model, the evaluation record is displayed.

[0041] Optionally, the method further includes:

[0042] Upon receiving a re-evaluation instruction, the at least one data processing result is re-evaluated using the second model.

[0043] Optionally, the step of evaluating the data processing results from at least one preset dimension using the second model to obtain at least one evaluation result includes:

[0044] Using preset format matching rules, the data processing results are scored from the format dimension to obtain a format score result;

[0045] By calculating the consistency score between the content and the topic of the data processing results, the data processing results are scored from the perspective of topic expression, and a topic expression score result is obtained.

[0046] The data processing results are sorted by confusion level, and the data processing results are scored from the fluency dimension to obtain a fluency score result;

[0047] Using preset repetition evaluation rules, the data processing results are scored from the repetition dimension to obtain repetition score results;

[0048] Using pre-set rhyme judgment rules, the data processing results are scored from the prosody dimension to obtain a prosody score result;

[0049] Using a pre-set rhetoric judgment method, the data processing results are scored from the perspective of rhetoric skills to obtain rhetoric score results;

[0050] The format score, the theme expression score, the fluency score, the repetition score, the rhythm score, and the rhetoric score are weighted and calculated according to a preset weight ratio to obtain the data processing result.

[0051] A third aspect of this application provides a data processing apparatus, which is applied to a model evaluation system and includes:

[0052] The task data receiving module is used to receive task data corresponding to data processing tasks.

[0053] A data processing result acquisition module is used to perform a data processing task through the at least one first model to obtain at least one data processing result;

[0054] A data processing result input module is used to input the at least one data processing result into the second model;

[0055] The data processing result evaluation module is used to evaluate the data processing result from at least one preset dimension using the second model, and obtain at least one evaluation result;

[0056] A data processing result determination module is used to determine a target data processing result from the at least one data processing result based on the evaluation result.

[0057] Optionally, the device further includes:

[0058] The model information receiving module is used to receive model information of at least one first model to be loaded;

[0059] The first model acquisition submodule is used to acquire the first model from the storage location corresponding to the local path when the model address is a local path;

[0060] The second model acquisition submodule is used to download the first model through the program interface address when the model address is the program interface address.

[0061] The parameter input submodule is used to input the corresponding model parameter value into the first model according to the model parameter type;

[0062] The model activation module is used to set the first model to an enabled state after the first model has been loaded.

[0063] Optionally, the device further includes:

[0064] The evaluation record module is used to store the data processing result corresponding to the first model, the data processing result, and the corresponding evaluation reason in the evaluation record corresponding to the first model;

[0065] The evaluation record display module is used to display the evaluation record when a request is received to view the evaluation record corresponding to the first model.

[0066] Optionally, the device further includes:

[0067] A re-evaluation module is used to re-evaluate the at least one data processing result using the second model upon receiving a re-evaluation instruction.

[0068] Optionally, the data processing result evaluation module includes:

[0069] The first evaluation submodule is used to score the data processing results from the format dimension using preset format matching rules, and obtain a format score result;

[0070] The second evaluation submodule is used to score the data processing results from the dimension of topic expression by calculating the consistency score between the content and topic of the data processing results, and to obtain the topic expression score result.

[0071] The third evaluation submodule is used to sort the data processing results by confusion level and score the data processing results from the fluency dimension to obtain a fluency score result.

[0072] The fourth evaluation submodule is used to score the data processing results from the repeatability dimension using preset repeatability evaluation rules, and obtain the repeatability score result.

[0073] The fifth evaluation submodule is used to score the data processing results from the prosody dimension using pre-set rhyme judgment rules, and obtain the prosody score result;

[0074] The sixth evaluation submodule is used to score the data processing results from the perspective of rhetorical skills using a pre-set rhetorical judgment method, and obtain the rhetorical score result.

[0075] The evaluation result acquisition submodule is used to perform weighted calculations on the format score result, the topic expression score result, the fluency score result, the repetition score result, the rhythm score result, and the rhetoric score result according to a preset weight ratio to obtain the data processing result.

[0076] Optionally, the device further includes:

[0077] The first weight ratio adjustment module is used to adjust the weight ratio according to the data processing result;

[0078] The second weight ratio adjustment module is used to adjust the weight ratio based on user feedback information;

[0079] The third weight ratio adjustment module is used to adjust the weight ratio by iteratively optimizing different combinations of the weight ratio through a preset optimization algorithm.

[0080] A fourth aspect of this application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps described in the first aspect of this application.

[0081] A fifth aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in the first aspect of this application.

[0082] Using the data processing method provided in this application, task data corresponding to a data processing task is received; the data processing task is executed through at least one first model to obtain at least one data processing result; the at least one data processing result is input into a second model; the data processing result is evaluated from at least one preset dimension through the second model to obtain at least one evaluation result; and a target data processing result is determined from the at least one data processing result based on the evaluation result.

[0083] In this application, at least one first model is preloaded in the model evaluation system. Task data is input into the first model to obtain the corresponding data processing results. Then, the data processing results are evaluated in multiple dimensions by a second model to obtain the final data processing result. This method can freely load the first model that needs to be evaluated, evaluate multiple first models at the same time, and evaluate the data processing results from multiple dimensions. In this way, the performance differences of different models can be accurately grasped through the evaluation results, and the optimal data processing result can be selected. Attached Figure Description

[0084] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0085] Figure 1 This is a schematic diagram of the structure of a model evaluation system proposed in an embodiment of this application;

[0086] Figure 2 This is an architecture diagram of a model evaluation system proposed in an embodiment of this application;

[0087] Figure 3 This is a flowchart of the model evaluation system proposed in one embodiment of this application;

[0088] Figure 4 This is a system evaluation flowchart proposed in one embodiment of this application;

[0089] Figure 5 This is a flowchart of a data processing method proposed in an embodiment of this application;

[0090] Figure 6 This is a flowchart of the model parameter setting method proposed in one embodiment of this application;

[0091] Figure 7 This is a schematic diagram of a data processing apparatus according to an embodiment of this application;

[0092] Figure 8This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0093] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0094] refer to Figure 1 , Figure 1 This is a schematic diagram of the model evaluation system structure proposed in the embodiments of this application, as shown below. Figure 1 As shown, the system includes:

[0095] A data processing unit is used to call at least one first model to perform a data processing task and obtain at least one data processing result, wherein the at least one first model is a different model; a result evaluation unit is used to evaluate the at least one data processing result respectively through a second model to obtain at least one evaluation result and output the evaluation result to the user page, wherein the evaluation result is used to compare the differences between the data processing results of different first models; a user feedback acquisition unit is used to acquire user feedback information on the evaluation result so that the second model can optimize the evaluation logic based on the feedback information.

[0096] In this embodiment, the first model is an artificial intelligence model that performs data processing tasks. It can execute corresponding data processing tasks based on the user's task requirements. For example, if a user needs to generate a poem using the model, the data processing task to be performed is the poem generation task. The user inputs the corresponding task data, such as the main body and format of the poem, and the first model can then generate the corresponding poem. This system can simultaneously use multiple different first models to perform data processing tasks, thereby obtaining multiple data processing results.

[0097] In this embodiment, the second model is a text content evaluation model trained using Transformer (a deep learning framework) based on Natural Language Processing (NLP). Taking poetry recognition as an example, when training the second model, processed poems conforming to standard formats are used as training material and input into the model to be trained. The parameters of the second model are adjusted to obtain a trained second model, which has the performance to evaluate and score elements such as the format, theme, and rhyme of poems, thus making a comprehensive evaluation of a poem. This evaluation score reflects the quality of the data processing results and thus reflects the performance of the first model. By comparing multiple evaluation scores, the differences between the data processing results of the first model can be determined.

[0098] The result evaluation unit includes:

[0099] The dimension evaluation subunit is used to evaluate the data processing results from at least one preset dimension to obtain the evaluation results.

[0100] In this embodiment, the second model evaluates the data processing results through at least one preset dimension. Each dimension is evaluated using a corresponding algorithm. For example, when evaluating poetry, the content of the poetry is evaluated from multiple dimensions, including format, theme expression, fluency, repetition, rhythm, and rhetorical techniques.

[0101] The user feedback acquisition unit includes at least one of the following: a first user feedback acquisition subunit, used to acquire the user's evaluation selection for each of the evaluation results, the evaluation selection including any of the following: evaluation too high, evaluation normal, evaluation too low; and a second user feedback acquisition subunit, used to acquire the user's evaluation content for each of the evaluation results, the evaluation content including the score for each of the evaluation results and evaluation information.

[0102] In this embodiment, the user can evaluate each evaluation result on the system's results display page. Options include evaluation too high, evaluation normal, evaluation too low, or the user can directly score the evaluation result and specify the corresponding evaluation information, including the evaluation reason, etc. The user feedback acquisition unit obtains the user's evaluation selection or evaluation information to optimize the parameters of the second model.

[0103] The system also includes:

[0104] The model loading unit is used to load the corresponding first model based on the model information input by the user.

[0105] The status setting unit is used to set the usage status of the first model, which includes an enabled status and a deactivated status; the model information display unit is used to display the model information of the first model; and the parameter configuration unit is used to configure the parameters of the first model according to the model configuration parameters.

[0106] In this embodiment, the user can load the first model into the system, specifying the model name, loading path, and other information. The system can then load the corresponding first model according to the loading path. The user can also set the usage status of each model, and the model information of each model will be displayed in the model list. Furthermore, the user can configure the parameters of the model independently by inputting the corresponding parameters.

[0107] The system further includes: a file receiving unit, configured to receive a task file through a first interface, the task file including task data corresponding to the data processing task; and a result file obtaining unit, configured to output a result file corresponding to the data processing task through a second interface, the result file including the data processing result of the first model and the evaluation result of the second model.

[0108] In this embodiment, the system can receive a task file containing task data through a preset first interface. After executing the data processing task, it can output a result file containing the corresponding result data through a second interface. The user can provide the task file and receive the corresponding result file.

[0109] The system also includes:

[0110] The evaluation result display unit is used to display the evaluation score corresponding to each of the first models and the evaluation reason corresponding to the evaluation score; the image generation unit is used to calculate the average score of the evaluation scores corresponding to all the first models after all the first models have completed the data processing task, and generate a corresponding score distribution map, and mark the average score in the score distribution map; the evaluation history display unit is used to display the data processing result corresponding to each data processing task executed by the first model, and the evaluation record corresponding to the data processing result, wherein the evaluation record includes the task content of the data processing task, the evaluation result corresponding to the data processing result, and the evaluation reason.

[0111] In this embodiment, the evaluation results can be recorded and statistically analyzed. Users can view the output results of each first model and second model recorded after each data processing task is executed on the corresponding page, as well as the detailed information of the evaluation results corresponding to each data processing task, including evaluation content, evaluation reasons, evaluation scores, etc. It can also calculate the average score of multiple evaluation results, generate a score distribution chart, and mark the average score on the score distribution chart, so that users can intuitively see the score of the data processing results output by each model.

[0112] refer to Figure 2 , Figure 2 This is an architecture diagram of a model evaluation system proposed in an embodiment of this application, as shown below. Figure 2 As shown, the front end of the model evaluation system includes a homepage, a model management page, and an evaluation history page. The homepage displays model functionality, model inference functionality, model evaluation functionality, and provides a visual representation of model evaluation scores. The model management page includes functions for model download, model configuration, model parameter settings, and model on / off. The evaluation history page displays the number of evaluations for each primary model and related data. The details page displays detailed data for each inference iteration of the primary model. The inference evaluation interface in the program interface can import user-defined content to be evaluated, and the export interface outputs the corresponding evaluation scores. This system relies on multiple primary models for inference, on secondary models for scoring, and on database-provided functions.

[0113] refer to Figure 3 , Figure 3 This is a flowchart of the model evaluation system proposed in one embodiment of this application, as follows: Figure 3 As shown, when users access the system, the homepage will open by default. The homepage primarily features functions such as setting and displaying inference content and second-model evaluation content. Users can switch to the model settings page to perform model-related operations, including downloading models, setting parameters, and enabling / disabling them. Users can view the evaluated content and its details on the historical evaluation page. When a user enters the homepage, the system displays the models currently set and activated by the user. The system will then use these activated models for inference. Users can choose between automatic evaluation, manual evaluation, or self-evaluation for data processing. In automatic evaluation, the system will automatically evaluate the data processing results using the second model after all first-model inferences have been completed. In manual evaluation, users click the evaluation button on the homepage. For self-evaluation, after all models have completed inference, users fill in the evaluation information (score) below each model and submit it to the second model for the larger model's reference.

[0114] After entering the management page, a list of models is displayed, including downloaded models, local models, and online models. Online models will display status information such as pending download, downloading, and download complete, as well as the download path and parameter configuration information. The local model list displays relevant information such as the local path, online path, and parameter configuration. Online models will display the online address and parameter configuration information. Both downloaded and local models require model parameter configuration.

[0115] After entering the historical evaluation page, the system will display the evaluation records of each first model in the evaluation list according to the evaluation time order. This includes model information of the models that have been evaluated. You can click on the details to enter the evaluation details page, which includes inference content, data processing results, user self-evaluation information, etc.

[0116] The system also supports user feedback, allowing users to rate the evaluation score of the second model. Based on user feedback, the second model optimizes its evaluation logic to better meet user requirements. The second model determines whether its performance evaluation score is high, low, or normal based on user ratings, and it can further optimize its parameters according to user feedback scores, ensuring that the second model's evaluation score better meets user needs.

[0117] refer to Figure 3 , Figure 3 This is a system evaluation flowchart proposed in one embodiment of this application, such as... Figure 3 As shown, when a user operates the system, they first add the model to be used on the model configuration page, enter the model name, local path (required for local model calls), API (Application Programming Interface) address (required for online model calls), model parameters and corresponding values, and then save the configuration. At this time, the system will enable the model and make API calls using the configured address and parameters.

[0118] Next, select the model you want to use, enable or disable the model on the model configuration page, and after configuration, click the Start Inference button to perform inference and display the data processing results on the homepage.

[0119] The homepage is the main page for user-input inference evaluation. It includes controls for user input of inference content, multi-model data processing result display, large-model evaluation methods, large-model evaluation scoring results, large-model evaluation reasons, and user self-evaluation. Users input the content to be inferred in the inference content control and click the "Start Inference" control, at which point the system begins inference. Multiple models infer simultaneously, and the results are displayed in real-time in the multi-model data processing result display until all models have completed inference.

[0120] After all the first models have completed inference, the system automatically evaluates the data processing results based on the instructions received from the large model evaluation method control, or waits for the user to click the manual evaluation button. After the evaluation is completed, the user can click the re-evaluation button to re-evaluate the data processing results. The system's self-evaluation control can receive custom content sent by the user, evaluate the custom content, and obtain the corresponding evaluation score.

[0121] After evaluation, the visualization interface will display the evaluation score and corresponding evaluation reasons below the data processing results of each model. Finally, an average evaluation score and a score distribution chart are generated to intuitively display the score rankings. Users can go to the evaluation history page to view all evaluation records, which will display information such as inference content, average inference score, and score for each model. Users can also use a local API to automatically perform inference and evaluation on local files. The local file content contains all the content the user wants to infer. Then, through the API, the final multi-model data processing results and evaluation scores are obtained.

[0122] refer to Figure 4 , Figure 4 This is a flowchart of a data processing method proposed in an embodiment of this application, which is applied to a model evaluation system. Figure 4 As shown, the method includes the following steps:

[0123] S11: Receive task data corresponding to the data processing task.

[0124] In this embodiment, the model evaluation system pre-loads at least one first model. These first models can be obtained locally or downloaded from a model library on the network via a program interface. Task data corresponding to the data processing task is received through the pre-loaded at least one first model.

[0125] For example, the first model is a language model that can generate verses, and the task data is "generate a five-character quatrain about the moon".

[0126] S12: Perform the data processing task through at least one first model to obtain at least one data processing result.

[0127] In this embodiment, the data processing result is the output obtained by the first model based on the input task data.

[0128] In this embodiment, multiple first models simultaneously receive task data and generate corresponding data processing results based on the task data. These first models are all pre-trained models that process the input task data using parameters learned during training to obtain the corresponding data processing results.

[0129] For example, the first model will generate a five-character quatrain containing the moon or the element of the moon based on the input task data "generate a five-character quatrain about the moon".

[0130] S13: Input the at least one data processing result into the second model.

[0131] In this embodiment, when multiple data processing results are obtained from multiple first models, the multiple data processing results are respectively input into the second model.

[0132] S14: The data processing results are evaluated from at least one preset dimension using the second model to obtain at least one evaluation result.

[0133] In this embodiment, the evaluation result is the evaluation score obtained after evaluating the data processing results of the first model.

[0134] In this embodiment, the input data processing results are evaluated from multiple dimensions using a second model to obtain evaluation results. Multiple aspects need to be considered during the evaluation, and thus multiple dimensions are set to evaluate the data processing results.

[0135] In this embodiment, taking poetry as an example, the second model is mainly based on Natural Language Processing (NLP) technology. It uses Transformer (a deep learning framework) technology to train the model, making it capable of understanding the language structure and style of poetry.

[0136] In this embodiment, taking poetry evaluation as an example, most current poetry evaluation methods use BLEU (bilingual evaluation understudy) with certain rules. However, when using rule-based evaluation alone, only the format can be evaluated, and other evaluation effects cannot be achieved. Although the BLEU method is fast, does not distinguish between languages, and has better applicability, it only considers the similarity of N-grams and does not include semantic-level evaluation.

[0137] In this embodiment, the content of the poem is evaluated from multiple dimensions, including format, thematic expression, fluency, repetition, rhythm, and rhetorical techniques.

[0138] In this embodiment, the specific steps for scoring the data processing result using at least one preset evaluation dimension to obtain the data processing result include:

[0139] S14-1: Using preset format matching rules, score the data processing results from the format dimension to obtain a format score result.

[0140] In this embodiment, different inference content has different format requirements. Matching rules can be set according to the actual situation. Using the preset format matching rules, the data processing results are scored from the format dimension to obtain the format score result. When building the second model, the set format matching rules are added to the model. During the training process, the second model will adjust the corresponding parameters of the format matching rules, thereby continuously improving the format matching rules.

[0141] In this implementation, taking poetry as an example, the second model supports the recognition of formats including five-character quatrains, seven-character quatrains, five-character regulated poems, and seven-character regulated poems. It also supports various ci (lyric poetry) titles such as *Zhegutian*, *Huanxisha*, *Manjianghong*, *Pusa Man*, *Dielianhua*, *Linjiangxian*, *Shuidiaogetou*, *Qingpingyue*, and *Niannujiao*. Furthermore, it supports different character counts for the same ci title, enabling better format verification and evaluation of different types of poetry. The model uses rule matching, where the content is segmented based on punctuation marks. It then determines whether each segment conforms to the corresponding format and character count. If it does, the data processing result is 100 points; otherwise, it is 0 points. For example, if each line of a five-character quatrain has five characters, and each line in the data processing result has five characters, it is considered a perfect score.

[0142] The corresponding evaluation formula is:

[0143] (1)

[0144] Where format_score represents the format score, f(x) represents the final score, and weight represents the weight. As can be seen from the formula, the score result is divided into two parts. If format_score is 0, the overall score is 0. If format_score is not 0, the final score is the sum of the scores of each dimension multiplied by the scores of each dimension.

[0145] S14-2: By calculating the consistency score between the content and the topic of the data processing result, the data processing result is scored from the topic expression dimension to obtain the topic expression score result.

[0146] In this embodiment, the consistency score between the content and the topic of the data processing result is also called the answer fidelity score. It is an evaluation index obtained using LLM (Large Language Model) to evaluate whether the topic of a piece of content meets the proposed requirements. The second model integrates the language recognition function of the large language model and continuously adjusts the parameters for recognizing the content during the training process, thereby accurately obtaining the answer fidelity score for each inference content.

[0147] In this embodiment, the data processing results are scored from the perspective of topic expression by calculating the answer fidelity score of the data processing results, thereby obtaining the corresponding topic expression score result.

[0148] In this embodiment, the second model is used to decompose the poetry composition result into multiple statements, and the consistency of each statement with the theme is examined. A "fidelity score" is calculated based on the ratio of the number of supported statements to the total number of statements. The score is awarded in an additive format, with 5 points added for each supported statement, up to a maximum of 100 points.

[0149] (2)

[0150] Here, `score` represents the final score, `w` represents the weight, `k` represents the number of sums, and `y` represents the score added for a match, where `y` is always equal to 5. The weighted sum is compared to 100, and the minimum is taken. Finally, the score is obtained by multiplying by the weight `w`.

[0151] In this embodiment, the second model fine-tunes the model to achieve a deeper fit with the theme. By using model fine-tuning, the content of the poems is compared with the user's questions and answers to evaluate the degree of relevance between the content and the theme. This includes the theme words and words related to the theme words. For example, if the theme is "Write a five-character quatrain with the Mid-Autumn Festival as the theme", then the theme word is "Mid-Autumn Festival". However, if the title contains "moon" or "Chang'e", it will also be evaluated as theme relevance.

[0152] S14-3: Sort the data processing results by confusion level, score the data processing results from the fluency dimension, and obtain the fluency score result.

[0153] In this embodiment, the basic idea of ​​perplexity ranking (PPL) is to calculate the probability of a sentence appearing. The higher the score, the lower the probability of its appearance. That is, the higher the score, the lower the probability of the sequence appearing, and the more difficult it is to produce fluent text. The lower the score, the easier it is to produce fluent text. Based on this idea, the fluency of the poem is calculated.

[0154] In this embodiment, the probability of a sentence appearing is calculated and raised to the power of 1 / N, as shown in the following formula:

[0155] (3)

[0156] Where w1 w2…wN represent N tokens in a sentence (tokens correspond to the dictionary in the language model; for Chinese, they may be a character or a word), and P is the probability of generating each word.

[0157] As can be seen from the above formula, when using PPL sorting analysis to analyze a poem, the lower the score, the more likely the poem is to be a fluent text and the higher its fluency.

[0158] The scoring formula is:

[0159] (4)

[0160] Where score is the final score, w represents the weight, and PPL is the result calculated as described above. The final score is obtained by multiplying by the weight w.

[0161] In this embodiment, fluency is mainly expressed using a formula. The final results are sorted by perplexity, and the fluency of the data processing results is judged by the standard of formula evaluation, which greatly improves the efficiency of evaluation.

[0162] S14-4: Using preset repeatability evaluation rules, score the data processing results from the repeatability dimension to obtain repeatability score results.

[0163] In this embodiment, a preset repetition evaluation rule is used to score the data processing results from the repetition dimension to obtain a repetition score result. The repetition evaluation rule is added when the second model is built. When training the model, the relevant parameters of the repetition evaluation rule can be iteratively adjusted, so as to better process and analyze the poems through the rule.

[0164] In this embodiment, the repetition rate is evaluated using a rule-based approach. The poem content is broken down character by character, and the degree of repetition is then determined. The higher the repetition rate, the lower the score. This evaluation dimension uses a deduction system, with a starting score of 100 points. Each instance of repetition deducts 5 points, with a maximum score of 100 points and a minimum score of 0 points. The formula is:

[0165] (5)

[0166] Where score is the final score, w represents the weight, k represents the number of words to be summed, and y represents the score for repeated words, with y=5 being a constant value. The score is calculated by subtracting the score for matching words from 100 points and comparing it with 0; the larger value is taken as the final result, which is then multiplied by the weight w to obtain the final score.

[0167] In this embodiment, the repetition degree is the measurement of the repetition rate of the poem content in the final result. For a poem, if there is too much repeated content, it may lead to a rather rigid result. Of course, it is not excluded that some good poems use repeated content, such as "Jiji fu jiji". Therefore, the method of subtracting the repeated number from the total score is adopted to subtract the repeated content and obtain the final result. This not only judges the repeated number but also optimizes the result in a form from high to low, avoiding the situation of overly weakening repeated words.

[0168] S14-5: Use the pre-set rhyme judgment rule to score the data processing result from the prosody dimension and obtain the prosody scoring result.

[0169] In this embodiment, the pre-set rhyme judgment rule is used to score the data processing result from the prosody dimension and obtain the prosody scoring result. When constructing the second model, the rhyme judgment rule is added, and relevant parameters can be continuously optimized as the model is trained, so as to more accurately perceive and evaluate the prosody and rhythm of a text content.

[0170] In this embodiment, the method of converting Chinese characters to pinyin is adopted. The last character of each sentence is converted to pinyin, and it is judged whether it meets the rhyme. According to the degree of satisfaction, a score is given. The higher the proportion of rhyming sentences that are satisfied, the higher the score. The formula is as follows:

[0171] (6)

[0172] Where score is the final score, w represents the weight, n represents the total number of rhymes that meet the requirements, t is the maximum number of rhymes for the current type of poem, and multiplying by 100 is to obtain a value within 100 after calculating the proportion, so as to keep the same unit as other dimensions for statistics. Multiply by the weight w to get the final score.

[0173] In this embodiment, the second model adopts the method of pinyin recognition, compares the pinyin of the last character, and judges whether it meets the rhyme, which can better evaluate the prosody.

[0174] S14-6: Use the pre-set rhetorical judgment method to score the data processing result from the rhetorical skill dimension and obtain the rhetorical scoring result.

[0175] In this embodiment, when generating an article or a poem, it is necessary to judge its rhetorical devices. Use the pre-set rhetorical judgment method to score the data processing result from the rhetorical skill dimension and obtain the rhetorical scoring result. During the training process of the second model, the parameters corresponding to the rhetorical judgment method will be continuously adjusted, so that the model can more accurately judge the rhetorical methods used in the text content.

[0176] In this embodiment, the second model translates the input poem into modern Chinese. Then, it determines the poem's rhetorical skill by judging whether the modern Chinese text contains metaphors, personification, or other rhetorical devices, and assigns a score accordingly. The more rhetorical devices used, the higher the score. Each rhetorical device adds 5 points, up to a maximum of 100 points. The formula is as follows:

[0177] (7)

[0178] Where score is the final score, w represents the weight, k represents the number of sums, and y represents the score added for rhetorical devices, where y=5. The result is multiplied by the weight w to obtain the final score.

[0179] In this embodiment, the second model uses the method of translating ancient poems into modern text to determine the rhetorical devices. This method simplifies the judgment logic and increases the accuracy of the judgment. The large model is trained to judge different rhetorical devices through fine-tuning. It can judge various devices such as metaphor, parallelism, and personification. By judging and evaluating different devices, the dimensions of poetry evaluation are increased.

[0180] S14-7: The format score, the theme expression score, the fluency score, the repetition score, the rhythm score, and the rhetoric score are weighted and calculated according to a preset weight ratio to obtain the data processing result.

[0181] In this embodiment, the format score, topic expression score, fluency score, repetition score, rhythm score, and rhetoric score are weighted according to preset weight ratios to obtain the data processing result. The data processing result is a comprehensive score.

[0182] In this embodiment, the second model supports specifying the weight of each evaluation dimension as a parameter. For example, if the theme expression is the most important, it can be set to the highest weight, with 100 points being the highest score and 0 points being the lowest. Currently, the sum of all parameter weights is 100. The system will first select the dimension with the highest weight, and then select the dimension with the next highest weight. If the sum of the preceding parameter settings reaches 100, the subsequent parameters will be automatically set to 0. If there are dimensions with the same weight, they will be selected according to the system's preset order. For example, for fluency and repetition, the system sets fluency to have a higher priority than repetition. If these two weights are set the same, the system will prioritize fluency for weight summation. If there are remaining weights, then repetition will be allocated.

[0183] In this embodiment, the initial weighting ratio is based on expert experience and opinions. Initially, the expert opinions are: topic expression: 40%, fluency: 25%, repetition: 15%, rhythm: 10%, and rhetorical skills: 10%.

[0184] In this embodiment, the method further includes:

[0185] S14-8: Adjust the weight ratio based on the data processing results.

[0186] In this embodiment, a feedback loop is established. After obtaining the data processing results, the model adjusts the proportion of the scoring criteria by automatically feeding back the evaluation score. The feedback loop can be implemented by methods such as reinforcement learning or genetic algorithms. Through these algorithms, the model can gradually optimize its internal parameters based on the evaluation score.

[0187] S14-9: Adjust the weight ratio based on user feedback.

[0188] In this embodiment, user feedback information refers to the user's autonomous rating of the data processing results of each first model, and also includes the user's rating of the second model.

[0189] In this embodiment, user feedback information is introduced to further adjust the proportion of the scoring criteria. Users can score the data processing results or provide other forms of feedback, and the user feedback information guides the parameter optimization of the second model.

[0190] S14-10: By using a preset optimization algorithm, iteratively optimize different combinations of the weight ratios and adjust the weight ratios.

[0191] In this embodiment, hyperparameter optimization techniques can be used to automatically adjust the weight ratios to maximize the overall score. For example, a Bayesian model can be used for optimization, and a corresponding acquisition function can be designed to iteratively optimize the model. Different ratio combinations can be tried automatically within a given search space to find the optimal solution.

[0192] In this embodiment, the initial weighting of each dimension is based on the expert opinion rating ratio. Then, hyperparameter optimization techniques (such as Bayesian optimization) are used to find an optimized set of weighting ratios. These optimized weighting ratios are then input into the second model, and further fine-tuned through a feedback loop (such as reinforcement learning). The feedback loop can automatically adjust the ratios based on the performance of the poems generated by the model in real or simulated environments to achieve better results.

[0193] In addition to the automatic feedback loop, user feedback is introduced to guide the optimization process. Users can rate the poems generated by the model or provide other forms of feedback. This user feedback serves as additional input data to adjust the optimization objectives or weights in the feedback loop. For example, if most users feel that the rhetoric of a poem is not good enough, the system can automatically increase the weight of the rhetoric score and adjust the weights of other parts accordingly.

[0194] For example, after continuous optimization, the resulting weight ratios are: topic expression: 38.3%, fluency: 30.5%, repetition: 10.6%, rhythm: 11.2%, and rhetorical skills: 9.4%.

[0195] S15: Based on the evaluation results, determine the target data processing result from the at least one data processing result.

[0196] In this embodiment, after obtaining the evaluation result corresponding to each data processing result, the evaluation result with the highest score is determined from multiple data processing results and used as the target data processing result.

[0197] For example, when the data processing task is to generate a poem, the poem with the highest score is selected as the poem to be used, and the first model that generates the poem is determined as the optimal model.

[0198] In this embodiment, a comprehensive optimization framework is constructed that integrates various optimization strategies such as hyperparameter optimization, feedback loops, and user feedback. Different optimization stages and priorities are set within this framework. By continuously optimizing the weights of each dimension of the evaluation text in the second model through multiple strategies, the second model can make more accurate evaluations of the data processing results, effectively improving the efficiency and accuracy of the second model's evaluation.

[0199] In another embodiment of this application, before receiving task data via at least one pre-loaded first model, the method further includes:

[0200] S21: Receive the model information of the first model to be loaded.

[0201] In this embodiment, the model information includes the model name, local path, program interface address, model parameters, and model parameter values.

[0202] In this embodiment, the system first receives model information of at least one first model to be loaded.

[0203] In this embodiment, the model information includes the model name, model address, model parameter type, and model parameter value.

[0204] S22: If the model address included in the model information is a local path, obtain the first model from the storage location corresponding to the local path.

[0205] In this embodiment, if the model address is a local path, it means that the model is stored locally, so the first model can be obtained from the storage location corresponding to the local path.

[0206] S23: If the model address included in the model information is a program interface address, download the first model through the program interface address.

[0207] In this embodiment, if the model address is the program interface address, the first model needs to be downloaded from the program interface address.

[0208] S24: Based on the model parameter types included in the model information, input the corresponding model parameter values ​​into the first model.

[0209] In this embodiment, the model parameter type specifies the specific type of the model parameters, and the model parameter value is the specific numerical value of the parameter.

[0210] In this embodiment, the model information includes the type of model parameters and the corresponding model parameter values. Based on the type of model parameters, the corresponding model parameter values ​​are loaded into the first model, and the first model is loaded into the model evaluation system.

[0211] S25: Set the first model to the enabled state.

[0212] In this embodiment, once the first model has been loaded, the first model is set to the startup state.

[0213] In this embodiment, it is also necessary to set the model parameters, referring to... Figure 5 , Figure 5 This is a flowchart of the model parameter setting process proposed in an embodiment of this application, as follows: Figure 5 As shown, users select the model whose parameters need to be set on the page, fill in the request header information, the specific request parameters, and the URL (network address) request parameters, and save the information. The system validates the parameters, and if the parameters are valid, stores them in the database. When using the model, the stored parameters are retrieved and loaded into the model. The following parameter data tables were designed for the model-related parameters. The fields of the main parameters are shown in Tables 1 and 2. Table 1 is the model parameter type table, and Table 2 is the model parameter value table.

[0214]

[0215] Table 1

[0216]

[0217] Table 2

[0218] In this embodiment, users can load the corresponding model into the system according to their needs and set the corresponding parameters accordingly, which can achieve flexible expansion of the first model and meet the user's needs.

[0219] In another embodiment of this application, after obtaining the performance evaluation score, the average score of all the data processing results is calculated.

[0220] For example, if there are three data processing results, 95, 96, and 97, then the average score is 96.

[0221] In this embodiment, after calculating the average score, a score distribution chart of these data processing results is generated. The average score of these performance evaluation scores is marked on the score distribution chart, which can intuitively show whether the performance of the first model is higher than the average.

[0222] In this embodiment, by generating the average score and the score distribution chart, the score of each model can be seen intuitively, which helps users to grasp the specific performance of each first model.

[0223] In another embodiment of this application, the method further includes:

[0224] S31: Store the data processing result corresponding to the first model, the data processing result, and the corresponding evaluation reason in the evaluation record corresponding to the first model.

[0225] In this embodiment, for each first model, the data processing results of the model, the data processing results, and the corresponding evaluation reasons are stored in the evaluation record corresponding to the first model, so that users can view them at any time.

[0226] S32: Upon receiving a request to view the evaluation record corresponding to the first model, the evaluation record is displayed.

[0227] In this embodiment, when the system receives a request to view the evaluation record of the first model, the evaluation record is displayed.

[0228] In this embodiment, users can view the evaluation records at any time to determine the overall performance of each first model and gain a comprehensive understanding of the model's performance.

[0229] In another embodiment of this application, the method further includes:

[0230] S41: Upon receiving a re-evaluation instruction, the at least one data processing result is re-evaluated using the second model.

[0231] In this embodiment, when the system receives a re-evaluation instruction, the model evaluation system re-evaluates the data processing results using a second model. Because parameters can be adjusted through feedback loops during the evaluation process, the scores may differ each time, making multiple evaluations more meaningful.

[0232] In this embodiment, the user can issue a re-evaluation command at any time to re-evaluate the evaluation score. Based on the results of multiple evaluations, the user can grasp the overall performance of each first model.

[0233] In another embodiment of this application, the system statistically analyzes the evaluation scores of the second models. If the score of a first model is lower than a preset score threshold in multiple evaluations, the first model is marked as a model not recommended for use, to remind the user that the first model has poor performance and cannot provide good service to the user. Alternatively, the system can sort the scores of each first model, and if a first model has a low ranking in multiple rankings, it is marked as a model not recommended. In this way, the system can help users filter out high-performance first models and eliminate low-performance first models, allowing users to quickly find high-performance models and improve the user experience.

[0234] Based on the same inventive concept, one embodiment of this application provides a data processing apparatus applied to a model evaluation system. (Reference) Figure 7 , Figure 7 This is a schematic diagram of a data processing apparatus 700 according to an embodiment of this application. Figure 7 As shown, the device includes:

[0235] The task data receiving module 701 is used to receive task data corresponding to the data processing task.

[0236] The data processing result acquisition module 702 is used to perform a data processing task through the at least one first model to obtain at least one data processing result;

[0237] The data processing result input module 703 is used to input the at least one data processing result into the second model;

[0238] The data processing result evaluation module 704 is used to evaluate the data processing result from at least one preset dimension using the second model, and obtain at least one evaluation result;

[0239] The data processing result determination module 705 is used to determine a target data processing result from the at least one data processing result based on the evaluation result.

[0240] Optionally, the device further includes:

[0241] The model information receiving module is used to receive model information of at least one first model to be loaded;

[0242] The first model acquisition submodule is used to acquire the first model from the storage location corresponding to the local path when the model address is a local path;

[0243] The second model acquisition submodule is used to download the first model through the program interface address when the model address is the program interface address.

[0244] The parameter input submodule is used to input the corresponding model parameter value into the first model according to the model parameter type;

[0245] The model activation module is used to set the first model to an enabled state after the first model has been loaded.

[0246] Optionally, the device further includes:

[0247] The evaluation record module is used to store the data processing result corresponding to the first model, the data processing result, and the corresponding evaluation reason in the evaluation record corresponding to the first model;

[0248] The evaluation record display module is used to display the evaluation record when a request is received to view the evaluation record corresponding to the first model.

[0249] Optionally, the device further includes:

[0250] A re-evaluation module is used to re-evaluate the at least one data processing result using the second model upon receiving a re-evaluation instruction.

[0251] Optionally, the data processing result evaluation module includes:

[0252] The first evaluation submodule is used to score the data processing results from the format dimension using preset format matching rules, and obtain a format score result;

[0253] The second evaluation submodule is used to score the data processing results from the dimension of topic expression by calculating the consistency score between the content and topic of the data processing results, and to obtain the topic expression score result.

[0254] The third evaluation submodule is used to sort the data processing results by confusion level and score the data processing results from the fluency dimension to obtain a fluency score result.

[0255] The fourth evaluation submodule is used to score the data processing results from the repeatability dimension using preset repeatability evaluation rules, and obtain the repeatability score result.

[0256] The fifth evaluation submodule is used to score the data processing results from the prosody dimension using pre-set rhyme judgment rules, and obtain the prosody score result;

[0257] The sixth evaluation submodule is used to score the data processing results from the perspective of rhetorical skills using a pre-set rhetorical judgment method, and obtain the rhetorical score result.

[0258] The evaluation result acquisition submodule is used to perform weighted calculations on the format score result, the topic expression score result, the fluency score result, the repetition score result, the rhythm score result, and the rhetoric score result according to a preset weight ratio to obtain the data processing result.

[0259] Optionally, the device further includes:

[0260] The first weight ratio adjustment module is used to adjust the weight ratio according to the data processing result;

[0261] The second weight ratio adjustment module is used to adjust the weight ratio based on user feedback information;

[0262] The third weight ratio adjustment module is used to adjust the weight ratio by iteratively optimizing different combinations of the weight ratio through a preset optimization algorithm.

[0263] Based on the same inventive concept, another embodiment of this application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the data processing method as described in any of the above embodiments of this application.

[0264] Based on the same inventive concept, another embodiment of this application provides an electronic device. Figure 8 This is a schematic diagram of an electronic device 800 according to an embodiment of this application, including a memory 802, a processor 801, and a computer program stored in the memory and executable on the processor. When executed by the processor, the program implements the steps in the data processing method described in any of the above embodiments of this application.

[0265] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0266] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0267] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0268] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0269] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0270] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0271] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0272] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0273] The above provides a detailed description of the model evaluation system, data processing method, apparatus, device, and medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A model evaluation system, characterized in that, The system includes: A data processing unit is configured to invoke at least one first model to execute a data processing task and obtain at least one data processing result, wherein the at least one first model is a different model; the first model is an artificial intelligence model that executes the data processing task; and the data processing result is text. The result evaluation unit is used to evaluate the at least one data processing result using the second model, obtain at least one evaluation result, and output the evaluation result to the user page. The evaluation result is used to compare the differences between the data processing results of different first models. The second model is a model for evaluating text. The user feedback acquisition unit is used to acquire user feedback information on the evaluation results, so that the second model can optimize the evaluation logic based on the feedback information. The result evaluation unit includes: The dimension evaluation subunit is used to evaluate the data processing results from at least one preset dimension to obtain the evaluation results; The system also includes: The evaluation result display unit is used to display the evaluation score corresponding to each of the first models, and the evaluation reason corresponding to the evaluation score; An image generation unit is configured to calculate the average score of the evaluation scores corresponding to all the first models after all the first models have completed the data processing task, and generate a corresponding score distribution map, and mark the average score in the score distribution map. The evaluation history display unit is used to display the data processing results corresponding to each data processing task executed by the first model, and the evaluation record corresponding to the data processing results. The evaluation record includes the task content of the data processing task, the evaluation result corresponding to the data processing result, and the evaluation reason.

2. The model evaluation system according to claim 1, characterized in that, The user feedback acquisition unit includes at least one of the following: The first user feedback acquisition subunit is used to acquire the user's evaluation selection for each of the evaluation results, and the evaluation selection includes any of the following: evaluation too high, evaluation normal, evaluation too low; The second user feedback acquisition subunit is used to acquire the user's evaluation content for each of the evaluation results, the evaluation content including the score and evaluation information for each evaluation result.

3. The model evaluation system according to claim 1, characterized in that, The system also includes: The model loading unit is used to load the corresponding first model based on the model information input by the user. A status setting unit is used to set the usage status of the first model, the usage status including an enabled status and a deprecated status; A model information display unit is used to display the model information of the first model; The parameter configuration unit is used to configure the parameters of the first model according to the model configuration parameters.

4. The model evaluation system according to claim 1, characterized in that, The system also includes: The file receiving unit is used to receive a task file through a first interface, wherein the task file includes task data corresponding to the data processing task. The result file acquisition unit is used to output the result file corresponding to the data processing task through the second interface. The result file includes the data processing result of the first model and the evaluation result of the second model.

5. A data processing method, characterized in that, The method is based on the model evaluation system according to any one of claims 1 to 4, and includes: Receive task data corresponding to the data processing task; The data processing task is executed by at least one first model to obtain at least one data processing result; the first model is an artificial intelligence model that executes the data processing task; the data processing result is text. The at least one data processing result is input into the second model; the second model is a model for evaluating text. The data processing results are evaluated from at least one preset dimension using the second model to obtain at least one evaluation result; Based on the evaluation results, a target data processing result is determined from the at least one data processing result.

6. The data processing method according to claim 5, characterized in that, Before obtaining at least one data processing result by performing the data processing task through at least one first model, the method further includes: Receive the model information of the first model to be loaded; If the model address included in the model information is a local path, the first model is obtained from the storage location corresponding to the local path; If the model address included in the model information is a program interface address, the first model is downloaded through the program interface address; Based on the model parameter types included in the model information, the corresponding model parameter values ​​are input into the first model; Set the first model to enabled.

7. The data processing method according to claim 5, characterized in that, The method further includes: The data processing results corresponding to the first model, the data processing results, and the corresponding evaluation reasons are stored in the evaluation record corresponding to the first model. Upon receiving a request to view the evaluation record corresponding to the first model, the evaluation record is displayed.

8. The data processing method according to claim 5, characterized in that, The method further includes: Upon receiving a re-evaluation instruction, the at least one data processing result is re-evaluated using the second model.

9. The data processing method according to claim 5, characterized in that, The step of evaluating the data processing results from at least one preset dimension using the second model to obtain at least one evaluation result includes: Using preset format matching rules, the data processing results are scored from the format dimension to obtain a format score result; By calculating the consistency score between the content and the topic of the data processing results, the data processing results are scored from the perspective of topic expression, and a topic expression score result is obtained. The data processing results are sorted by confusion level, and the data processing results are scored from the fluency dimension to obtain a fluency score result; Using preset repetition evaluation rules, the data processing results are scored from the repetition dimension to obtain repetition score results; Using pre-set rhyme judgment rules, the data processing results are scored from the prosody dimension to obtain a prosody score result; Using a pre-set rhetoric judgment method, the data processing results are scored from the perspective of rhetoric skills to obtain rhetoric score results; The format score, the theme expression score, the fluency score, the repetition score, the rhythm score, and the rhetoric score are weighted and calculated according to a preset weight ratio to obtain the data processing result.

10. The data processing method according to claim 9, characterized in that, The method further includes: Based on the evaluation results, the weighting ratios are adjusted. The weight ratios are adjusted based on user feedback. The weight ratio is adjusted by iteratively optimizing different combinations of the weight ratios using a preset optimization algorithm.

11. A data processing apparatus, characterized in that, The apparatus is based on the model evaluation system according to any one of claims 1 to 4, comprising: The task data receiving module is used to receive task data corresponding to data processing tasks. A data processing result acquisition module is used to execute a data processing task through the at least one first model to obtain at least one data processing result; the first model is an artificial intelligence model that executes the data processing task; the data processing result is text; A data processing result input module is used to input the at least one data processing result into a second model; the second model is a model for evaluating text. The data processing result evaluation module is used to evaluate the data processing result from at least one preset dimension using the second model, and obtain at least one evaluation result; A data processing result determination module is used to determine a target data processing result from the at least one data processing result based on the evaluation result.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 5 to 10.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 5 to 10.

Citation Information

Patent Citations

  • Assessment method and device of large language model and computer equipment

    CN118535443A

  • Image evaluation model training method and device, image processing method and device, computer equipment and readable storage medium

    CN118747297A