Model capability classification evaluation method and device of large model, equipment and medium

By building a refined capability classification system and standardized evaluation process, the problem of inconsistent evaluation standards of large language models is solved, the accuracy and reliability of evaluation results are achieved, and the evaluation efficiency and comparability are improved.

CN120408117APending Publication Date: 2025-08-01SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510708240.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The lack of a systematic capability classification framework in the prior art has led to inconsistent evaluation standards for large language models, difficulty in making horizontal comparisons, and uneven input data quality affects the reliability of the evaluation results.

Method used

A large model capability classification evaluation method is provided. By determining the task evaluation type, generating formatted test task data, calling model parameter loading functions and performing manual evaluation, a refined capability classification system and standardized evaluation process are constructed.

Benefits of technology

A unified standard for the assessment of large-scale language models is realized, the accuracy and reliability of evaluation results are improved, labor and time costs are reduced, and the efficiency and comparability of the evaluation process are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408117A_ABST
    Figure CN120408117A_ABST
Patent Text Reader

Abstract

The invention discloses a model capability classification evaluation method and device for a large model, equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: determining a task evaluation type of a to-be-evaluated large model, generating corresponding test task data based on a sub-capability evaluation item and a specific test scene, and carrying out the preformatting processing of the test task data, the formatted target test task data is obtained; inputting the target test task data into the to-be-evaluated large model, and calling a model parameter loading function, so that the to-be-evaluated large model loads corresponding model parameters and then performs task processing on the target test task data to obtain a test result index; and performing manual evaluation on the test result index to obtain a corresponding model capability evaluation result, and optimizing the to-be-evaluated large model by using the model capability evaluation result. And the model capability of the large model under different test tasks in different scenes can be accurately evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method, device, equipment and medium for classifying and evaluating the model capabilities of large models. Background Art

[0002] With the rapid development of large language models, the evaluation of model capabilities has become increasingly important. In the prior art, there is a lack of a systematic ability classification framework, resulting in inconsistent model evaluation criteria and making it difficult to conduct horizontal comparisons. In current evaluation methods, the quality of input data directly affects the reliability of evaluation results. The test data formats from different sources are not unified and the quality varies, leading to deviations in evaluation results.

[0003] Therefore, accurately evaluating the model capabilities of large models under different scenarios and different test tasks is a technical problem to be solved in this field. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for classifying and evaluating the model capabilities of large models, which can accurately evaluate the model capabilities of large models under different scenarios and different test tasks. The specific solutions are as follows:

[0005] In the first aspect, the present application discloses a method for classifying and evaluating the model capabilities of a large model, including:

[0006] Determine the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction;

[0007] Generate corresponding test task data based on the sub-ability evaluation items and the specific test scenario, and perform pre-formatting processing on the test task data to obtain the target test task data after formatting processing;

[0008] Input the target test task data into the large model to be evaluated, and call the model parameter loading function, so that the large model to be evaluated can load the corresponding model parameters and perform task processing on the target test task data to obtain corresponding test result indicators; where the model parameter loading function is a function for loading configuration parameters corresponding to the test task data under the specific test scenario;

[0009] Perform manual evaluation on the test result indicators to obtain the corresponding model ability evaluation result, so as to optimize the large model to be evaluated using the model ability evaluation result.

[0010] Optionally, before determining the task evaluation type of the large model to be evaluated, it further includes:

[0011] Construct the to-be-selected ability evaluation directions including content generation ability, language understanding ability, reasoning ability, knowledge Q&A ability, mathematical ability, programming ability, and multi-modal ability;

[0012] Among them, the content generation ability includes any one or several sub-ability evaluation items in content writing, content expansion, content rewriting, content continuation, content imitation, and other generation types; the language understanding ability includes any one or several sub-ability evaluation items in information extraction, content summarization, multilingual translation, multi-turn dialogue, intent recognition, Chinese semantic understanding, traditional culture understanding, and text error correction; the knowledge Q&A ability includes any one or several sub-ability evaluation items in life common sense, natural science, social science, humanities, medicine, geography, audio-visual entertainment Q&A, and encyclopedia knowledge completion; the reasoning ability includes any one or several sub-ability evaluation items in logical reasoning, common sense reasoning, causal reasoning, rule reasoning, and hypothesis testing; the mathematical ability is the ability for mathematical calculation-related tasks; the programming ability is the ability for code generation and analysis; the multi-modal ability is the comprehensive ability of audio, video, and image.

[0013] Optionally, determine the task evaluation type of the to-be-evaluated large model, including:

[0014] Select the ability evaluation direction of the to-be-evaluated large model from the to-be-selected ability evaluation directions, and determine several sub-ability evaluation items of the to-be-evaluated large model according to the selected ability evaluation direction to construct the task evaluation type of the to-be-evaluated large model.

[0015] Optionally, perform pre-formatting processing on the test task data to obtain the target test task data after formatting processing, including:

[0016] Perform context-aware cleaning on the test task data to obtain the content-type labeled task data after cleaning;

[0017] Perform fault tolerance processing on the non-standard HTML task data in the test task data to obtain the non-standard HTML task data after cleaning;

[0018] Perform encoding conversion and garbled code repair on the encoded data in the test task data to obtain the encoded data after cleaning;

[0019] Perform standardization processing on the content-type labeled task data after cleaning, the non-standard HTML task data after cleaning, and the encoded data after cleaning to obtain the target test task data after formatting processing.

[0020] Optionally, perform fault tolerance processing on the non-standard HTML task data in the test task data to obtain the non-standard HTML task data after cleaning, including:

[0021] Perform layer-by-layer structural error repair processing on fragmented HTML task data or unstructured text HTML task data using regular expressions and heuristic rules to obtain the first cleaned non-standard HTML task data;

[0022] Perform multi-level filtering processing on special character task data or non-text task data within a preset encoding range, including character deletion, character replacement, character retention, and binary data entropy value detection, to obtain the second cleaned non-standard HTML task data;

[0023] Determine the cleaned non-standard HTML task data based on the first cleaned non-standard HTML task data and the second cleaned non-standard HTML task data.

[0024] Optionally, perform standardization processing on the cleaned content-based tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the target test task data after formatting processing, including:

[0025] Perform conversion processing on the non-UTF-8 task data in the cleaned content-based tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the first standardized task data after transcoding standardization;

[0026] Perform unified mapping conversion on the punctuation marks in the cleaned content-based tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the second standardized task data after mapping standardization;

[0027] Perform unified Arabic numeral format conversion processing on the different digital format task data in the cleaned content-based tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the third standardized task data after digital format standardization;

[0028] Perform multi-time zone normalization and fuzzy date parsing processing on the time task data in the cleaned content-based tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the fourth standardized task data after time standardization;

[0029] Determine the target test task data after formatting processing based on the first standardized task data, the second standardized task data, the third standardized task data, and the fourth standardized task data.

[0030] Optionally, perform manual evaluation on the test result indicators to obtain the corresponding model ability evaluation results, including:

[0031] Determine the standard task execution result corresponding to the test task data;

[0032] Determine whether the test result metrics corresponding to the test task data are consistent with the standard task execution results;

[0033] Based on the judgment result and the pre-saved consistent scores and / or inconsistent scores, determine the model ability evaluation result for evaluating the large model to be evaluated.

[0034] In a second aspect, the present application discloses a model ability classification evaluation device for a large model, including:

[0035] A type determination module for determining the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction;

[0036] A data processing module for generating corresponding test task data based on the sub-ability evaluation items and specific test scenarios, and performing pre-formatting processing on the test task data to obtain the target test task data after formatting processing;

[0037] A task execution module for inputting the target test task data into the large model to be evaluated, and calling the model parameter loading function, so that after the large model to be evaluated loads the corresponding model parameters, it performs task processing on the target test task data to obtain the corresponding test result metrics; where the model parameter loading function is a function for loading configuration parameters corresponding to the test task data in a specific test scenario;

[0038] A result evaluation module for performing manual evaluation on the test result metrics to obtain the corresponding model ability evaluation result, so as to optimize the large model to be evaluated using the model ability evaluation result.

[0039] In a third aspect, the present application discloses an electronic device, including:

[0040] A memory for storing a computer program;

[0041] A processor for executing the computer program to implement the steps of the model ability classification evaluation method for the large model disclosed above.

[0042] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the model ability classification evaluation method for the large model disclosed above.

[0043] It can be seen that the present application discloses a method for classifying and evaluating the model capabilities of a large model, including: determining the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction; generating corresponding test task data based on the sub-ability evaluation items and specific test scenarios, and performing pre-formatting processing on the test task data to obtain the target test task data after formatting processing; inputting the target test task data into the large model to be evaluated, and invoking the model parameter loading function, so that the large model to be evaluated can load the corresponding model parameters and perform task processing on the target test task data to obtain corresponding test result indicators; where the model parameter loading function is a function for loading configuration parameters corresponding to the test task data under the specific test scenario; performing manual evaluation on the test result indicators to obtain the corresponding model ability evaluation result, so as to optimize the large model to be evaluated using the model ability evaluation result. Thus, it can be seen that the pre-data formatting processing module effectively improves the quality of the input data. Through data cleaning, standardization, and structuring processing, data noise and bias are reduced, ensuring the accuracy and reliability of the evaluation results. The standardized evaluation process and the optimization of manual scoring evaluation make the evaluation process more efficient, reduce labor and time costs, and improve the throughput and speed of the evaluation. Through the ability classification system and standardized evaluation process proposed by the present invention, a unified standard for evaluating the capabilities of large language models is achieved, improving the comparability of evaluation results between different models. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0045] Figure 1 It is a flowchart of a method for classifying and evaluating the model capabilities of a large model disclosed in the present application;

[0046] Figure 2 It is a flowchart of a specific method for classifying and evaluating the model capabilities of a large model disclosed in the present application;

[0047] Figure 3 It is a schematic structural diagram of a device for classifying and evaluating the model capabilities of a large model disclosed in the present application;

[0048] Figure 4 It is a structural diagram of an electronic device disclosed in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] With the rapid development of large language models, the evaluation of model capabilities has become increasingly important. In the prior art, there is a lack of a systematic ability classification framework, resulting in inconsistent model evaluation criteria and making it difficult to conduct horizontal comparisons. The present invention proposes a complete classification system and method flow, filling this technical gap.

[0051] In the existing evaluation methods, the quality of the input data directly affects the reliability of the evaluation results. The test data from different sources have inconsistent formats and uneven quality, leading to deviations in the evaluation results.

[0052] Therefore, the present invention provides a model ability classification and evaluation scheme for large models, which can accurately evaluate the model capabilities of large models under different scenarios and different test tasks.

[0053] Referring to Figure 1 As shown, the embodiments of the present invention disclose a model ability classification and evaluation method for large models, including:

[0054] Step S11: Determine the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction.

[0055] In this embodiment, before determining the task evaluation type of the large model to be evaluated, it further includes: constructing selectable ability evaluation directions including content generation ability, language understanding ability, reasoning ability, knowledge Q&A ability, mathematical ability, programming ability, and multi-modal ability; where the content generation ability includes any one or several sub-ability evaluation items among content writing, content expansion, content rewriting, content continuation, content imitation, and other generation types; the language understanding ability includes any one or several sub-ability evaluation items among information extraction, content summarization, multilingual translation, multi-turn dialogue, intention recognition, Chinese semantic understanding, traditional culture understanding, and text error correction; the knowledge Q&A ability includes any one or several sub-ability evaluation items among common sense of life, natural science, social science, humanities, medicine, geography, audio-visual entertainment Q&A, and encyclopedia knowledge completion; the reasoning ability includes any one or several sub-ability evaluation items among logical reasoning, common sense reasoning, causal reasoning, rule reasoning, and hypothesis testing; the mathematical ability is the ability for tasks related to mathematical calculations; the programming ability is the ability for code generation and analysis; the multi-modal ability is the comprehensive ability of audio, video, and graphics.

[0056] It is understandable that before adjusting the large model through the Fine-tuning technology to adapt to specific ability classification tasks, an evaluation direction for the abilities to be selected is constructed. Among them, the corresponding relationships of each ability section in the evaluation direction for the abilities to be selected are as follows:

[0057] The content generation ability specifically includes any one or several of the following ability evaluation items: content writing, content expansion, content rewriting, content continuation, content imitation, and other generation types (generating poems, novels, couplets, news, descriptions, commercial advertisements), etc.

[0058] The language understanding ability specifically includes any one or several of the following ability evaluation items: information extraction (entity extraction, keyword extraction), content summarization, content summary, multilingual translation, multi-turn dialogue (except for mathematics and code types), intention recognition, sentiment analysis, news topic classification, content evaluation, understanding of traditional Chinese culture, Chinese semantic understanding, text error correction, analyzing a given passage (background text) or answering questions, and determining whether two questions are about the same thing, etc.

[0059] The knowledge Q&A ability specifically includes any one or several of the following ability evaluation items: common sense of life, natural science (physics, chemistry, earth science, astronomy, biology, computer science, engineering, etc.), social science (economics, politics, law, ethics, sociology, psychology, education, management, anthropology, folklore, journalism, communication, etc.), humanities (literature, history, philosophy, art), medicine, geography, multiple-choice questions (except for mathematics / programming multiple-choice questions), audio-visual entertainment Q&A, encyclopedia Q&A, knowledge supplementation, etc.

[0060] The reasoning ability specifically includes any one or several of the following ability evaluation items: logical reasoning, common sense reasoning, causal reasoning (determining the causal relationship between variables, such as whether A causes B), rule reasoning (logical puzzles), hypothesis testing (inferring based on... / deducing based on.... / judging the truth or falsehood of the second question based on the first question), prediction-related, etc.

[0061] The mathematical ability is specifically the task ability related to mathematical calculations.

[0062] The programming ability specifically includes that any content with code can be determined as programming ability or content related to programming.

[0063] The multi-modal ability is specifically the comprehensive ability of audio-visual (image recognition, object localization, image-text matching, video summarization, text-to-image, speech synthesis).

[0064] In this embodiment, determining the task evaluation type of the large model to be evaluated includes: selecting the ability evaluation direction of the large model to be evaluated from the to-be-selected ability evaluation directions, and determining several sub-ability evaluation items of the large model to be evaluated according to the selected ability evaluation direction, so as to construct the task evaluation type of the large model to be evaluated. In this way, on each of the above-mentioned preset sub-ability evaluation items and the to-be-selected ability evaluation directions (ability classification system) constructed based on each ability evaluation direction, select the current ability evaluation direction of the current large model to be evaluated, and several sub-ability evaluation items on the current ability evaluation direction, so as to obtain the task evaluation type of the current large model to be evaluated. For example, if the current large model to be evaluated is a Chinese semantic understanding large model and only the Chinese semantic understanding ability is evaluated this time, then the selected ability evaluation direction is language understanding ability, and the sub-ability evaluation item is Chinese semantic understanding; if the current large model to be evaluated is a Chinese semantic understanding large model and this time it is evaluated for Chinese semantic understanding ability and content generation ability based on Chinese semantic understanding, then the selected ability evaluation directions are language understanding ability and content generation ability, and the sub-ability evaluation items are Chinese semantic understanding, content writing, content expansion, content rewriting, etc. It can be understood that in order to conduct targeted ability evaluation on the large model, one or more ability evaluation directions and one or more sub-ability evaluation items can be selected according to the above-mentioned pre-designed ability classification system according to the actual evaluation requirements, so as to obtain a targeted task evaluation type.

[0065] Step S12: Generate corresponding test task data based on the sub-ability evaluation items and specific test scenarios, and perform pre-formatting processing on the test task data to obtain the target test task data after formatting processing.

[0066] In this embodiment, if the specific test scenario is to only evaluate the Chinese semantic understanding ability, then according to the sub-ability item being Chinese semantic understanding, design test tasks to cover sub-abilities such as polysemy, idioms, ancient Chinese, and context reasoning. Specific test task examples are shown in Table 1 below:

[0067] Table 1

[0068]

[0069] In this embodiment, if the specific test scenario is to evaluate several ability evaluation directions at the same time, for example: Chinese semantic understanding ability, content generation ability, and multi-modal ability, then according to the sub-ability items being Chinese semantic understanding ability, content generation ability, and multi-modal comprehensive ability, similarly design test tasks to cover sub-abilities such as polysemy, idioms, ancient Chinese, and context reasoning. The example of the specific test task regarding Chinese semantic understanding is shown in Table 1 above, and the example of the specific test task regarding content generation ability is shown in Table 2 below:

[0070] Table 2

[0071]

[0072] It can be seen from this that through different sub - ability evaluation items and specific test scenarios, the input - reference answers in the above table are correspondingly generated as the corresponding test task data. The above - mentioned test task data can be collected from channels such as educational question banks, literary works, social media, or high - quality reference answers can be manually written (covering reasonable answers).

[0073] In this embodiment, after the test task data is generated, the test task data is input into the data pre - processing module. The data pre - processing module is responsible for the preliminary sorting and processing of the original data. The main tasks include data cleaning, duplicate removal, format conversion, etc., to ensure the consistency and standardization of the data, laying a solid foundation for subsequent ability classification.

[0074] In this embodiment, context - aware cleaning is performed on the test task data to obtain the cleaned content - type labeled task data. It can be understood that context - aware cleaning of the test task data is performed based on DOM (Document Object Model) tree parsing, retaining content - type labels, such as <code>、 <pre>, only strip style and script tags, and automatically complete unclosed tags through context speculation to obtain the cleaned content-type tag task data.

[0075] In this embodiment, fault tolerance processing is performed on the non-standard HTML task data in the test task data to obtain the cleaned non-standard HTML task data; it can be understood that for non-standard HTML (HyperText Markup Language), such as fragmented text, regular expressions + heuristic rules are used for fault tolerance processing to avoid damaging the original data structure to obtain the cleaned non-standard HTML task data.

[0076] Specifically, use regularization expressions and heuristic rules to perform layer-by-layer structure error repair processing on fragmented HTML task data or unstructured text HTML task data to obtain the first cleaned non-standard HTML task data; perform multi-level filtering processing including character deletion, character replacement, character retention, and binary data entropy value detection on special character task data or non-text task data according to the preset encoding range to obtain the second cleaned non-standard HTML task data; determine the cleaned non-standard HTML task data based on the first cleaned non-standard HTML task data and the second cleaned non-standard HTML task data. It can be understood that a Unicode (Universal Character Set) multi-level filtering matrix is constructed to classify fragmented HTML task data or unstructured text HTML task data layer by layer to obtain the first cleaned non-standard HTML task data; directly delete control characters, replace private area characters with placeholders, calculate the entropy value of text fragments, regard high-entropy value fragments (such as Base64-encoded images) as binary data, and use entropy value detection for automatic binary data identification and removal to obtain the second cleaned non-standard HTML task data.

[0077] In this embodiment, the encoded data in the test task data is subjected to encoding conversion and garbled code repair to obtain the cleaned encoded data; the cleaned content-type label task data, the cleaned non-standard HTML task data, and the cleaned encoded data are standardized to obtain the target test task data after formatting. Specifically, common encodings such as UTF-8 (Unicode Transformation Format-8-bit), GBK (GuoBiaoKuozhan), and ISO-8859 (International Standard Organization-8859) are detected in sequence, and byte pattern matching is used to improve the accuracy. For texts with incorrect encodings, they are restored through byte recombination + language model error correction (such as BERT (Bidirectional Encoder Representations from Transformers)) to obtain the cleaned encoded data. Then, the following standardization processing methods are used to process all the above-cleaned data. Specifically:

[0078] The non-UTF-8 task data in the cleaned content-type label task data, the cleaned non-standard HTML task data, and the cleaned encoded data is converted to obtain the first standardized task data after transcoding standardization;

[0079] The punctuation marks in the cleaned content-type label task data, the cleaned non-standard HTML task data, and the cleaned encoded data are uniformly mapped and converted to obtain the second standardized task data after mapping standardization;

[0080] The different digital format task data in the cleaned content-type label task data, the cleaned non-standard HTML task data, and the cleaned encoded data is uniformly converted to the Arabic numeral format to obtain the third standardized task data after digital format standardization;

[0081] The time task data in the cleaned content-type label task data, the cleaned non-standard HTML task data, and the cleaned encoded data is normalized across multiple time zones and fuzzy date parsing is performed to obtain the fourth standardized task data after time standardization;

[0082] The target test task data after formatting is determined based on the first standardized task data, the second standardized task data, the third standardized task data, and the fourth standardized task data.

[0083] The processing processes of the above four types of standardized task data are as follows:

[0084] Unified text encoding (UTF-8): Real-time transcoding pipeline, automatically detects and converts non-UTF-8 text, preserving the original encoding metadata. Principle: Selects the optimal encoding based on the BOM marker or character distribution probability.

[0085] Normalize punctuation: Construct a punctuation mapping table (e.g., map Chinese commas to English commas), combined with context rules (to avoid incorrect conversion of decimal points). Principle: Relies on the bidirectional maximum matching algorithm to avoid ambiguity.

[0086] Unify the digital representation format: Regular expressions identify different formats (scientific notation / Chinese numerals), unified into Arabic numerals + units (e.g., "12,000 converted to 12000"). Principle: Parses compound expressions based on the digital morphological syntax tree.

[0087] Standardize the date and time format: Automatic normalization of multiple time zones (converted to UTC+0), fuzzy date parsing (e.g., "last Q3" → "2023-07-01"). Principle: Combines a rule engine and a timestamp interpolation algorithm.

[0088] In addition, the standardized task data can also be enhanced through the data augmentation sub-module to obtain the enhanced task data for the large model ability evaluation test. The specific enhancement steps are as follows:

[0089] Generate variants through synonym replacement: Based on semantic similar word replacement in the ConceptNet knowledge graph, avoiding changing the original meaning.

[0090] Transform the syntactic structure: Use the dependency syntax tree for active / passive voice conversion and sentence splitting and merging.

[0091] Generate multilingual parallel data: Adopt back-translation: Chinese to English to German to Chinese, generating multilingual pairs with consistent semantics.

[0092] Automatically generate negative samples: Adversarial generation: Constructs plausible incorrect samples through GPT-3.5 to enhance the model's robustness.

[0093] Step S13: Input the target test task data into the large model to be evaluated, and call the model parameter loading function so that the large model to be evaluated can load the corresponding model parameters and perform task processing on the target test task data to obtain the corresponding test result metrics; where the model parameter loading function is the function of loading the configuration parameters corresponding to the test task data for executing specific test scenarios.

[0094] In this embodiment, the large model to be evaluated includes a basic model layer, and the basic model layer adopts a hierarchical model architecture. Therefore, the optimal basic model is selected according to the task domain of the target test task data.

[0095] For the text domain, BERT-wwm and RoBERTa-large (Chinese optimized version) are adopted. Among them, BERT-wwm is a Chinese whole-word masking model, which is good at dealing with word segmentation boundary problems; RoBERTa-large is dynamic masking + larger batch training, which improves semantic representation ability.

[0096] For the multi-modal domain, Flamingo-80B is adopted, where Flamingo-80B supports cross-modal alignment of text and images and is suitable for tasks such as image description and visual question answering.

[0097] For the code domain, CodeLlama-34B is adopted, where CodeLlama-34B is optimized for code generation and understanding and supports multiple languages such as Python / Java.

[0098] In this way, different models in the above domains are managed through the Model Zoo, the model parameter loading function is called, and the adaptation module is loaded in real time according to the task requirements (such as switching to the code model to process programming tasks) to achieve plug-and-play.

[0099] When the large model to be evaluated loads the corresponding model parameters, a 3D parallel strategy is executed. Among them, data parallelism means splitting the data batches to multiple GPUs (Graphics Processing Unit), such as 8-card training, and each card processes 1 / 8 of the data. Model parallelism means splitting the model layers to different devices, such as distributing the attention heads and FFN layers of the Transformer to different GPUs. Pipeline parallelism: Execute in segments by layer, such as GPU1 processes layers 1-5, and GPU2 processes layers 6-10, reducing the video memory occupancy. During the execution of the 3D parallel strategy, 16-bit floating-point is used to save gradients, saving 40% of the video memory compared to FP32 while maintaining numerical stability. At the same time, the accumulation step is automatically increased according to the remaining video memory, such as adjusting from 4 steps to 8 steps, and the effective batch size is increased to 1024.

[0100] During the training process of the large model, since the large model is a hierarchical fine-tuning architecture, the parameters of the general layer keep the underlying parameters of the pre-trained model (such as the first 6 layers of BERT) fixed to retain general language knowledge; the domain adaptation layer LoRA (Low-Rank Adaptation) inserts a low-rank adaptation matrix in the Transformer layer, such as rank = 64, and only fine-tunes the adapter parameters to reduce the amount of calculation. Different output heads are dynamically activated according to the task type (such as classification, generation) to avoid task interference.

[0101] In this embodiment, after the above model training is completed, the target test task data is input into the large model to be evaluated, and the model parameters corresponding to the task domain are called to process the target test task data to obtain the corresponding test result metrics.

[0102] Step S14: Perform manual evaluation on the test result metrics to obtain the corresponding model ability evaluation result, so as to optimize the large model to be evaluated by using the model ability evaluation result.

[0103] In this embodiment, determine the standard task execution result corresponding to the test task data; determine whether the test result metrics corresponding to the test task data are consistent with the standard task execution result; based on the judgment result and the pre-saved result consistency score and / or the score for inconsistent results, determine the model ability evaluation result for evaluating the large model to be evaluated. It can be understood that a certain proportion of samples are randomly selected from the classified test result metrics for manual review. The data set is divided into multiple subsets for multiple training and testing to evaluate the stability of the model. Conduct in-depth analysis on the misclassified samples to find out the reasons for the errors. Among them, the quality inspection process is as follows:

[0104] (1) Sample extraction: Randomly extract classified data samples according to a preset ratio.

[0105] (2) Manual review: Professional quality inspection personnel review the extracted samples according to the classification criteria.

[0106] (3) Result comparison: Compare the manual review result with the model classification result, and record indicators such as accuracy and precision. Among them, the accuracy is the ratio of the number of correctly classified samples to the total number of samples; the precision is the ratio of the number of samples correctly classified as a certain class to the total number of samples classified as that class.

[0107] (4) Error analysis: Mark and analyze the misclassified samples to find out the reasons for the misclassification.

[0108] (5) Performance evaluation: Calculate the comprehensive performance of the model according to the evaluation indicators.

[0109] After obtaining the model ability evaluation result, perform the subsequent feedback process. The feedback process is as follows:

[0110] (1) Regular report: Regularly generate a quality inspection report, including content such as classification accuracy and error analysis.

[0111] (2) Problem feedback: Timely feedback the problems found in the quality inspection process to the model training and classification marking teams.

[0112] (3) Standard optimization: Adjust and optimize the classification criteria according to the quality inspection results to improve the classification accuracy.

[0113] (4)Process improvement: Address the issues found in quality inspection and improve the data preprocessing, model training, and classification and labeling processes.

[0114] (5)Model iteration: Iteratively optimize the model based on the feedback results to improve the classification performance.

[0115] As Figure 2 shown, the specific model ability classification and evaluation process is as follows:

[0116] Data preprocessing is used to clean, standardize, and structure the original data to generate a high-quality training set.

[0117] Model training: Based on the preprocessed data, train a basic large model (such as BERT, GPT, etc.) to learn general features.

[0118] Model fine-tuning: Further adjust the model parameters on specific task data to improve domain adaptability.

[0119] Classification prediction: Input the input data into the fine-tuned model and output the preliminary classification results (such as task type labels).

[0120] Manual verification: If the verification is successful, that is, the classification result is manually confirmed to be correct, proceed to the next stage; if the verification fails, that is, the manual label is incorrect, trigger a feedback mechanism, such as adjusting the model or data.

[0121] Ability classification: According to the verified prediction results, classify the data into preset ability dimensions, such as Chinese semantic understanding, multi-modal generation.

[0122] Ability label result: Output the structured ability labels.

[0123] Result summary: Statistically analyze the evaluation indicators of each ability dimension and generate a final report.

[0124] It can be seen that through the ability classification system and standardized evaluation process proposed by the present invention, a unified standard for evaluating the capabilities of large language models has been achieved, greatly improving the comparability of evaluation results between different models and providing a reliable reference basis for the industry. The pre-data formatting processing module effectively improves the quality of input data. Through data cleaning, standardization, and structuring, data noise and bias are reduced, ensuring the accuracy and reliability of evaluation results. The standardized evaluation process and the optimization of manual scoring evaluations make the evaluation process more efficient, reducing labor and time costs and increasing the throughput and speed of evaluation. The seven major categories of ability classification frameworks and the detailed classification of sub-abilities ensure that the capabilities of large language models in multiple dimensions can be comprehensively evaluated without missing any important ability indicators. Through the refined evaluation criteria and quality verification process, the present invention can help model developers discover and address deficiencies in the model's specific capabilities, thereby optimizing the model targeted and improving the overall performance. The modular design enables the ability classification system to have good scalability, and new ability dimensions can be easily added to adapt to the ever-evolving technological and market demands. The optimizations for Chinese characteristics, such as evaluation items for understanding Chinese culture and Chinese semantic understanding, help improve the model's performance in the Chinese context and promote the localization application of large language models in the Chinese market. Through bias filtering in the quality verification process, the present invention helps identify and eliminate potential biases in the data, thereby reducing discrimination and bias problems that may occur in the model's application.

[0125] It can be seen that the present application discloses a method for classifying and evaluating the model capabilities of a large model, including: determining the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction; generating corresponding test task data based on the sub-ability evaluation items and specific test scenarios, and performing pre-formatting processing on the test task data to obtain the target test task data after formatting processing; inputting the target test task data into the large model to be evaluated, and invoking the model parameter loading function so that the large model to be evaluated can load the corresponding model parameters and perform task processing on the target test task data to obtain corresponding test result indicators; where the model parameter loading function is a function for loading configuration parameters corresponding to the test task data in the specific test scenario; performing manual evaluation on the test result indicators to obtain the corresponding model ability evaluation result, so as to optimize the large model to be evaluated using the model ability evaluation result. Thus, it can be seen that the pre-data formatting processing module effectively improves the quality of the input data. Through data cleaning, standardization, and structuring processing, data noise and deviation are reduced, ensuring the accuracy and reliability of the evaluation results. The standardized evaluation process and manual scoring evaluation optimization make the evaluation process more efficient, reducing labor and time costs, and improving the throughput and speed of the evaluation. Through the ability classification system and standardized evaluation process proposed by the present invention, a unified standard for evaluating the capabilities of large language models is achieved, improving the comparability of evaluation results between different models.

[0126] Referring to Figure 3 as shown, the present invention also correspondingly discloses a device for classifying and evaluating the model capabilities of a large model, including:

[0127] A type determination module 11, configured to determine the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction;

[0128] A data processing module 12, configured to generate corresponding test task data based on the sub-ability evaluation items and specific test scenarios, and perform pre-formatting processing on the test task data to obtain the target test task data after formatting processing;

[0129] A task execution module 13, configured to input the target test task data into the large model to be evaluated, and invoke the model parameter loading function so that the large model to be evaluated can load the corresponding model parameters and perform task processing on the target test task data to obtain corresponding test result indicators; where the model parameter loading function is a function for loading configuration parameters corresponding to the test task data in the specific test scenario;

[0130] A result evaluation module 14 is used to perform manual evaluation on test result metrics to obtain corresponding model ability evaluation results, so as to optimize the large model to be evaluated by using the model ability evaluation results.

[0131] It can be seen that this application discloses determining the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction; corresponding test task data generated based on the sub-ability evaluation items and specific test scenarios, and performing pre-formatting processing on the test task data to obtain the target test task data after formatting processing; inputting the target test task data into the large model to be evaluated, and calling the model parameter loading function, so that after the large model to be evaluated loads the corresponding model parameters, it performs task processing on the target test task data to obtain corresponding test result metrics; where the model parameter loading function is a function for loading configuration parameters corresponding to the test task data under the specific test scenario; performing manual evaluation on the test result metrics to obtain corresponding model ability evaluation results, so as to optimize the large model to be evaluated by using the model ability evaluation results. Thus, it can be seen that the pre-data formatting processing module effectively improves the quality of the input data. Through data cleaning, standardization, and structuring processing, it reduces data noise and bias, ensuring the accuracy and reliability of the evaluation results. The standardized evaluation process and manual scoring evaluation optimization make the evaluation process more efficient, reducing labor and time costs, and improving the throughput and speed of the evaluation. Through the ability classification system and standardized evaluation process proposed by the present invention, a unified standard for evaluating the capabilities of large language models is achieved, improving the comparability of evaluation results between different models.

[0132] Furthermore, the embodiment of this application also discloses an electronic device Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation to the scope of use of this application.

[0133] Figure 4 It is a structural schematic diagram of an electronic device 20 provided by the embodiment of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the model ability classification evaluation method of the large model disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0134] In this embodiment, the power supply 23 is used to provide operating voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and specific limitations thereof are not imposed herein; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type thereof can be selected according to specific application requirements, and specific limitations are not imposed herein.

[0135] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0136] In addition, the memory 22, as a carrier for resource storage, may be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0137] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, so as to implement the operation and processing of the massive data 223 in the memory 22 by the processor 21. It can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the model ability classification and evaluation method of the large model executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of completing other specific tasks. The data 223 may include not only the data transmitted by the external device received by the electronic device, but also the data collected by its own input / output interface 25, etc.

[0138] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the model ability classification and evaluation method of the large model disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0139] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0140] Those skilled in the art may further realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application. The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, removable disk, CD-ROM (compact disc read-only memory), or any other form of storage medium known in the technical field.

[0141] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0142] The above provides a detailed introduction to the solution provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.< / pre> < / code>

Claims

1. A method for classifying and evaluating the model capabilities of a large model, characterized in that Including: Determine the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction; Based on the sub-ability evaluation items and the corresponding test task data generated by the specific test scenario, and perform pre-formatting processing on the test task data to obtain the target test task data after formatting processing; Input the target test task data into the large model to be evaluated, and call the model parameter loading function, so that the large model to be evaluated can load the corresponding model parameters and then process the target test task data to obtain the corresponding test result indicators; where the model parameter loading function is a function for loading the configuration parameters corresponding to the test task data under the specific test scenario; Perform manual evaluation on the test result indicators to obtain the corresponding model ability evaluation result, so as to optimize the large model to be evaluated by using the model ability evaluation result.

2. The method for classifying and evaluating the model capabilities of the large model according to claim 1, wherein, Before determining the task evaluation type of the large model to be evaluated, it further includes: Construct a to-be-selected ability evaluation direction including content generation ability, language understanding ability, reasoning ability, knowledge Q&A ability, mathematical ability, programming ability, and multi-modal ability; Among them, the content generation ability includes any one or several sub-ability evaluation items in content writing, content expansion, content rewriting, content continuation, content imitation, and other generation types; the language understanding ability includes any one or several sub-ability evaluation items in information extraction, content summarization, multilingual translation, multi-turn dialogue, intention recognition, Chinese semantic understanding, traditional culture understanding, and text error correction; the knowledge Q&A ability includes any one or several sub-ability evaluation items in life common sense, natural science, social science, humanities, medicine, geography, audio-visual entertainment Q&A, and encyclopedia knowledge completion; the reasoning ability includes any one or several sub-ability evaluation items in logical reasoning, common sense reasoning, causal reasoning, rule reasoning, and hypothesis testing; the mathematical ability is the ability for mathematical calculation-related tasks; the programming ability is the ability for code generation and analysis; the multi-modal ability is the comprehensive ability of audio, video, and image.

3. The method for classifying and evaluating the model capabilities of the large model according to claim 2, wherein Determining the task evaluation type of the large model to be evaluated includes: Select the ability evaluation direction of the large model to be evaluated from the to-be-selected ability evaluation direction, and determine several sub-ability evaluation items of the large model to be evaluated according to the selected ability evaluation direction, so as to construct the task evaluation type of the large model to be evaluated.

4. The method for classifying and evaluating the model capabilities of the large model according to claim 1, wherein Performing pre-formatting processing on the test task data to obtain the target test task data after formatting processing includes: Perform context-aware cleaning on the test task data to obtain the content-type labeled task data after cleaning; Perform fault tolerance processing on the non-standard HTML task data in the test task data to obtain the non-standard HTML task data after cleaning; Perform encoding conversion and garbled repair on the encoded data in the test task data to obtain the encoded data after cleaning; Perform standardization processing on the cleaned content-type tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the target test task data after formatting processing.

5. The method for classifying and evaluating the model capabilities of the large model according to claim 4, wherein The fault tolerance processing of the non-standard HTML task data in the test task data to obtain the cleaned non-standard HTML task data includes: Using regular expressions and heuristic rules to perform layer-by-layer structural error repair processing on fragmented HTML task data or unstructured text HTML task data to obtain the first cleaned non-standard HTML task data; Perform multi-level filtering processing including character deletion, character replacement, character retention, and binary data entropy value detection on special character task data or non-text task data within a preset encoding range to obtain the second cleaned non-standard HTML task data; Determine the cleaned non-standard HTML task data based on the first cleaned non-standard HTML task data and the second cleaned non-standard HTML task data.

6. The method for classifying and evaluating the model capabilities of the large model according to claim 4, wherein, The standardization processing of the cleaned content-type tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the target test task data after formatting processing includes: Perform conversion processing on the non-UTF-8 task data in the cleaned content-type tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the first standardized task data after transcoding standardization; Perform unified mapping conversion on the punctuation marks in the cleaned content-type tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the second standardized task data after mapping standardization; Perform unified Arabic numeral format conversion processing on the different digital format task data in the cleaned content-type tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the third standardized task data after digital format standardization; Perform multi-time zone normalization and fuzzy date parsing processing on the time task data in the cleaned content-type tag task data, the cleaned non-standard HTML task data, and the cleaned encoding data to obtain the fourth standardized task data after time standardization; Determine the target test task data after formatting processing based on the first standardized task data, the second standardized task data, the third standardized task data, and the fourth standardized task data.

7. The method for classifying and evaluating the model capabilities of the large model according to claim 1, characterized in that, The manual evaluation of the test result indicators to obtain the corresponding model ability evaluation result includes: Determine the standard task execution result corresponding to the test task data; Judge whether the test result indicators corresponding to the test task data are consistent with the standard task execution result; Based on the judgment result and the pre-saved result consistency score and / or the score for inconsistent results, determine the model ability evaluation result for evaluating the large model to be evaluated.

8. An apparatus for classifying and evaluating the model capabilities of a large model, characterized in that, including: A type determination module for determining the task evaluation type of the large model to be evaluated, where the task evaluation type includes the ability evaluation direction of the large model to be evaluated and several sub-ability evaluation items in the ability evaluation direction; A data processing module for generating corresponding test task data based on the sub-ability evaluation items and a specific test scenario, and performing pre-formatting processing on the test task data to obtain target test task data after formatting processing; A task execution module for inputting the target test task data into the large model to be evaluated and invoking the model parameter loading function, so that after the large model to be evaluated loads the corresponding model parameters, it performs task processing on the target test task data to obtain corresponding test result indicators; where the model parameter loading function is a function for loading configuration parameters corresponding to the test task data in the specific test scenario; A result evaluation module for performing manual evaluation on the test result indicators to obtain corresponding model ability evaluation results, so as to optimize the large model to be evaluated by using the model ability evaluation results.

9. An electronic device, characterized in that, Comprising: A memory for storing computer programs; A processor for executing the computer program to implement the steps of the method for classifying and evaluating the model ability of the large model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing computer programs; wherein, when the computer program is executed by the processor, it implements the steps of the method for classifying and evaluating the model ability of the large model according to any one of claims 1 to 7.

Citation Information

Cited By

  • Model performance evaluation method and device, electronic equipment and medium

    CN121190915A

  • A method, apparatus, electronic device, and medium for evaluating the performance of a remote sensing model.

    CN121190915B