Method, device and equipment for improving capability of large model and medium

By constructing an evaluation system and screening and decontaminating high-quality instruction data, combined with supervised fine-tuning strategies, the generality and target capability of the large language model were improved, solving the problem of insufficient model performance in specific and cross-domain fields, and achieving stronger generalization ability and adaptability.

CN120873600APending Publication Date: 2025-10-31SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510992040.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Large language models lack generality in specific and cross-domain applications. Existing supervised fine-tuning methods rely on imbalanced data and evaluation set leakage, resulting in insufficient generalization ability and unstable performance.

Method used

Construct an evaluation system, including developing evaluation sets and unseen evaluation sets, screening and synthesizing high-quality instruction data, performing decontamination processing, and optimizing the core skills of the model through supervised fine-tuning, and periodically evaluating and adjusting the training strategy.

Benefits of technology

It significantly improves the model's performance in core skills such as knowledge retrieval, reasoning, mathematics, code generation, and instruction following, and enhances the model's versatility, generalization ability, adaptability, and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873600A_ABST
    Figure CN120873600A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device and equipment for improving the capability of a large model and a medium. An evaluation system is constructed, and the evaluation system comprises a development evaluation set and an unseen evaluation set and is used for evaluating the performance of the model on different core skills; constructing instruction data, namely screening samples from the public data set, synthesizing new data and carrying out depollution treatment to ensure that the data covers target capability and does not leak the content of the evaluation set; performing supervision fine tuning, and optimizing core skills of the model based on the instruction data, the core skills including knowledge recall, reasoning, mathematics, code generation and instruction following; and regularly evaluating model performance and adjusting a training strategy according to an evaluation result. Through a systematic evaluation system, high-quality instruction data construction and a supervision fine tuning strategy, the expression of a large-scale language model in core skills such as knowledge recall, reasoning, mathematics, code generation and instruction following is remarkably improved, and the universality and target ability of the large model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and in particular to a method, apparatus, device, and medium for enhancing the capabilities of large models. Background Technology

[0002] In recent years, Large Language Models (LLMs) have made significant progress, expanding their capabilities from simple text generation to solving complex tasks. However, in practical applications, the models' ability to target specific domains and their generalization capabilities across domains still need improvement.

[0003] Traditional supervised fine-tuning (SFT) methods typically rely on public or synthetic datasets. However, these datasets often suffer from uneven distribution and contaminated evaluation sets, leading to insufficient generalization ability of the models in practical applications. For example, in knowledge recall, the model may perform well in answering questions about common knowledge, but its recall ability drops significantly for less common or specialized knowledge. In reasoning, the model may be accurate in handling simple logical reasoning problems, but its accuracy drops drastically when faced with complex multi-step reasoning tasks. In mathematical and code generation, the model's generated results may contain errors or not conform to standards. In instruction following, the model may fail to accurately understand the user's instruction intent, resulting in generated results that do not meet the user's expectations. These problems severely limit the effectiveness and value of large language models in practical applications.

[0004] To address the aforementioned issues, this invention proposes a method for enhancing the general and target capabilities of large-scale models based on instruction data. Through systematic data construction, decontamination processing, and a multi-stage training strategy, the core skills of the model are significantly improved. By addressing both data and training aspects, this approach aims to comprehensively enhance the model's general and target capabilities, enabling it to perform exceptionally well in various application scenarios. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and medium for enhancing the capabilities of large models, which can improve the general and target capabilities of large models.

[0006] According to one aspect of the present invention, a method for enhancing the capabilities of a large model is provided, comprising:

[0007] Construct an evaluation system, which includes an developed evaluation set and an unseen evaluation set, to evaluate the model’s performance on different core skills;

[0008] Constructing instruction data involves selecting samples from public datasets, synthesizing new data, and performing decontamination processing to ensure that the data covers the target capabilities without revealing the evaluation set content;

[0009] Supervised fine-tuning is performed to optimize the core skills of the model based on the instruction data. The core skills include knowledge retrieval, reasoning, mathematics, code generation, and instruction following.

[0010] Regularly evaluate model performance and adjust training strategies based on the evaluation results.

[0011] Optionally, the construction of the evaluation system includes:

[0012] Define core skills and identify the core skills that need to be improved. These core skills include knowledge retrieval, reasoning, mathematics, code generation, instruction following, and security.

[0013] Select benchmark tests, and choose a set of benchmark test tasks for each of the core skills, wherein:

[0014] The benchmark task is divided into a development evaluation set and an unseen evaluation set. The development evaluation set is used to guide the training process, and the unseen evaluation set is used to finally evaluate the generalization ability of the model.

[0015] Optionally, the construction instruction data includes:

[0016] Data is sifted from publicly available datasets, including WildChat and OpenAssistant; instruction data is synthesized for specific skills, including generating mathematical problems or code generation tasks through templates;

[0017] The n-gram matching algorithm is used to detect and remove samples that overlap with the evaluation set to prevent training data from leaking the evaluation set content; each instruction data is reviewed to ensure that it meets the requirements of the target capability.

[0018] A unified data format is used to ensure that each instruction contains a clear task description and a corresponding correct answer; a dialogue history is constructed for multi-turn dialogue tasks.

[0019] Optionally, the supervised fine-tuning includes:

[0020] High-quality samples are selected from the constructed instruction data for supervised fine-tuning to optimize the distribution of training data and ensure balanced coverage of each core skill. During the training process, the model performance is evaluated periodically using the development evaluation set, and the data distribution is dynamically adjusted.

[0021] Based on the performance of the evaluation set, optimize the data mixing ratio and training parameters; and perform targeted optimization for specific skills such as mathematics and code generation.

[0022] Optionally, the method further includes:

[0023] We constructed separate data mixtures for different skills and found the optimal mixture ratio for the performance of a single skill through experiments.

[0024] The optimal data mixing ratios for each of the aforementioned skills are combined to form the initial mixed data;

[0025] During the iteration process, contaminated data is removed, and excessively large datasets are downsampled to balance model performance.

[0026] Optionally, the method further includes:

[0027] During training, the random seed is changed and the training is performed multiple times. The best single training result is selected as the final model to reduce performance fluctuations caused by randomness.

[0028] Optionally, the periodic evaluation of model performance and adjustment of the training strategy based on the evaluation results includes:

[0029] After each training phase, the model performance is evaluated using a development evaluation set, and the training strategy is adjusted based on the evaluation results.

[0030] Use an unseen evaluation set to evaluate the model's generalization ability to ensure that the model performs stably on unseen data.

[0031] According to another aspect of the present invention, a large-scale model capability enhancement device is provided, comprising:

[0032] The first construction unit is used to construct an evaluation system, which includes a developed evaluation set and an unseen evaluation set, for evaluating the model's performance on different core skills.

[0033] The second construction unit is used to construct instruction data, including screening samples from public datasets, synthesizing new data, and performing decontamination processing to ensure that the data covers the target capabilities without leaking the evaluation set content.

[0034] An optimization unit is used for supervised fine-tuning, optimizing the core skills of the model based on the instruction data. The core skills include knowledge retrieval, reasoning, mathematics, code generation, and instruction following.

[0035] The evaluation unit is used to periodically evaluate model performance and adjust training strategies based on the evaluation results.

[0036] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0037] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the large model capability enhancement method according to any embodiment of the present invention.

[0038] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the large model capability enhancement method described in any embodiment of the present invention.

[0039] This invention provides a method, apparatus, device, and medium for enhancing the capabilities of large-scale language models. The method includes constructing an evaluation system comprising a developed evaluation set and an unseen evaluation set to assess the model's performance on different core skills; constructing instruction data, including selecting samples from public datasets, synthesizing new data, and performing decontamination processing to ensure data coverage of target capabilities without revealing the evaluation set content; performing supervised fine-tuning to optimize the model's core skills based on the instruction data, including knowledge retrieval, reasoning, mathematics, code generation, and instruction following; and periodically evaluating model performance and adjusting training strategies based on evaluation results. This invention, through a systematic evaluation system, high-quality instruction data construction, and supervised fine-tuning strategies, significantly improves the performance of large-scale language models on core skills such as knowledge retrieval, reasoning, mathematics, code generation, and instruction following, thereby enhancing the general and target capabilities of large-scale models.

[0040] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart of a method for enhancing the capabilities of a large model according to an embodiment of the present invention;

[0043] Figure 2 This is a schematic diagram of the structure of a large-scale model capability enhancement device provided in an embodiment of the present invention;

[0044] Figure 3 This is a schematic diagram of the structure of an electronic device that implements the large-model capability enhancement method of the present invention. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0046] like Figure 1 As shown, this embodiment of the invention provides a method for enhancing the capabilities of a large model, which may include the following steps:

[0047] S110. Construct an evaluation system, which includes developing an evaluation set and an unseen evaluation set, to assess the model’s performance on different core skills.

[0048] By constructing a multi-layered evaluation framework that includes both development and unseen evaluation sets, the model's performance across various core skills, such as knowledge retrieval, reasoning, mathematical reasoning, code generation, instruction following, and security, is comprehensively assessed. The development evaluation set is used for model evaluation and adjustment during training, helping the model continuously improve its performance. The unseen evaluation set, on the other hand, evaluates the model's generalization ability after training, i.e., its performance on unseen data, preventing the model from over-relying on training data and ensuring its adaptability to new tasks and scenarios.

[0049] S120. Construct instruction data, including selecting samples from public datasets, synthesizing new data, and performing decontamination processing to ensure that the data covers the target capabilities without leaking the evaluation set content.

[0050] To enhance the model's target capabilities, high-quality instruction data is constructed from multiple perspectives. Filtering high-quality data from publicly available datasets provides abundant training material; synthesizing instruction data specific to certain skills allows for targeted improvement of the model's performance on those skills. Decontamination is a crucial step; an n-gram matching algorithm is used to detect and remove samples overlapping with the evaluation set, preventing the model from "memorizing" evaluation set content during training. Simultaneously, manual data review ensures the data meets the target capability requirements, improving data reliability. Standardizing the data format and constructing a coherent contextual history for multi-turn dialogues helps the model better understand and process instruction data, improving training effectiveness.

[0051] S130. Supervised fine-tuning is the core skill for optimizing the model based on instruction data. The core skills include knowledge retrieval, reasoning, mathematics, code generation, and instruction following.

[0052] Based on the constructed instruction data, a supervised fine-tuning strategy is employed to optimize the model's core skills. In the first stage, high-quality samples are selected for supervised fine-tuning to optimize the training data distribution, ensuring each core skill is adequately trained. Simultaneously, the model's performance is periodically evaluated using a development evaluation set, dynamically adjusting the data distribution. In the second stage, based on the performance on the development evaluation set, the data mixing ratio and training parameters are further optimized, with targeted optimizations performed for specific skills such as mathematics and code generation, significantly improving the model's professional capabilities and meeting practical application needs.

[0053] S140. Regularly evaluate model performance and adjust training strategies based on evaluation results.

[0054] After each training phase, the model's performance is evaluated using a development evaluation set. By analyzing the evaluation results, the model's strengths and weaknesses in different core skills can be identified in a timely manner, allowing for targeted adjustments to the training strategy, such as changing data distribution or optimizing training parameters. Finally, the model's generalization ability is evaluated using an unseen evaluation set to ensure stable performance on new tasks, improve the model's reliability and effectiveness in practical applications, and achieve a comprehensive improvement in both the model's general capabilities and its ability to achieve specific objectives.

[0055] In this embodiment of the invention, the evaluation system is constructed, including:

[0056] Define core skills and identify the core skills that need to be improved. Core skills include knowledge retrieval, reasoning, mathematics, code generation, instruction following, and security.

[0057] Select benchmark tests, choosing a set of benchmark test tasks for each core skill, where:

[0058] The benchmark task is divided into a development evaluation set and an unseen evaluation set. The development evaluation set is used to guide the training process, while the unseen evaluation set is used to evaluate the generalization ability of the model in the final stage.

[0059] In model training and evaluation, the first step is to identify the core skills that need to be improved. Core skills represent the model's performance capabilities in different domains and tasks.

[0060] Knowledge retrieval refers to a model's ability to accurately extract relevant information from its learned knowledge system. For example, when a user asks questions about historical events or scientific knowledge, the model can correctly provide the corresponding knowledge content.

[0061] Reasoning refers to the ability of a model to logically deduce and analyze based on known information to arrive at reasonable conclusions. Solving logic puzzles and making causal inferences based on given conditions are examples of reasoning ability.

[0062] Mathematics encompasses various mathematical operations and problem-solving skills, including arithmetic, algebra, geometry, and more. Examples include solving mathematical word problems and performing numerical calculations.

[0063] Code generation refers to the ability of a model to generate corresponding code based on a given requirement or problem description. This requires the model to understand the syntax and semantics of a programming language and to be able to translate the problem into an effective code implementation.

[0064] Instruction following refers to the model's ability to accurately understand user-issued instructions and complete the corresponding tasks according to those instructions. For example, performing text editing or data processing operations based on instructions.

[0065] Security refers to the model's ability to ensure data security and prevent the generation of harmful or inappropriate content during information processing and user interaction. For example, preventing the model from generating content containing sensitive information, false information, or malicious code.

[0066] For each defined core skill, a suitable set of benchmark test tasks needs to be selected:

[0067] Benchmarking tasks for knowledge recall include MMLU (Massive Multitask Language Understanding) and PopQA. These tests cover a wide range of knowledge domains and evaluate a model's knowledge recall capability by the accuracy of its answers to these questions.

[0068] The benchmark tasks for reasoning are BigBenchHard and DROP, which design various complex reasoning problems to examine the model's ability in logical reasoning, semantic understanding, and other aspects.

[0069] Mathematical benchmark tests, such as GSM8K (Grade School Math 8K) and MATH, contain mathematical problems of varying difficulty levels to assess a model’s mathematical computation and problem-solving abilities.

[0070] Instruction-following benchmark tasks, such as IFEval and AlpacaEval, use a series of instruction tasks to test whether the model can accurately understand and execute user instructions.

[0071] The selected benchmark tasks are divided into a development evaluation set and an unseen evaluation set. The development evaluation set is used to monitor and guide training in real time during model training. At each stage of training, the model is evaluated using the development evaluation set, and training strategies are adjusted based on the evaluation results, such as adjusting the learning rate and optimizing the model structure, to gradually improve the model's performance on core skills. After model training is complete, the unseen evaluation set is used for final evaluation. Since the tasks in the unseen evaluation set have not been encountered by the model during training, they can more realistically reflect the model's generalization ability, that is, the model's performance when faced with unseen tasks. If the model can also achieve good results on the unseen evaluation set, it indicates that the model has strong generalization ability and can handle various different tasks in real-world applications.

[0072] In this embodiment of the invention, constructing instruction data includes:

[0073] Data was sifted from publicly available datasets, including WildChat and Open Assistant; instruction data was synthesized for specific skills, including generating math problems or code generation tasks using templates.

[0074] The n-gram matching algorithm is used to detect and remove samples that overlap with the evaluation set to prevent training data from leaking the evaluation set content; each instruction data is reviewed to ensure that it meets the requirements of the target capability.

[0075] Standardize the data format to ensure that each instruction contains a clear task description and the corresponding correct answer; construct dialogue history for multi-turn dialogue tasks.

[0076] Public datasets such as WildChat Open Assistant contain a wealth of text data, which can serve as the foundation for constructing instruction data. Filtering suitable data from these datasets allows us to leverage abundant existing information, saving on data collection costs and time. For example, WildChat may contain various natural language dialogues, some of which can be processed and converted into instruction data; Open Assistant may provide user-model interaction commands and responses. Filtering data from these provides suitable samples for model training.

[0077] For specific skills, such as mathematical problems or code generation tasks, templates can be used to generate corresponding instruction data. Taking mathematical problems as an example, different types of mathematical problem templates can be designed, such as addition, subtraction, multiplication, and division operations, equation solving, etc., and then a large number of different mathematical problems and their corresponding answers can be generated based on the templates. For code generation tasks, some common programming requirement templates can be defined, such as sorting algorithm implementation, file read / write operations, etc., and then corresponding code examples and task descriptions can be generated.

[0078] To ensure the accuracy and fairness of model evaluation, it is necessary to prevent the training data from revealing the contents of the evaluation set. If the training data contains samples from the evaluation set, the model may memorize the answers to these samples during training, thus obtaining falsely high scores during evaluation and failing to accurately reflect the model's generalization ability. The n-gram matching algorithm is a commonly used text matching method that segments text into a sequence of n consecutive characters or words. By comparing the n-gram sequences in the training data and the evaluation set, it identifies the overlapping portions and removes these overlapping samples from the training data.

[0079] In addition to removing samples that overlap with the evaluation set, each instruction data point also needs to be reviewed, as public datasets and synthetic data may contain errors, incompleteness, or fail to meet the target capability requirements. The review process involves checking whether the task description of the instruction data is clear and unambiguous, whether the answer is correct and reasonable, and whether it truly improves the model's performance on the target capability.

[0080] To facilitate model training and processing, the instruction data needs to be formatted uniformly. Each instruction should include a clear task description and a corresponding correct answer. For example, for a math problem instruction, the task description could be "Calculate the result of 2+3," and the correct answer would be "5." For code generation instruction, the task description could be "Write a Python function to sort a list in ascending order," and the correct answer would be a piece of Python code that performs that function.

[0081] In practical applications, many scenarios involve multi-turn dialogues. To enable models to better handle multi-turn dialogue tasks, it is necessary to construct a dialogue history for each instruction. The dialogue history records previous conversations, allowing the model to understand the current task and context based on this historical information, thereby providing a more accurate response.

[0082] In this embodiment of the invention, supervised fine-tuning includes:

[0083] High-quality samples are selected from the constructed instruction data for supervised fine-tuning to optimize the distribution of training data and ensure balanced coverage of each core skill. During the training process, the model performance is evaluated periodically using the development evaluation set, and the data distribution is dynamically adjusted.

[0084] Based on the performance of the evaluation set, optimize the data mixing ratio and training parameters; and perform targeted optimization for specific skills such as mathematics and code generation.

[0085] High-quality samples are selected from the constructed instruction data for supervised fine-tuning. These high-quality samples are characterized by clear task descriptions and accurate, reasonable answers, effectively guiding the model to learn correct knowledge and skills. Training with these high-quality samples avoids the model learning incorrect or ambiguous information, improving training efficiency and effectiveness.

[0086] To ensure the model performs well across all core skills, the distribution of training data needs to be optimized. This means ensuring a relatively balanced number of samples covering each core skill in the training data, avoiding overtraining in some skills while undertraining in others. For example, if there are too many samples related to knowledge recall and too few samples related to mathematical skills in the training data, the model may perform well in knowledge recall but be weak in mathematical ability.

[0087] During training, the model's performance needs to be evaluated periodically using a development evaluation set. This set reflects the model's mastery of each core skill at the current training stage. Based on the evaluation results, the distribution of the training data is dynamically adjusted. If the model is found to be performing poorly on a particular core skill, the proportion of samples related to that skill in the training data can be increased to strengthen training for that skill.

[0088] Based on the model's performance on the evaluation set, the data mixing ratio is optimized. The data mixing ratio refers to the proportion of instruction data from different sources and of different types during training. By adjusting this ratio, the most suitable combination of training data for model learning can be found. Simultaneously, training parameters such as the learning rate and batch size also need to be optimized. The learning rate determines the step size of parameter updates during model training; a suitable learning rate allows the model to converge to the optimal solution more quickly. The batch size affects the efficiency and stability of model training.

[0089] For specific skills such as mathematics and code generation, due to their high level of specialization and complexity, targeted optimization may be necessary. Specialized training tasks and methods can be designed for these skills, relevant training data can be added, or training strategies more suitable for these skills can be adopted. For example, for code generation skills, more code examples and programming standards can be introduced to allow the model to learn better code structures and programming habits; for mathematical skills, more complex mathematical problems and problem-solving approaches can be designed to improve the model's mathematical reasoning and computational abilities.

[0090] In embodiments of the present invention, the method may further include:

[0091] We constructed separate data mixtures for different skills and found the optimal mixture ratio for the performance of a single skill through experiments.

[0092] The optimal data mixing ratios for each skill are combined to form the initial mixed data;

[0093] During the iteration process, contaminated data is removed, and excessively large datasets are downsampled to balance model performance.

[0094] In model training, different skills often require different types of data to improve. Therefore, it is necessary to construct data mixtures for each skill separately. For example, for knowledge retrieval skills, encyclopedic data, Q&A community data, etc., can be combined in different proportions; for code generation skills, open-source code library data, programming competition problem data, etc., can be mixed.

[0095] To ensure the model performs optimally on a single skill, experiments are needed to determine the optimal mixing ratio of different data sources. With other training conditions fixed, multiple training runs and evaluations are performed on different mixing ratios. For example, for the knowledge recall skill, the mixing ratio of encyclopedia data and Q&A community data can be set to 2:8, 5:5, 8:2, etc. Then, an evaluation set is developed to assess the model's knowledge recall ability at each ratio, identifying the optimal mixing ratio for the model's performance.

[0096] Once the optimal data mixing ratio for each skill is determined, these ratios are combined to form the initial mixed data. This ensures that each skill receives appropriate data support during model training, thereby achieving balanced development across multiple skills. Data processing during the iteration process.

[0097] During iterative training, some contaminated data may exist—data that overlaps with the evaluation set or contains errors or noise. This data can interfere with the model's learning, leading to inaccurate evaluation results and reduced generalization ability. Therefore, appropriate methods (such as n-gram matching algorithms) are needed to detect and remove this contaminated data to ensure the quality of the training data.

[0098] Excessive data volume from certain sources can cause models to overemphasize these data points during training, neglecting other data and thus affecting model balance. In such cases, downsampling is necessary. Downsampling involves randomly selecting a portion of data from a large dataset to reduce its size and achieve relative balance with other datasets. For example, if the amount of data in an open-source codebase far exceeds the amount of data in programming competition problems, downsampling of the open-source codebase data can balance the influence of both during training.

[0099] In embodiments of the present invention, the method may further include:

[0100] During training, the random seed is changed and the training is performed multiple times. The best single training result is selected as the final model to reduce performance fluctuations caused by randomness.

[0101] Many operations during model training involve randomness. Taking neural networks as an example, during weight initialization, the connection weights of each neuron are randomly assigned. Different initial weights result in different starting points for the model's training, potentially leading to variations in the final training results. In the data processing stage, the training data is often shuffled to disrupt its order, preventing the model from learning data in a fixed sequence and improving learning efficiency. The shuffling process relies on random numbers; different random number sequences will result in different data orders, thus affecting model training. By repeatedly changing the random seed during training, multiple possible training paths can be explored, leading to various different training results.

[0102] If training is performed only once, the final model performance may fluctuate significantly due to initial conditions and randomness during training. A model trained in one instance might perform well on some tasks but poorly on another. Multiple training iterations yield multiple models, each with varying performance. Selecting the best single-training result from these results as the final model can, to some extent, mitigate poor results caused by randomness and increase the probability of obtaining a high-performance model.

[0103] Multiple training sessions are equivalent to testing and optimizing the model under different "environments," allowing the model to find better solutions in various possible training paths, improving the model's robustness, and enabling it to maintain relatively stable performance when facing different input data and tasks.

[0104] In this embodiment of the invention, periodically evaluating model performance and adjusting the training strategy based on the evaluation results includes:

[0105] After each training phase, the model performance is evaluated using a development evaluation set, and the training strategy is adjusted based on the evaluation results.

[0106] Use an unseen evaluation set to evaluate the model's generalization ability to ensure that the model performs stably on unseen data.

[0107] A development evaluation set is used to assess model performance at the end of each training phase. The training process is divided into multiple phases, with the model being adjusted based on the results of the previous phase and new training data at each phase. After a phase of training is completed, the model is tested using the development evaluation set to obtain performance data on core skills such as knowledge retrieval, reasoning, mathematics, code generation, and instruction following. This data directly reflects the model's strengths and weaknesses; for example, low accuracy in mathematical calculation tasks indicates insufficient training in mathematical ability. Based on the evaluation results, training strategies are adjusted, including changing the distribution of training data to increase the proportion of mathematically relevant training data; and adjusting training parameters such as the learning rate and training epochs to optimize the model towards performance improvement, continuously refining its performance.

[0108] The unseen evaluation set is used after model training is complete; its data did not appear during the training process. Using the unseen evaluation set for evaluation allows us to examine whether the model can apply the knowledge and skills learned during training to entirely new scenarios. If the model performs stably on the unseen evaluation set, with accuracy, recall, and other metrics approaching those of the developed evaluation set, it indicates strong generalization ability and the capacity to handle various types of data in real-world applications. Poor performance suggests overfitting or insufficient learning, requiring adjustments to the training strategy or optimization of the model structure.

[0109] This invention significantly improves the general and target capabilities of large-scale language models through a systematic evaluation system, high-quality instruction data construction, and a multi-stage supervised fine-tuning strategy, and has the following beneficial effects:

[0110] Comprehensive improvement of model performance: By clarifying core skill objectives through a multi-level evaluation framework and optimizing training data and fine-tuning strategies in a targeted manner, the model's performance in key tasks such as knowledge retrieval, reasoning, mathematics, code generation, and instruction following is significantly improved, making it more practical and adaptable in diverse scenarios.

[0111] Enhanced generalization ability: By decontaminating the training data and scientifically dividing it into evaluation sets and unseen evaluation sets, the "memory" problem of the model on the evaluation set content is effectively avoided, ensuring that it performs well on unseen data, thereby greatly improving the model's generalization ability.

[0112] Efficiently utilize data resources: Employ data mixing experiments and optimization strategies to reasonably allocate the proportion of data for different skills, avoid uneven data distribution or redundancy, maximize data utilization, and ensure a balance between the quality and scale of training data through downsampling and decontamination processing.

[0113] Flexible adaptation to specific domain needs: By optimizing specific skills (such as mathematical calculations, code generation, etc.), the model can better meet the practical application needs of professional fields such as medicine, finance, and programming, and provide more accurate and reliable services to industry users.

[0114] Reduce training costs and complexity: Batch aggregation optimization and iterative training strategies reduce reliance on a single training result. By selecting the best model through multiple training runs, performance fluctuations caused by randomness are reduced, while training efficiency and stability are improved.

[0115] This invention promotes the implementation and innovation of technology. It provides comprehensive technical support and guidance for the post-training of large-scale language models, which not only improves the core skills of the models, but also lays the foundation for their wide application in open fields. It has important technical value and market prospects.

[0116] like Figure 2 As shown, this embodiment of the invention provides a large-scale model capability enhancement device, which includes:

[0117] The first construction unit 210 is used to construct an evaluation system, which includes a development evaluation set and an unseen evaluation set, to evaluate the model’s performance on different core skills.

[0118] The second construction unit 220 is used to construct instruction data, including screening samples from public datasets, synthesizing new data, and performing decontamination processing to ensure that the data covers the target capability without leaking the evaluation set content.

[0119] The optimization unit 230 is used for supervised fine-tuning, optimizing the core skills of the model based on instruction data. The core skills include knowledge retrieval, reasoning, mathematics, code generation, and instruction following.

[0120] Evaluation unit 240 is used to periodically evaluate model performance and adjust training strategies based on evaluation results.

[0121] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the capability enhancement device for large models. In other embodiments of the present invention, the capability enhancement device for large models may include more or fewer components than illustrated, or combine some components, or split some components, or arrange different components. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0122] The information interaction and execution process between the various units in the above-mentioned device are based on the same concept as the method embodiment of the present invention, and the specific details can be found in the description of the method embodiment of the present invention, and will not be repeated here.

[0123] Figure 3 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0124] like Figure 3As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0125] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0126] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for enhancing the capabilities of large models.

[0127] In some embodiments, the large model capability enhancement method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the large model capability enhancement method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the large model capability enhancement method by any other suitable means (e.g., by means of firmware).

[0128] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0129] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0130] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0132] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0133] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0134] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for enhancing the capabilities of a large model, characterized in that, include: Construct an evaluation system, which includes an developed evaluation set and an unseen evaluation set, to evaluate the model’s performance on different core skills; Constructing instruction data involves selecting samples from public datasets, synthesizing new data, and performing decontamination processing to ensure that the data covers the target capabilities without revealing the evaluation set content; Supervised fine-tuning is performed to optimize the core skills of the model based on the given instruction data. These core skills include knowledge retrieval, reasoning, mathematics, code generation, and instruction following. Regularly evaluate model performance and adjust training strategies based on the evaluation results.

2. The method according to claim 1, characterized in that, The constructed evaluation system includes: Define core skills and identify the core skills that need to be improved. These core skills include knowledge retrieval, reasoning, mathematics, code generation, instruction following, and security. Select benchmark tests, and choose a set of benchmark test tasks for each of the core skills, wherein: The benchmark task is divided into a development evaluation set and an unseen evaluation set. The development evaluation set is used to guide the training process, and the unseen evaluation set is used to finally evaluate the generalization ability of the model.

3. The method according to claim 1, characterized in that, The construction instruction data includes: Data is sifted from publicly available datasets, including WildChat and Open Assistant; instruction data is synthesized for specific skills, including generating mathematical problems or code generation tasks through templates; The n-gram matching algorithm is used to detect and remove samples that overlap with the evaluation set to prevent training data from leaking the evaluation set content; each instruction data is reviewed to ensure that it meets the requirements of the target capability. A unified data format is used to ensure that each instruction contains a clear task description and a corresponding correct answer; a dialogue history is constructed for multi-turn dialogue tasks.

4. The method according to claim 1, characterized in that, The aforementioned monitoring and fine-tuning includes: High-quality samples are selected from the constructed instruction data for supervised fine-tuning to optimize the distribution of training data and ensure balanced coverage of each core skill. During the training process, the model performance is evaluated periodically using the development evaluation set, and the data distribution is dynamically adjusted. Based on the performance of the evaluation set, optimize the data mixing ratio and training parameters; and perform targeted optimization for specific skills such as mathematics and code generation.

5. The method according to claim 1, characterized in that, The method also includes: We constructed separate data mixtures for different skills and found the optimal mixture ratio for the performance of a single skill through experiments. The optimal data mixing ratios for each of the aforementioned skills are combined to form the initial mixed data; During the iteration process, contaminated data is removed, and excessively large datasets are downsampled to balance model performance.

6. The method according to claim 1, characterized in that, The method also includes: During training, the random seed is changed and the training is performed multiple times. The best single training result is selected as the final model to reduce performance fluctuations caused by randomness.

7. The method according to claim 1, characterized in that, The periodic evaluation of model performance and adjustment of training strategies based on the evaluation results include: After each training phase, the model performance is evaluated using a development evaluation set, and the training strategy is adjusted based on the evaluation results. Use an unseen evaluation set to evaluate the model's generalization ability to ensure that the model performs stably on unseen data.

8. A capability enhancement device for a large model, characterized in that, include: The first construction unit is used to construct an evaluation system, which includes a developed evaluation set and an unseen evaluation set, for evaluating the model's performance on different core skills. The second construction unit is used to construct instruction data, including screening samples from public datasets, synthesizing new data, and performing decontamination processing to ensure that the data covers the target capabilities without leaking the evaluation set content. An optimization unit is used for supervised fine-tuning, optimizing the core skills of the model based on the instruction data. The core skills include knowledge retrieval, reasoning, mathematics, code generation, and instruction following. The evaluation unit is used to periodically evaluate model performance and adjust training strategies based on the evaluation results.

9. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the capability enhancement method for the large model according to any one of claims 1-7.

10. A computer-readable medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to execute the capability enhancement method for the large model as described in any one of claims 1-7.