Instruction data matching experiment method, system and equipment and storage medium
By applying genetic algorithms to automatically optimize data set proportions in large model fine-tuning, the problem of inefficient manual proportions is solved, efficient and automatic data proportions are achieved, and model adaptability and application scope are improved.
Patent Information
- Application Number
- CN202510291789.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
During the fine-tuning of large models, the ratio of instruction data is manually set, which is inefficient and difficult to accurately capture the optimal ratio strategy. Especially when facing large-scale data sets, the experimental complexity and workload have increased significantly.
Genetic algorithms are used to automatically and efficiently determine the optimal proportioning scheme for each data set during fine-tuning of the large model. Through the proportion ratio of the gene-encoded data set, the initial population is randomly generated, and the fitness evaluation, selection, crossover and mutation operations are performed, and the population is iteratively updated until the termination condition is reached, and the optimal proportioning scheme is output.
It improves the efficiency of fine-tuning of large models, reduces the number and time of manual repeated experiments, enhances model adaptability, broadens the scope of application of large models, reduces labor costs, and has good economic and social benefits.
Smart Images

Figure CN120218166A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to an instruction data ratio experiment method, system, device, and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, large models are increasingly widely used in fields such as natural language processing. As a key means to improve model performance, the large model fine-tuning technology has received extensive attention. Large model fine-tuning refers to the process of further training a model on the basis of a pre-trained large model using new, specific task-related data sets. The main purpose of this fine-tuning technology is to enable the model to adapt to new, specific tasks or fields without having to train an entirely new model from scratch. Usually, a data set corresponds to one or more tasks and fields that the model can handle. In order for a large model to complete multiple tasks, generally multiple data sets need to jointly participate in the fine-tuning of the large model, and then it is necessary to determine the ratio of each data set, that is, the instruction data ratio.
[0003] However, in actual operation, the current instruction data ratio mainly depends on the ratio strategy set in advance by humans and conducts large-scale ratio experiments according to human experience. Developers need to initially set the ratio of different data sources according to their own experience, and then find the best ratio scheme through repeated experiments. Due to the limitation of subjective experience in the manually set ratio strategy, it is difficult to accurately capture the optimal ratio strategy. In addition, this method that relies on repeated human experiments not only takes a long time and has low efficiency, but also when facing large-scale data sets, the complexity and workload of the experiments will increase exponentially. Therefore, how to improve the accuracy and efficiency of large model fine-tuning data ratio has become a technical problem to be solved urgently.
[0004] Based on this, the present invention provides an instruction data ratio experiment method, system, device, and storage medium. Summary of the Invention
[0005] The present invention provides an instruction data ratio experiment method, system, device, and storage medium to at least solve the problem that the ratio of instruction data in the large model fine-tuning process in the prior art depends on manual setting and has low efficiency.
[0006] In a first aspect, an embodiment of the present application provides an instruction data ratio experiment method, and the method includes: S1: Determine each data set participating in the large model fine-tuning, as well as the task and field scope corresponding to each data set, and set the parameters of the genetic algorithm; S2: Use the ratio as a gene to encode each data set and randomly generate an initial population; S3: For each individual in the initial population, according to the data ratio scheme it represents, use the corresponding proportion of the dataset for the fine-tuning training of the large model; S4: Evaluate the large model after fine-tuning training using predefined evaluation metrics to obtain the fitness value corresponding to each individual; S5: According to the fitness value corresponding to each individual, use the selection operation to select the individual with the highest fitness value from the current population to generate the next generation population; pair the selected individuals, and then perform the crossover operation with a set crossover probability; for the newly generated individuals, perform gene mutation with a set mutation probability to generate a new generation population; S6: Replace the current population with the new generation population generated through genetic operations, and repeat steps S3 to S5 until the set termination condition is reached, output the individual with the highest fitness value in the new generation population to obtain the optimal ratio scheme.
[0007] Further, in step S2, the parameters of the genetic algorithm include population size, crossover probability, mutation probability, and maximum number of iterations.
[0008] Further, in step S2, the sum of the ratio proportions of all datasets is equal to 1.
[0009] Further, there are three datasets, namely dataset A, dataset B, and dataset C. The gene encoding of an individual is represented as [pA, pB, pC], where pA is the ratio proportion of dataset A, pB is the ratio proportion of dataset B, and pC is the ratio proportion of dataset C, satisfying pA + pB + pC = 1.
[0010] Further, in step S4, the predefined evaluation metrics include at least one of the following metrics: Training set Loss, accuracy, recall rate, F1 value.
[0011] Further, the selection operation adopts the roulette wheel selection, tournament selection, or ranking selection method; The crossover operation adopts the single-point crossover, multi-point crossover, or uniform crossover method.
[0012] Further, the set termination condition is the maximum number of iterations or the fitness value of the individuals in the population tends to be stable.
[0013] In a second aspect, an instruction data ratio experiment system applied to the instruction data ratio experiment method described in the above aspects is also provided in an embodiment of the present application. The system includes: A dataset determination module, configured to determine each dataset participating in the fine-tuning of the large model, as well as the task and domain scope corresponding to each dataset, and set the parameters of the genetic algorithm; An initial population generation module, which is used to perform gene encoding on each data set using the mixing ratio as a gene and randomly generate an initial population; A model training module, which is used to, for each individual in the initial population, according to the data mixing scheme represented by it, use the corresponding proportion of the data set for fine-tuning training of the large model; A model evaluation module, which is used to evaluate the large model after fine-tuning training using predefined evaluation metrics to obtain the fitness value corresponding to each individual; A new generation selection module, which is used to, according to the fitness value corresponding to each individual, select the individual with the highest fitness value from the current population through a selection operation for generating the next generation population; pair the selected individuals, and then perform a crossover operation with a set crossover probability; for the newly generated individuals, perform gene mutation with a set mutation probability to generate a new generation population; An optimal ratio output module, which is used to replace the current population with the new generation population generated through genetic operations, repeatedly execute the model training module, the model evaluation module, and the new generation selection module until a set termination condition is reached, output the individual with the highest fitness value in the new generation population, and obtain the optimal mixing ratio scheme.
[0014] In a third aspect, an electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the instruction data mixing experiment method described in the above aspects are implemented.
[0015] In a fourth aspect, a storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the instruction data mixing experiment method described in the above aspects are implemented.
[0016] As can be seen from the above technical solutions, the present invention has the following advantages: In the instruction data mixing experiment method provided in this application, the number of times and time of manual repeated experiments are reduced. Especially when facing large-scale data sets, the efficient search ability of the genetic algorithm can greatly reduce the experimental complexity and workload and speed up the process of large model fine-tuning; the model adaptability is enhanced: by optimizing the data mixing, the large model can better adapt to the needs of different industries and different tasks, broaden the application scope of the large model, and provide strong support for the intelligent transformation of various industries; the labor cost is reduced: the dependence on the experience of professionals is reduced, and the waste of resources and time cost caused by improper manual setting are reduced, having good economic and social benefits. Description of the Drawings
[0017] To more clearly illustrate the technical solution of the present invention, the accompanying drawings required for the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 It is a flowchart of the experimental method for the instruction data ratio. Detailed implementation manners
[0019] In the following, the experimental method, system, device and storage medium for the instruction data ratio will be described in detail, and various embodiments of the present disclosure will be described more comprehensively. The present disclosure can have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0020] In the following, the term "comprise" or "may comprise" that can be used in various embodiments of the present disclosure indicates the presence of the disclosed functions, operations or elements, and does not limit the addition of one or more functions, operations or elements. In addition, as used in various embodiments of the present disclosure, the terms "comprise", "have" and their cognates are only intended to indicate a specific feature, number, step, operation, element, component or combination of the foregoing items, and should not be construed as first excluding the existence or addition of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items.
[0021] In various embodiments of the present disclosure, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the listed words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.
[0022] Expressions (such as "first", "second", etc.) used in various embodiments of the present disclosure may modify various constituent elements in the various embodiments, but do not limit the corresponding constituent elements. For example, the above expressions do not limit the order and / or importance of the elements. The above expressions are only used for the purpose of distinguishing one element from other elements. For example, the first user device and the second user device indicate different user devices, although both are user devices. For example, without departing from the scope of the various embodiments of the present disclosure, the first element may be referred to as the second element, and similarly, the second element may also be referred to as the first element.
[0023] It should be noted that: If a description "connects" one component to another component, the first component can be directly connected to the second component, and a third component can be "connected" between the first component and the second component. Conversely, when a component is "directly connected" to another component, it can be understood that there is no third component between the first component and the second component.
[0024] The term "user" used in various embodiments of the present disclosure may refer to a person using an electronic device, which may be a monitoring person, or a tester, or an operator.
[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] The embodiment of the present application provides an instruction data ratio experiment method, which solves the technical problems that the ratio of instruction data in the fine-tuning process of the existing large model depends on manual setting and is inefficient. To solve the above problems, the present invention introduces a genetic algorithm, an intelligent optimization algorithm, to automatically and efficiently determine the optimal ratio scheme of each data set during the fine-tuning of the large model, thereby improving the adaptability and performance of the large model in a specific task or field, reducing labor costs and experimental complexity, and providing a new, scientific and effective technical means for the application and optimization of the large model.
[0027] Next, the technical solutions proposed in the embodiments of the present application will be described in detail through the accompanying drawings.
[0028] Figure 1 It is a flowchart of an instruction data ratio experiment method provided by an embodiment of the present application. As Figure 1 shown, an instruction data ratio experiment method provided by an embodiment of the present application includes six stages: initialization, generating an initial population, fitness evaluation, genetic operations, iterative update, and outputting the optimal ratio scheme, and specifically includes the following steps: Initialization: S1: Determine each data set participating in the fine-tuning of the large model, as well as the corresponding tasks and domain scopes of each data set, and set the parameters of the genetic algorithm.
[0029] The parameters of the genetic algorithm include population size, crossover probability, mutation probability, and maximum number of iterations. The population size determines the number of individuals in each generation of the population. The crossover probability and mutation probability are used to control the intensity of genetic operations. The maximum number of iterations limits the running time of the algorithm to prevent infinite loops.
[0030] Generate the initial population: S2: Use the mixing ratio as genes to encode each dataset. According to the set population size, generate a sufficient number of initial individuals and randomly generate the initial population. It should be noted that each individual in the initial population represents a data mixing scheme, and the way of its gene encoding is the mixing ratio of each dataset during the fine-tuning process. The sum of the mixing ratios of all datasets is equal to 1.
[0031] Fitness evaluation: S3: For each individual in the initial population, according to the data mixing scheme it represents, use the corresponding proportion of the dataset for the fine-tuning training of the large model.
[0032] S4: After the fine-tuning training is completed, use predefined evaluation metrics to evaluate the performance of the fine-tuned large model in a specific task or domain range, and obtain the fitness value corresponding to each individual. The higher the fitness value, the better the performance of the large model under this data mixing scheme.
[0033] Genetic operations: S5: Selection operation: According to the fitness value corresponding to each individual, use the selection operation to select the individual with the highest fitness value from the current population to generate the next generation population. Among them, individuals with higher fitness have a higher probability of being selected, so as to pass excellent genes to the next generation. Crossover operation: Pair the selected individuals and then perform the crossover operation with a set crossover probability. Through the crossover operation, exchange some gene segments of the two individuals to generate new individuals and increase the diversity of the population. Mutation operation: For the newly generated individuals, perform gene mutation with a set mutation probability. Randomly change one or more gene values in the individual gene encoding to further expand the search range of the population and avoid the algorithm falling into a local optimum. Normalization operation: To ensure that the individuals have physical meanings (data mixing ratios), normalize the new individuals obtained from the above operations to generate a new generation population.
[0034] Iterative update: S6: Replace the current population with the new generation population generated by the genetic operations, and repeat steps S3 to S5 to perform fitness evaluation and genetic operations on the new generation population until the set maximum number of iterations is reached or the fitness values of the individuals in the population tend to be stable and no longer increase significantly.
[0035] Output the optimal mixing scheme: During the iteration process, record the individual with the highest fitness and its corresponding fitness value in each generation of the population. When the algorithm terminates, select the individual with the highest fitness from the recorded individuals and output the data mixing scheme represented by its gene encoding as the final optimal mixing scheme.
[0036] The above-mentioned instruction data ratio experiment method improves the experimental efficiency, reduces the number of times and time of manual repeated experiments. Especially when facing large-scale data sets, the efficient search ability of the genetic algorithm can greatly reduce the experimental complexity and workload, and speed up the process of large model fine-tuning; enhances the model adaptability: by optimizing the data ratio, the large model can better adapt to the needs of different industries and different tasks, broadens the application scope of the large model, and provides strong support for the intelligent transformation of various industries; reduces the labor cost: reduces the dependence on the experience of professional personnel, reduces the waste of resources and time cost caused by improper manual setting, and has good economic and social benefits.
[0037] In an exemplary embodiment, there are three data sets, namely data set A, data set B, and data set C. The gene encoding of an individual is represented as [pA, pB, pC], where pA is the ratio of data set A, pB is the ratio of data set B, and pC is the ratio of data set C, satisfying pA + pB + pC = 1.
[0038] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another instruction data ratio experiment method is provided. In step S4, the predefined evaluation metrics include at least one of the following metrics: training set Loss, accuracy, recall rate, F1 value.
[0039] Based on the above embodiment, in order to further improve the instruction data ratio experiment method provided in the above embodiment, as an implementable manner, in an embodiment, the selection operation adopts roulette wheel selection, tournament selection or ranking selection method; the crossover operation adopts single-point crossover, multi-point crossover or uniform crossover method.
[0040] Exemplarily, determine the data sets and their task fields participating in fine-tuning, and set the genetic algorithm parameters. The data sets are A, B, and C respectively, and the population size is set to , the crossover probability is set to , the mutation probability is set to , the maximum number of iterations is set to .
[0041] Randomly generate an initial population. The gene encoding of each individual is [pA, pB, pC], satisfying pA + pB + pC = 1. Among them, the gene encoding of individual 1 is: [0.4, 0.3, 0.3], the gene encoding of individual 2 is: [0.2, 0.5, 0.3], and the gene encoding of individual 3 is: [0.6, 0.2, 0.2].
[0042] For each individual, fine-tuning training is carried out according to its data ratio plan, and the fitness value is calculated using a predefined evaluation metric. Assume the evaluation metric is accuracy.
[0043] For example, the data ratio plan of individual 1 is [0.4, 0.3, 0.3]: Use 40% of dataset A, 30% of dataset B, and 30% of dataset C for fine-tuning.
[0044] If the accuracy of the fine-tuned model on the validation set is 85%, then the fitness value of individual 1 is 0.85.
[0045] Select excellent individuals according to the fitness value. Using the roulette wheel selection method, calculate the selection probability of each individual:
[0046] where, represents the selection probability of the th individual, represents the fitness value of the th individual, represents the fitness value of the th individual, .
[0047] Assume the fitness values of individual 1, individual 2, and individual 3 in the population are: 0.85, 0.8, 0.9 respectively. Then the selection probabilities are: Individual 1:
[0048] Individual 2:
[0049] Individual 3:
[0050] Randomly select individuals according to the selection probability to generate the next generation population.
[0051] Pair the selected individuals and perform crossover operation with a set crossover probability Assume single-point crossover is used and the crossover point is randomly selected. For example: Individual 1 and individual 2 are paired, and the crossover point is 1: Before crossover: Individual 1 [0.4, 0.3, 0.3], Individual 2 [0.2, 0.5, 0.3]; After crossover: Individual 1 [0.2, 0.3, 0.3], Individual 2 [0.4, 0.5, 0.3]; For the newly generated individuals, perform gene mutation with a set mutation probability Assume the mutation point is randomly selected and the mutation value is randomly generated. For example: The second gene of individual 1 [0.2, 0.3, 0.3] mutates, and the new value is 0.4: After mutation, individual 1 [0.2, 0.4, 0.4].
[0052] Replace the current population with the new generation of population generated through genetic operations, and repeat steps 2 to 5 until the set termination conditions are reached. For example, the maximum number of iterations is 100, or the fitness values of individuals in the population tend to be stable.
[0053] Record the individual with the highest fitness in each generation of the population and its corresponding fitness value. When the algorithm terminates, select the individual with the highest fitness as the optimal ratio plan. Suppose the final optimal individual is [0.3, 0.4, 0.3] and the fitness value is 0.92, then the optimal ratio plan is: Dataset A: 30%, Dataset B: 40%, Dataset C: 30%.
[0054] The present invention includes six stages: initialization, generation of the initial population, fitness evaluation, genetic operations, iterative update, and output of the optimal ratio plan. By determining the datasets involved in fine-tuning and the corresponding task fields, setting the parameters of the genetic algorithm, randomly generating the initial population and performing fitness evaluation, and performing genetic operations such as selection, crossover, and mutation based on the fitness values, iteratively updating the population, and finally outputting the optimal data ratio plan. The present invention can improve the data ratio efficiency, reduce the labor cost, and provide a scientific and effective technical means for the application and optimization of large models.
[0055] The present invention also provides an instruction data ratio experiment system, which includes: A dataset determination module for determining each dataset involved in the fine-tuning of the large model, as well as the tasks and field ranges corresponding to each dataset, and setting the parameters of the genetic algorithm.
[0056] An initial population generation module for using the ratio as genes to encode each dataset, generating a sufficient number of initial individuals according to the set population size, and randomly generating the initial population; A model training module for using the corresponding ratio of the dataset for the fine-tuning training of the large model for each individual in the initial population according to the data ratio plan it represents; it should be noted that each individual in the initial population represents a data ratio plan, and its gene encoding method is the ratio of each dataset in the fine-tuning process.
[0057] A model evaluation module for evaluating the fine-tuned large model using predefined evaluation metrics after the fine-tuning training is completed, and obtaining the fitness value corresponding to each individual; the higher the fitness value, the better the performance of the large model under this data ratio plan.
[0058] A new generation selection module is used to select the individual with the highest fitness value from the current population through a selection operation according to the fitness value corresponding to each individual, for generating the next generation population; pair the selected individuals, and then perform a crossover operation with a set crossover probability; for the newly generated individuals, perform gene mutation with a set mutation probability to generate a new generation population.
[0059] An optimal ratio output module is used to replace the current population with the new generation population generated through genetic operations, repeatedly execute the model training module, the model evaluation module, and the new generation selection module until the set termination condition is reached, output the individual with the highest fitness value in the new generation population, and obtain the optimal ratio scheme.
[0060] The set termination condition is the maximum number of iterations or the fitness value of individuals in the population tends to be stable; during the iteration process, record the individual with the highest fitness value in each generation population and its corresponding fitness value; when the algorithm terminates, select the individual with the highest fitness value from the recorded individuals, and use the data ratio scheme represented by its gene encoding as the final optimal ratio scheme for output.
[0061] In the embodiments of the instruction data ratio experiment system provided in the above embodiments of the present disclosure, the instruction data ratio experiment system and the instruction data ratio experiment methods in the above various embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the instruction data ratio experiment system, reference can be made to the embodiments of the above instruction data ratio experiment methods.
[0062] The instruction data ratio experiment method provided in the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown in the figure, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.
[0063] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0064] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0065] Among them, the processor may be the nerve center and command center of the electronic device. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0066] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated access and reduces the waiting time of the processor, thus improving the system efficiency.
[0067] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.
[0068] The internal memory can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. The internal memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0069] The above-mentioned electronic device implements the determination of each data set participating in the fine-tuning of the large model, as well as the corresponding tasks and domain scopes of each data set, and sets the parameters of the genetic algorithm in the experimental method for the instruction data ratio; uses the ratio as the gene to encode each data set and randomly generates an initial population; for each individual in the initial population, according to the data ratio scheme it represents, uses the corresponding proportion of the data set for the fine-tuning training of the large model; uses predefined evaluation metrics to evaluate the large model after fine-tuning training to obtain the fitness value corresponding to each individual; according to the fitness value corresponding to each individual, selects the individual with the highest fitness value from the current population through selection operation for generating the next generation population; pairs the selected individuals and then performs crossover operation with a set crossover probability; for the newly generated individuals, performs gene mutation with a set mutation probability to generate a new generation population; replaces the current population with the new generation population generated through genetic operations, repeats the above steps until the set termination condition is reached, outputs the individual with the highest fitness value in the new generation population, and obtains the optimal ratio scheme. The experimental method for the instruction data ratio improves the experimental efficiency, reduces the number of times and time of manual repeated experiments. Especially when facing large-scale data sets, the efficient search ability of the genetic algorithm can greatly reduce the experimental complexity and workload and speed up the process of large model fine-tuning; enhances the model adaptability: by optimizing the data ratio, enables the large model to better adapt to the needs of different industries and different tasks, broadens the application scope of the large model, and provides strong support for the intelligent transformation of various industries; reduces the labor cost: reduces the dependence on the experience of professional personnel, reduces the waste of resources and time cost caused by improper manual setting, and has good economic and social benefits.
[0070] In the storage medium provided by this application, there is a program product that can implement the experimental method for the instruction data ratio.
[0071] The experimental method for instruction data ratio includes: determining each dataset involved in the fine-tuning of the large model, as well as the corresponding tasks and domain scopes for each dataset, and setting the parameters of the genetic algorithm; using the ratio as genes to encode each dataset and randomly generating an initial population; for each individual in the initial population, according to the data ratio scheme it represents, using the corresponding proportion of the dataset for the fine-tuning training of the large model; using predefined evaluation metrics to evaluate the large model after fine-tuning training to obtain the fitness value corresponding to each individual; according to the fitness value corresponding to each individual, using the selection operation to select the individual with the highest fitness value from the current population to generate the next generation population; pairing the selected individuals and then performing crossover operations with a set crossover probability; for the newly generated individuals, performing gene mutation with a set mutation probability to generate a new generation population; replacing the current population with the new generation population generated through genetic operations, repeating the above steps until the set termination condition is reached, and outputting the individual with the highest fitness value in the new generation population to obtain the optimal ratio scheme.
[0072] The experimental method for instruction data ratio improves the experimental efficiency: reduces the number of times and time of manual repeated experiments. Especially when facing large-scale datasets, the efficient search ability of the genetic algorithm can greatly reduce the experimental complexity and workload, and speed up the process of large model fine-tuning.
[0073] The experimental method for instruction data ratio enhances the model adaptability: by optimizing the data ratio, enabling the large model to better adapt to the needs of different industries and different tasks, broadening the application scope of the large model, and providing strong support for the intelligent transformation of various industries.
[0074] The experimental method for instruction data ratio reduces the labor cost: reduces the dependence on the experience of professional personnel, reduces the waste of resources and time cost caused by improper manual setting, and has good economic and social benefits.
[0075] In some possible implementation manners, the experimental method for instruction data ratio of the present disclosure can be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to make the terminal device execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.
[0076] The storage medium of the present disclosure may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0077] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0078] For those of ordinary skill in the art, designing different forms of control circuits does not require creative labor according to the teachings of the present invention. These changes, modifications, substitutions, and variations to the embodiments still fall within the protection scope of the present invention without departing from the principles and spirit of the present invention.
Claims
1. A command data matching experimental method, characterized in that: The method comprises: S1: Determine the various data sets involved in fine-tuning the large model, as well as the tasks and domains corresponding to each data set, and set the parameters of the genetic algorithm; S2: Use the matching ratio as the gene to encode each data set and randomly generate the initial population; S3: For each individual in the initial population, according to the data ratio scheme it represents, the corresponding proportion of the data set is used for fine-tuning training of the large model; S4: Use the predefined evaluation indicators to evaluate the large model after fine-tuning training to obtain the fitness value corresponding to each individual; S5: According to the fitness value corresponding to each individual, the individual with the highest fitness value is selected from the current population by selection operation to generate the next generation population; the selected individuals are paired, and then a crossover operation is performed with a set crossover probability; for the newly generated individuals, a gene mutation is performed with a set mutation probability to generate a new generation population; S6: Replace the current population with the new generation population generated by genetic operation, and repeat steps S3 to S5 until the set termination condition is reached, output the individual with the highest fitness value in the new generation population, and obtain the optimal matching scheme.
2. The instruction data matching experimental method according to claim 1, characterized in that: In step S2, the parameters of the genetic algorithm include population size, crossover probability, mutation probability, and maximum number of iterations.
3. The instruction data matching experimental method according to claim 2, characterized in that: In step S2, the sum of the matching ratios of all data sets is equal to 1.
4. The instruction data matching experimental method according to claim 3, characterized in that: There are three data sets, namely data set A, data set B, and data set C. The genetic code of the individuals is expressed as [pA, pB, pC], where pA is the matching ratio of data set A, pB is the matching ratio of data set B, and pC is the matching ratio of data set C, satisfying pA+pB+pC=1.
5. The instruction data matching experimental method according to claim 1, characterized in that: In step S4, the predefined evaluation index includes at least one of the following indexes: Training set Loss, accuracy, recall, and F1 value.
6. The instruction data matching experimental method according to claim 5, characterized in that: The selection operation adopts roulette selection, tournament selection or ranking selection method; The crossover operation adopts single-point crossover, multi-point crossover or uniform crossover.
7. The instruction data matching experimental method according to claim 6, characterized in that: The set termination condition is the maximum number of iterations or the fitness values of individuals in the population tending to be stable.
8. An instruction data matching experiment system applied to the instruction data matching experiment method according to any one of claims 1 to 7, characterized in that: The system comprises: The data set determination module is used to determine the various data sets involved in the fine-tuning of the large model, as well as the tasks and domains corresponding to each data set, and to set the parameters of the genetic algorithm; The initial population generation module is used to use the matching ratio as the gene, encode the genes of each data set, and randomly generate the initial population; The model training module is used to use the corresponding proportion of the data set for fine-tuning the large model for each individual in the initial population according to the data ratio scheme represented by it; The model evaluation module is used to evaluate the large model after fine-tuning training using predefined evaluation indicators to obtain the fitness value corresponding to each individual; The next generation selection module is used to select the individual with the highest fitness value from the current population according to the fitness value corresponding to each individual, and use the selection operation to generate the next generation population; pair the selected individuals, and then perform a crossover operation with a set crossover probability; perform gene mutation on the newly generated individuals with a set mutation probability to generate a new generation population; The optimal matching ratio output module is used to replace the current population with the new generation population generated by genetic operation, repeatedly execute the model training module, model evaluation module and new generation selection module until the set termination condition is reached, output the individual with the highest fitness value in the new generation population, and obtain the optimal matching ratio solution.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the instruction data matching experimental method as described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the instruction data matching experimental method as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Model iteration method and device based on evolutionary algorithm, equipment and storage medium
CN121145997A