Model training method, device and equipment for industry large model and medium

Through automated data ratio generation and hyperparameter optimization processes, the problems of data ratio reliance on experience and low hyperparameter search efficiency in large-scale model training in the industry have been solved, which has improved model training efficiency and resource utilization, and enhanced analysis accuracy and conclusion reliability.

CN120705578APending Publication Date: 2025-09-26SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510825366.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In existing industry large-scale model training solutions, data allocation relies on manual experience, hyperparameter search efficiency is low, resource utilization is low, and the model generalization ability cannot be effectively optimized and computing resources are seriously wasted.

Method used

By creating an experimental space corresponding to the industry's large models, generating data ratios and optimizing hyperparameters, and using the Pipeline automated process to perform multi-combination parallel training, the optimal configuration files and models are screened out.

Benefits of technology

It realizes the automatic generation of data ratios and hyperparameter optimization, improves the efficiency of model training, reduces the waste of computing resources and human resources, and enhances the analysis accuracy and conclusion reliability of the model in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705578A_ABST
    Figure CN120705578A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device for an industry large model, equipment and a medium, and relates to the technical field of computers, and the method comprises the steps: creating a space which corresponds to the industry large model and is used for carrying out a matching experiment, and configuring information; determining a plurality of data matching combinations based on the target label information corresponding to the space; initiating a first search parameter batch task corresponding to the industry large model for various data matching combinations and super parameter range configuration files through a pre-constructed Pipeline in the space so as to screen out a current optimal target super parameter range configuration file and a first trained model; initiating a second search parameter batch task corresponding to the first trained model on the basis of the target super parameter range configuration file and a plurality of data matching combinations, so as to screen out a current optimal target data matching combination and a second trained model; and judging whether a preset Pipeline iteration termination condition is met at present or not. According to the invention, joint optimization of the hyper-parameter and training data matching can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a model training method, device, equipment and medium for large industry models. Background Art

[0002] With the rapid development of artificial intelligence technology, the demand for training and tuning large industry models (such as specialized models in finance, healthcare, and industry) is increasing. However, existing solutions still have significant shortcomings in model parameter tuning and optimizing the ratio of training data:

[0003] (1) Data ratio depends on experience: In traditional methods, the ratio of training data (such as the proportion of data of different label categories) usually relies on manual experience, which lacks scientific basis and easily leads to overfitting or underfitting of the model, affecting the generalization ability.

[0004] (2) Low efficiency of hyperparameter search: The optimization of model hyperparameters (such as learning rate, batch size, etc.) often uses grid search or random search, which has high computational cost and is difficult to cover complex parameter spaces. Especially when the data ratio changes dynamically, the coupling relationship between hyperparameters and data ratio is difficult to effectively analyze.

[0005] (3) Serious waste of resources: Insufficient parallelization capabilities make it impossible to initiate training tasks for multiple sets of data ratios and hyperparameter combinations at the same time, resulting in idle computing resources or repeated consumption. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a model training method, device, equipment and medium for large industry models, which can realize the automatic generation of data ratios and the joint optimization of hyperparameters and training data ratios, thereby improving the efficiency of hyperparameter search and model training, reducing the waste of computer resources and human resources, thereby improving resource utilization, and improving the analysis accuracy and conclusion reliability of large industry models in complex scenarios. The specific scheme is as follows:

[0007] In a first aspect, the present application provides a model training method for a large industry model, comprising:

[0008] Create a space for matching experiments corresponding to the industry model, and trigger information configuration operations on the determined model experiment space to obtain the configured space;

[0009] Performing data matching based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information to determine multiple data matching combinations;

[0010] In the configured space, by starting the pre-built Pipeline, the first parameter search batch task corresponding to the industry large model is initiated for various data ratio combinations and hyperparameter range configuration files, and the currently optimal target hyperparameter range configuration file and the corresponding first trained model are screened based on the corresponding first parameter search task results;

[0011] Initiate a second parameter search batch task corresponding to the first trained model based on the target hyperparameter range configuration file and the multiple data ratio combinations, and screen out the current optimal target data ratio combination and the corresponding second trained model based on the corresponding second parameter search task results;

[0012] Based on the current target data ratio combination and the second trained model, determine whether the preset Pipeline iteration termination condition is currently met, and if so, determine that the model parameter adjustment operation and training data ratio tuning operation corresponding to the industry large model have been completed.

[0013] Optionally, the method further includes:

[0014] If the preset Pipeline iteration termination condition is not met at present, the process jumps back to the step of initiating the first parameter search batch task corresponding to the industry large model for various data ratio combinations and hyper-parameter range configuration files.

[0015] Optionally, creating a space for performing a matching experiment corresponding to the industry macro model and triggering an information configuration operation for the determined model experiment space includes:

[0016] Create a space for proportioning experiments corresponding to the industry model to obtain a model experiment space;

[0017] The experiment name, label information, experiment space information display list, administrator information and accessible list corresponding to the model experiment space are configured to determine the configured space.

[0018] Optionally, the generating of data ratios based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information to determine multiple data ratio combinations includes:

[0019] Determining labeled training data sets corresponding to various target label information based on the multiple target label information corresponding to the configured space, and performing random balanced sampling on each labeled training data set to obtain balanced sampled data sets with consistent sample numbers;

[0020] Configuring the upper and lower limits and the step length of the ratio corresponding to each target tag information to determine the ratio range information corresponding to each target tag information;

[0021] Obtaining a preset ratio of the number of samples corresponding to each balanced sampling data set in the total training data set;

[0022] Data ratio generation is performed based on the ratio range information corresponding to the various target label information and the balanced sampled data sets, as well as the preset quantity proportions corresponding to each balanced sampled data set, to determine several data ratio combinations corresponding to multiple label combinations.

[0023] Optionally, the first parameter search batch task corresponding to the industry large model is initiated for various data ratio combinations and super-parameter range configuration files, and the currently optimal target super-parameter range configuration file and the corresponding first trained model are screened out based on the corresponding first parameter search task results, including:

[0024] For any of the data ratio combinations, based on the current data ratio combination and different hyper-parameter range configuration files, initiate the parameter search and batch running subtasks corresponding to the industry large model;

[0025] In the process of running each of the parameter search and batch running subtasks of the industry large model, the corresponding balanced sampling data set is obtained based on the current data ratio combination, and data is randomly extracted from the balanced sampling data set based on the current data ratio combination to determine the first training set;

[0026] Based on the first training set and the super-parameter range configuration files corresponding to each of the parameter search and batch running subtasks of the industry large model, the industry large model is trained to determine a corresponding trained model;

[0027] Performing model prediction on the trained model, and performing evaluation based on the corresponding first prediction result and the first target evaluation data set corresponding to each of the parameter search and batch running subtasks of the industry large model, to determine the first parameter search subtask results corresponding to each of the parameter search and batch running subtasks of the industry large model;

[0028] Determine a first parameter search task result corresponding to the current data matching combination based on each of the first parameter search subtask results;

[0029] Based on the first preset screening rule and the first parameter search task results corresponding to the various data ratio combinations, the current optimal target super-parameter range configuration file and the first trained model corresponding to the target super-parameter range configuration file are screened out.

[0030] Optionally, the initiating a second parameter search batch running task corresponding to the first trained model based on the target hyperparameter range configuration file and the multiple data ratio combinations, and screening out the current optimal target data ratio combination and the corresponding second trained model based on the corresponding second parameter search task results, includes:

[0031] Based on the target hyperparameter range configuration file, and in combination with various data ratio combinations, initiate parameter search and batch running subtasks corresponding to the first trained model;

[0032] During the execution of any of the parameter search and batch running subtasks of the first trained model, determining a corresponding second training set based on the data ratio combination corresponding to the current parameter search and batch running subtask of the first trained model and the corresponding balanced sampling data set;

[0033] Performing model training on the first trained model based on the second training set and the target hyperparameter range configuration file to determine the trained first trained model;

[0034] Performing model prediction on the first trained model after training, and evaluating based on a second prediction result and a second target evaluation data set corresponding to the current parameter search batch subtask of the first trained model to determine a corresponding second parameter search subtask result;

[0035] Determine a second parameter search task result corresponding to the target super-parameter range configuration file based on each of the second parameter search subtask results;

[0036] Based on the second preset screening rule and the second parameter search task result, the current optimal target data ratio combination and the second trained model corresponding to the target data ratio combination are screened out.

[0037] Optionally, the determining whether a preset Pipeline iteration termination condition is currently satisfied based on the current target data ratio combination and the second trained model includes:

[0038] Determine whether the current target data ratio combination is consistent with the data ratio combination corresponding to the target hyperparameter range configuration file, or determine whether the current second trained model converges to determine whether the preset Pipeline iteration termination condition is currently met.

[0039] In a second aspect, the present application provides a model training device for a large industry model, comprising:

[0040] The experimental space creation module is used to create a space corresponding to the industry large model for conducting ratio experiments, and trigger information configuration operations on the determined model experimental space to obtain the configured space;

[0041] a data ratio generation module, configured to generate data ratios based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information, so as to determine multiple data ratio combinations;

[0042] A first parameter search module is configured to initiate a first parameter search batch task corresponding to the industry large model for various data ratio combinations and hyperparameter range configuration files within the configured space by starting a pre-built pipeline, and to screen out the currently optimal target hyperparameter range configuration file and the corresponding first trained model based on the corresponding first parameter search task results;

[0043] A second parameter search module is used to initiate a second parameter search batch task corresponding to the first trained model based on the target super-parameter range configuration file and the multiple data ratio combinations, and screen out the current optimal target data ratio combination and the corresponding second trained model based on the corresponding second parameter search task results;

[0044] The iteration termination module is used to determine whether the preset Pipeline iteration termination condition is currently met based on the current target data ratio combination and the second trained model, and if so, determine whether the model parameter adjustment operation and training data ratio tuning operation corresponding to the industry large model have been completed.

[0045] In a third aspect, the present application provides an electronic device, comprising:

[0046] Memory, used to store computer programs;

[0047] A processor is used to execute the computer program to implement the steps of the aforementioned model training method for large industry models.

[0048] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the aforementioned model training method for large industry models.

[0049] It can be seen that in this application, a space for performing a matching experiment corresponding to the industry large model is created, and an information configuration operation is triggered on the determined model experiment space to obtain a configured space;

[0050] Data ratio generation is performed based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information to determine a plurality of data ratio combinations; within the configured space, by starting a pre-built Pipeline, a first parameter search batch task corresponding to the industry large model is initiated for various data ratio combinations and super-parameter range configuration files, and the currently optimal target super-parameter range configuration file and the corresponding first trained model are screened out based on the corresponding first parameter search task result; a second parameter search batch task corresponding to the first trained model is initiated based on the target super-parameter range configuration file and the multiple data ratio combinations, and the currently optimal target data ratio combination and the corresponding second trained model are screened out based on the corresponding second parameter search task result; based on the current target data ratio combination and the second trained model, it is determined whether the preset Pipeline iteration termination condition is currently met, and if so, it is determined that the model parameter adjustment operation and the training data ratio tuning operation corresponding to the industry large model have been completed. That is to say, in this application, a space corresponding to the industry large model for matching experiments is first created and configured, and then the data matching is generated for the various target label information and the corresponding labeled training data set corresponding to the space. After that, the pre-built Pipeline is started in the space, and the first parameter search batch task corresponding to the industry large model is initiated for the various generated data matching combinations and super-parameter range configuration files, and the optimal target super-parameter range configuration file and the first trained model are screened out based on the corresponding first parameter search task results. After that, the second parameter search batch task corresponding to the first trained model is initiated based on the target super-parameter range configuration file and the various data matching combinations, and the current optimal target data matching combination and the second trained model are screened out. When the preset Pipeline iteration termination condition is met, the model parameter adjustment operation and the training data matching tuning operation corresponding to the industry large model are determined to be completed. In this way, the automatic generation of data matching and the joint optimization of hyperparameters and training data matching can be realized, thereby improving the efficiency of hyperparameter search and model training, reducing the waste of computer resources and human resources, and thus improving resource utilization, and improving the analysis accuracy and conclusion reliability of the industry large model in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0052] Figure 1A flow chart of a model training method for a large industry model provided in this application;

[0053] Figure 2 A schematic diagram of a specific ratio generation process provided in this application;

[0054] Figure 3 A specific Pipeline iterative process diagram provided for this application;

[0055] Figure 4 A schematic diagram of the structure of a model training device for large industry models provided in this application;

[0056] Figure 5 This is a structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0058] Existing solutions still have significant deficiencies in terms of model parameter adjustment and training data ratio optimization: (1) Data ratio relies on experience: In traditional methods, the ratio of training data (such as the ratio of data of different label categories) is usually set based on manual experience, which lacks scientific basis and easily leads to model overfitting or underfitting, affecting generalization ability. (2) Inefficient hyperparameter search: The optimization of model hyperparameters (such as learning rate, batch size, etc.) often uses grid search or random search, which has high computational cost and is difficult to cover complex parameter space. Especially when the data ratio changes dynamically, the coupling relationship between hyperparameters and data ratio is difficult to effectively analyze. (3) Serious waste of resources: Insufficient parallelization capability makes it impossible to initiate training tasks for multiple sets of data ratios and hyperparameter combinations at the same time, resulting in idle computing resources or repeated consumption.

[0059] To this end, this application provides a model training solution for large industry models, which can realize the automatic generation of data ratios and the joint optimization of hyperparameters and training data ratios, thereby improving the efficiency of hyperparameter search and model training.

[0060] See also Figure 1 As shown, an embodiment of the present invention discloses a model training method for a large industry model, comprising:

[0061] Step S11: Create a space for performing a matching experiment corresponding to the industry large model, and trigger an information configuration operation on the determined model experiment space to obtain a configured space.

[0062] Specifically, in this embodiment, it is first necessary to create a space for matching experiments corresponding to the industry big model and configure the relevant information of the space, where the industry big model is specifically a big model related to the financial, medical or industrial industries. That is, create a space for matching experiments corresponding to the industry big model to obtain a model experiment space; configure the experiment name, label information, experiment space information display list, administrator information and accessible list corresponding to the model experiment space to determine the configured space. Specifically, create a space for matching experiments. Within the scope of this space, you can do the following things. At the same time, this space can restrict the permissions of accessible personnel. A model experiment includes the following information:

[0063] (1) Experiment name: This is a required field. Placehold: Enter the experiment name. Input rules: 2-30 characters, can only start with a letter or Chinese character, and can only contain letters, Chinese characters, numbers, ".", "_" or "-";

[0064] (2) Tag classification version number: This is a required option. Select the tag classification version number from the drop-down menu.

[0065] (3) Chinese name of label category: This is a required option. Select a label category. After the user enters a label, the system recommends labels that meet the conditions. Multiple labels can be selected (the corresponding threshold can be set independently) or a drop-down selection can be made. Label source: Labels maintained in the instruction dataset label system;

[0066] (4) Label value Chinese name: This is a required option. Placehold: Select a label value (label value is displayed as [parent label - child label]). After the user enters the label, the system recommends labels that meet the conditions. Multiple labels can be selected (the corresponding threshold can be set independently) or a drop-down selection can be made. Label source: Labels maintained in the instruction dataset label system;

[0067] (5) Description: This is an optional item. Placehold: Enter the description, up to 256 characters.

[0068] (6) Experiment owner: The system automatically obtains the current user name;

[0069] (7) Space member: This is a required option. Placehold: Enter a descriptive username. After the user enters a username, the system recommends usernames that meet the requirements. Multiple users can be selected (and the corresponding threshold can be set independently).

[0070] Furthermore, regarding model experiments, in this embodiment, the experimental space can also be displayed in a list, specifically a card list, which includes the experiment name, description information, labels (the card displays one line by default, the excess is displayed with ellipsis, multiple are separated by commas, and all labels are displayed when the mouse is hovered), user name, update time (the first update time = creation time, and the subsequent update time after editing is used.), and operation bar; click the area in the card to enter the corresponding experimental space.

[0071] In addition, this embodiment also features model experiment editing: the last saved data is displayed, allowing users to edit and resubmit the data. The card list also displays the updated data. This embodiment also allows for transfer of ownership of the model experiment space. A pop-up window appears, allowing users to select another user as the owner of the current experiment space.

[0072] Step S12: generating data ratios based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information to determine multiple data ratio combinations.

[0073] In this embodiment, after determining the configured space, the system uses the selected tags for that space, the datasets with those tags, and the proportion of each dataset in the total dataset to generate various data ratio combinations (for example, data with tag A: data with tag B = 0.4:0.6, which is one combination; other combinations could be 0.3:0.7, 0.2:0.8, and so on). Based on these various data ratio combinations, a data solution is generated for each combination (for example, data with tag A: data with tag B = 0.4:0.6).

[0074] Combine Figure 2As shown, it should be understood that, regarding the generation of the ratio, in this embodiment, first, based on the multiple target label information corresponding to the configured space, the labeled training data sets corresponding to the various target label information are determined, and random balanced sampling is performed on each labeled training data set to obtain each balanced sampled data set with the same sample quantity; then, the upper and lower limits and the step size of the ratio corresponding to each target label information are configured to determine the ratio range information corresponding to each target label information; the preset number ratio of the number of samples corresponding to each balanced sampled data set in the total training data set is obtained; based on the ratio range information corresponding to the various target label information and the balanced sampled data set, as well as the preset number ratio corresponding to each balanced sampled data set, data ratio generation is performed to determine several data ratio combinations corresponding to the multiple label combinations. That is, in this embodiment, before determining the data ratio combination, it is necessary to set the data ratio range (upper and lower limits and step size) of the data set, balance the number of samples of different data sets in the combination through random balanced sampling, and determine the number ratio of the data sets. Among them, regarding random balanced sampling, for example: the data set corresponding to label 1 has 5,000 data items, and the data set corresponding to label 2 has 15,000 data items. It is necessary to randomly extract 5,000 data items from the data set corresponding to label 2, and discard the others, so that the number of samples in the two data sets is consistent.

[0075] Furthermore, when generating data ratios in the system page in this embodiment, the following operations are performed first in the initial state of the page:

[0076] (1) The [Set Data Ratio Range] column displays all the tags selected when creating the experimental space by default. Each tag can set the data set, upper and lower limits of the ratio, and step size. The maximum height of the page area is fixed, and a vertical scroll bar is displayed if the height is too high.

[0077] (2) Multiple data sets can be selected for each label. Here, the training data set is selected, which is a labeled instruction data set.

[0078] (3) The upper and lower limits and step size of the matching ratio can be manually entered. They are float (i.e. floating point data) and retain two decimal places. The step size of the matching ratio of each label to the corresponding data set may be different.

[0079] (4) After clicking the [Execute] button in the upper right corner of the page, the corresponding data will be generated in the [Proportion Result] at the bottom of the page. If the execution time is relatively long, the front end will give a placeholder image to indicate that it is being executed;

[0080] (5) In the [Matching] result column, before the execution results are obtained, the tabs [Matching Combination] and [View Data] below it are both visible.

[0081] After the execution result is returned, perform the following operations:

[0082] (1) [Ratio Combination] displays a list of all data ratio combinations and the data ratio scheme for each combination (a script for splitting the data according to the label ratio after random and balanced sampling of the data corresponding to each label and integrating the data after splitting). The ratio scheme can be downloaded locally;

[0083] (2) In the [Ratio Combination] list, there may be many labels. In this case, you can freeze the ratio scheme column and add a horizontal scroll bar to the entire list;

[0084] (3) The result of each execution may not be what you want, and you may need to modify the content of [Set Data Ratio Range] and execute again. After determining that the execution result meets the requirements, click the [Confirm Submit] button in the upper right corner of the page. After [Confirm Submit], a pop-up window will prompt whether to confirm the submission. Once submitted, it cannot be executed again. After clicking Confirm, the [Execute] button position will be replaced with the prompt text: Submission successful.

[0085] (4) [Proportioning Plan] The downloaded file is a JSON (JavaScript Object Notation, a lightweight data exchange format) file.

[0086] It is understandable that the JSON file finally downloaded contains the determined multiple data ratio combinations, and the model training can be performed based on this file later.

[0087] Step S13: In the configured space, by starting the pre-built Pipeline, the first parameter search batch task corresponding to the industry large model is initiated for various data ratio combinations and super-parameter range configuration files, and the current optimal target super-parameter range configuration file and the corresponding first trained model are screened out based on the corresponding first parameter search task results.

[0088] In this embodiment, combined with Figure 3 As shown in the figure, after the data ratio is generated, the pre-built pipeline is launched to perform a model parameter search. Specifically, a set of data ratios and model hyperparameter ranges are given to find the optimal hyperparameter configuration. This stage actually initiates a batch run task. Each task will have multiple jobs, each of which includes model training, model prediction, and evaluation. During training, a data ratio combination is selected and a hyperparameter range configuration file is uploaded. (Since hyperparameters are not set to specific values, but rather multiple values ​​are given for each hyperparameter, this is a hyperparameter range configuration file.) After the batch run task is completed, the results of multiple jobs are displayed. Using the automated evaluation results in each job result, the optimal hyperparameter configuration used in the corresponding model training task is found.

[0089] Specifically, when executing the batch task corresponding to the model parameter search, this embodiment initiates the parameter search batch subtask corresponding to the industry large model for any of the data ratio combinations based on the current data ratio combination and the different super-parameter range configuration files; then, in the process of running each of the parameter search batch subtasks of the industry large model, the corresponding balanced sampled data set is obtained based on the current data ratio combination, and data is randomly extracted from the balanced sampled data set based on the current data ratio combination to determine the first training set; the industry large model is trained based on the first training set and the super-parameter range configuration files corresponding to each of the parameter search batch subtasks of the industry large model. Model training to determine the corresponding trained model; perform model prediction on the trained model, and evaluate it based on the corresponding first prediction result and the first target evaluation data set corresponding to each of the parameter search batch subtasks of the industry large model to determine the first parameter search subtask results corresponding to each of the parameter search batch subtasks of the industry large model; determine the first parameter search task result corresponding to the current data ratio combination based on each of the first parameter search subtask results; based on the first preset screening rule and the first parameter search task results corresponding to various data ratio combinations, screen out the current optimal target super-parameter range configuration file and the first trained model corresponding to the target super-parameter range configuration file.

[0090] It's important to understand that creating a new batch task creates multiple jobs, each consisting of model training, model prediction, and automated evaluation. There are multiple jobs in a batch task, and the business scheduling layer queues them. The basic information for the task is as follows:

[0091] (1) Select the matching combination. This is a required option. You must select the data matching combination generated by the [Generate Matching Scheme] module. The number is limited to 1. Once the current batch task is submitted, the system will perform data sorting according to the data matching scheme.

[0092] (2) Model selection is a required option. Select a model that already exists in the platform warehouse;

[0093] (3) Evaluation data is a required option. Multiple data sets can be selected (and can be removed after selection). One example data set is extracted from each data set for list display. The instruction data set with labels is used.

[0094] (4) Evaluation method, which is a required option. The evaluation method of each evaluation data set can be set differently. After the model prediction results come out, the similarity between the prediction results and the reference answers will be compared;

[0095] (5) Arena parameters: The interface provides a multi-line text box, supports a maximum input of 10240, and supports setting default values;

[0096] (6) Model hyperparameters: The interface provides a multi-line text box, supports a maximum input of 10240, and supports setting default values.

[0097] Regarding the batch task list, the basic information of the list is as follows:

[0098] (1) Search bar: You can search based on the creation time and the task ID (Identity Document) of the batch task. The search results will only display the matching data. Clicking the [Reset] button will clear the search conditions in the bar and the list will return to its initial state.

[0099] (2) List header: By default, the list is sorted in reverse order of update time, with the latest one displayed at the top, including serial number, task ID, number of completed jobs, number of failed jobs, number of running jobs, job running time, creator, creation time, and operation. Specifically:

[0100] ①The serial number and task ID are automatically generated by the system;

[0101] ② The number of completed jobs is used to count the number of jobs corresponding to the current batch task that have successfully run;

[0102] ③ The number of failed jobs is used to count the number of jobs corresponding to the current batch task that have failed to run;

[0103] ④ Number of running jobs, used to count the number of jobs corresponding to the current batch task that are in the running state (including those in the queue);

[0104] ⑤ Runtime, which is the time from the start of the first job to the completion of the last job;

[0105] ⑥ Creator: the person who created the batch task;

[0106] ⑦ Creation time: the date and time when the batch task is created and submitted for saving, which can be accurate to the second;

[0107] ⑧ For all batch tasks, you can view the job details (switch to the current page) and execute deletion (soft deletion). When deleting, a second warning will be given: "Are you sure you want to delete this batch task? The model will also be deleted when deleting the batch task. Please operate with caution! Task deletion requires the business scheduling side to terminate the task and release model resources."

[0108] ⑨ Click on the task ID of any row of data to link to the current batch task details page (when the current page is switched, the task details page will only display the full information saved last time, and editing is not allowed).

[0109] (3) Click the task ID to view the details of the batch task. The job list header displays: job number, hyperparameter configuration, average score, task status, operation, and more.

[0110] ①Job number, automatically generated by the system;

[0111] ② Hyperparameter configuration: used to display part of the hyperparameter content. The content beyond that is terminated with an ellipsis. Hovering the mouse can view the entire content.

[0112] ③ Running time, which is the time from the start to the end of the job. There is no running time;

[0113] ④ Average score. Only jobs in the successful run state have scores. Scores in other states are displayed with a horizontal bar.

[0114] ⑤Task status, including running, successful, and failed. Only when the training, prediction, and evaluation stages are all successful can the task be considered successful.

[0115] ⑥ Operation, including No Operation (for jobs in the Running, Failed, or Queued states) and View Evaluation Results (for jobs in the Successful Running state);

[0116] (4) Click the Download button to download the evaluation results of all current jobs and get an EXCEL file. The first sheet in the file displays the job list, and the following sheets use the job number as the sheet name to display the evaluation results of each job.

[0117] (5) Check the evaluation results. The upper left corner of the page displays the job number of the current job and the average score of the evaluation results, as well as the evaluation result header (including the dataset ID, the ID of the data in the dataset, the question, the reference answer, the reasoning answer, and the evaluation result). Specifically:

[0118] ① Dataset ID, number ID, question, and reference answer are all information contained in the data itself;

[0119] ② The inference answer is the inference answer obtained by running the problem based on the model prediction result made based on the last checkpoint after the model training of the current job is completed;

[0120] ③Evaluation results: This is the similarity score between the reference answer and the inference answer obtained based on the evaluation method selected when creating the batch task;

[0121] ④The average score in the upper left corner of the page is the mean of all evaluation results.

[0122] Step S14: Initiate a second parameter search batch task corresponding to the first trained model based on the target hyperparameter range configuration file and the multiple data ratio combinations, and screen out the current optimal target data ratio combination and the corresponding second trained model based on the corresponding second parameter search task results.

[0123] Combine Figure 3 As shown, in this embodiment, after performing the first parameter search task corresponding to the model parameter search and selecting the currently optimal target hyperparameter range configuration file and the corresponding first trained model, a batch task corresponding to the parameter search is initiated. Given the target hyperparameter range configuration file and multiple data ratio combinations, the optimal data ratio combination is found. This stage also initiates a batch task, each of which has multiple jobs, each of which also includes model training, model prediction, and evaluation. However, during training, a fixed set of hyperparameters is set or uploaded (rather than a range). The system can default to recommending the target hyperparameter range configuration file obtained in the previous model parameter search stage, and users can also change it based on their needs. Furthermore, a range of data ratio combinations is defined (which can be different from or the same as the range in the model parameter search stage). After the batch task is completed, the results of multiple jobs are displayed. Using the automated evaluation results in each job result, the optimal data ratio combination used in the corresponding model training task is found (the top three optimal combinations can be displayed, or a different number can be set based on actual needs). The implementation scheme for creating new batch tasks and batch task lists in this stage is consistent with that in the model parameter search stage.

[0124] It should be understood that regarding the execution of the batch task corresponding to the parameter search, in this embodiment, the parameter search batch subtask corresponding to the first trained model is first initiated based on the target super-parameter range configuration file and in combination with various data ratio combinations; then, in the process of running any of the parameter search batch subtasks of the first trained model, the corresponding second training set is determined based on the data ratio combination corresponding to the current parameter search batch subtask of the first trained model and the corresponding balanced sampling data set; the first trained model is trained based on the second training set and the target super-parameter range configuration file to determine the trained first trained model; the trained first trained model is model predicted, and evaluated based on the second prediction result and the second target evaluation data set corresponding to the current parameter search batch subtask of the first trained model to determine the corresponding second parameter search subtask result; the second parameter search task result corresponding to the target super-parameter range configuration file is determined based on the results of each of the second parameter search subtasks; the current optimal target data ratio combination and the second trained model corresponding to the target data ratio combination are screened out based on the second preset screening rule and the result of the second parameter search task.

[0125] Step S15: Based on the current target data ratio combination and the second trained model, determine whether the preset Pipeline iteration termination condition is currently met, and if so, determine whether the model parameter adjustment operation and training data ratio adjustment operation corresponding to the industry large model have been completed.

[0126] Specifically, after completing a round of model parameter search and ratio parameter search batch tasks, this embodiment determines whether to terminate the training by judging whether convergence has occurred, that is, by judging whether the current target data ratio combination is consistent with the data ratio combination corresponding to the target hyperparameter range configuration file, or judging whether the current second post-training model has converged, to determine whether the preset Pipeline iteration termination condition is currently met. If the preset Pipeline iteration termination condition is not currently met, then jump back to the step of initiating the first parameter search batch task corresponding to the industry large model for various data ratio combinations and hyperparameter range configuration files. You can check whether the ratio combination screened out in the ratio search stage is the same as the ratio combination in the model parameter search stage to determine whether convergence has occurred. If they are the same, the purpose of the experiment is achieved and the experiment is terminated. If they are not the same, a new round of model parameter search and ratio search batch tasks will continue to be executed until termination.

[0127] In addition, regarding Pipeline, this embodiment uses Pipeline builds to automatically execute multiple rounds of model parameter search and parameter ratio search batch tasks. One model experiment only runs one Pipeline build. Specifically, in Pipeline, the relevant information about creating a new batch task is:

[0128] (1) Basic information of the task:

[0129] ① Select a matching combination. This is a required option. You must select the data matching combination generated by the [Generate Matching Plan] module. Only one matching combination is allowed. Once the current batch task is submitted, the system will perform data sorting according to the data matching plan.

[0130] ②Model selection is a required option. Select a model already in the platform warehouse;

[0131] ③Evaluation data is a required option. Multiple data sets can be selected (and can be removed after selection). One example data set is extracted from each data set for list display. The instruction data set with labels is used.

[0132] ④Evaluation method, which is a required option. The evaluation method of each evaluation data set can be set differently. After the model prediction results come out, the similarity between the prediction results and the reference answers is compared;

[0133] ⑤arena parameters, the interface provides a multi-line text box, supports a maximum input of 10240, and supports setting default values;

[0134] ⑥Model hyperparameters: The interface provides a multi-line text box, supports a maximum input of 10240, and supports setting default values;

[0135] ⑦ Pipeline termination condition: the optimal hyperparameter or ratio combination of the current round is exactly the same as that of the previous round, or the difference between the highest average score of round N+1 (N≥3) and that of round N is ≤ the preset threshold (can be manually filled in the interface, unit %).

[0136] (2) Task environment configuration:

[0137] ① Supports selection of cluster, GPU (Graphics Processing Unit) type, number of replicas, and number of cards per replica;

[0138] ② The current batch task will have multiple jobs (including training, prediction, and evaluation), and the business scheduling layer will queue them.

[0139] (3) Once you select [Submit] on the page, the queue will be executed.

[0140] (4) Pipeline termination conditions:

[0141] ① The top 3 optimal data ratio combinations generated by the ratio search are the same as the given ratio combinations of the model search.

[0142] ② The model hyperparameters obtained in the N+1 round of model parameter search do not change compared to the model hyperparameters obtained in the N round of model parameter search.

[0143] At the same time, regarding the execution of batch tasks, you can view the task list and running status. Click the task ID to view the details. The details page does not allow editing.

[0144] In this way, this embodiment provides an industry large model parameter adjustment solution based on dynamic data matching, which plans and adjusts the distribution ratio of data classified by labels in different scenarios on the platform to achieve the output and optimization of the industry large model. In the experimental design, the sample size of the control group and the experimental group can be reasonably allocated to ensure the reliability of the results. A variety of matching schemes are formed by adjusting the upper and lower limits of the matching and the step size. A fixed matching combination sets multiple sets of parameters. After the model continuously runs batches of tasks, the highest score of the job is obtained. The corresponding parameters in this job are the optimal parameters for this matching combination. Reasonable data matching and parallel batch job ranking will give the optimal model training parameters, which directly affect the model generalization ability, the accuracy of the analysis results and the scientific nature of the experimental conclusions. The beneficial effects are as follows:

[0145] (1) Improve parameter tuning efficiency and accuracy: Through dynamic data ratio generation and hyperparameter range definition, multiple groups of parallel experiments are automatically initiated to quickly screen the optimal configuration, shortening the experimental cycle by more than 50%. Combined with the alternating iteration of model search and ratio search, global optimization of the parameter space is achieved, improving model accuracy and generalization ability;

[0146] (2) Reduce labor costs: Fully automated management of experimental processes (such as ratio scheme generation, task queuing and scheduling, and evaluation result aggregation) reduces manual operation errors and repetitive work. Supports multi-person collaboration and authority control to improve team experimental collaboration efficiency;

[0147] (3) Enhance model robustness: Through random balanced sampling and dynamic ratio adjustment of the data sets corresponding to each label, the model bias problem caused by uneven data distribution is alleviated, and the prediction stability in long-tail scenarios is improved;

[0148] (4) Resource utilization optimization: Based on the task queuing mechanism and dynamic allocation of hardware resources at the business scheduling layer, efficient utilization of GPU clusters is achieved to avoid resource idleness or conflicts;

[0149] (5) Scientific decision support: Provide visual evaluation results (such as score ranking, ratio combination comparison), support data-driven parameter adjustment decisions, enhance the interpretability of experimental conclusions, avoid invalid iterations through Pipeline termination conditions (such as parameter convergence threshold), and ensure the scientificity and reliability of experimental results;

[0150] (6) Strong scalability: During implementation, modular design is carried out for different stages (such as model experiment management and ratio scheme generation) to support flexible expansion to other different industry scenarios and adapt to diverse labeling systems and evaluation requirements.

[0151] It can be seen that in this application, a space corresponding to the industry large model for matching experiments is first created and configured, and then the data matching is generated for the various target label information and the corresponding labeled training data set corresponding to the space. After that, the pre-built Pipeline is started in the space, and the first parameter search and batch running task corresponding to the industry large model is initiated for the various data matching combinations and super-parameter range configuration files generated, and the optimal target super-parameter range configuration file and the first trained model are screened out based on the corresponding first parameter search task results. After that, the second parameter search and batch running task corresponding to the first trained model is initiated based on the target super-parameter range configuration file and the various data matching combinations, and the current optimal target data matching combination and the second trained model are screened out, and when the preset Pipeline iteration termination conditions are met, the model parameter adjustment operation and the training data matching tuning operation corresponding to the industry large model are determined to be completed. In this way, the automatic generation of data matching and the joint optimization of hyperparameters and training data matching can be realized, thereby improving the efficiency of hyperparameter search and model training, reducing the waste of computer resources and human resources, and thus improving resource utilization, and improving the analysis accuracy and conclusion reliability of the industry large model in complex scenarios.

[0152] See also Figure 4 As shown, the embodiment of the present application also discloses a model training device for a large industry model, including:

[0153] The experimental space creation module 11 is used to create a space for performing a matching experiment corresponding to the industry large model, and trigger an information configuration operation on the determined model experimental space to obtain a configured space;

[0154] A data ratio generating module 12 is configured to generate data ratios based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information, so as to determine multiple data ratio combinations;

[0155] A first parameter search module 13 is configured to initiate, within the configured space, a first parameter search batch task corresponding to the industry large model for various data ratio combinations and hyperparameter range configuration files by starting a pre-built pipeline, and to screen out the currently optimal target hyperparameter range configuration file and the corresponding first trained model based on the corresponding first parameter search task results;

[0156] A second parameter search module 14 is configured to initiate a second parameter search batch task corresponding to the first trained model based on the target super-parameter range configuration file and the multiple data ratio combinations, and to screen out the currently optimal target data ratio combination and the corresponding second trained model based on the corresponding second parameter search task results;

[0157] The iteration termination module 15 is used to determine whether the preset Pipeline iteration termination condition is currently met based on the current target data ratio combination and the second trained model, and if so, determine whether the model parameter adjustment operation and training data ratio adjustment operation corresponding to the industry large model have been completed.

[0158] In this application, the automatic generation of data ratios and the joint optimization of hyperparameters and training data ratios can be achieved, thereby improving the efficiency of hyperparameter search and model training, reducing the waste of computer resources and human resources, and thus improving resource utilization, and improving the analysis accuracy and conclusion reliability of large industry models in complex scenarios.

[0159] In some specific embodiments, the model training device for the industry large model may further include:

[0160] The step jump module is used to jump back to the step of initiating the first parameter search batch task corresponding to the industry large model for various data ratio combinations and hyperparameter range configuration files if the preset Pipeline iteration termination condition is not currently met.

[0161] In some specific embodiments, the experimental space creation module 11 may specifically include:

[0162] A space creation unit is used to create a space corresponding to the industry large model for conducting a proportioning experiment, so as to obtain a model experiment space;

[0163] The space configuration unit is used to configure the experiment name, label information, experiment space information display list, administrator information and accessible list corresponding to the model experiment space to determine the configured space.

[0164] In some specific embodiments, the data ratio generation module 12 may specifically include:

[0165] A balanced sampling unit is configured to determine labeled training data sets corresponding to various target label information based on the multiple target label information corresponding to the configured space, and to perform random balanced sampling on each labeled training data set to obtain balanced sampled data sets with consistent sample numbers;

[0166] a ratio range determination unit, configured to configure the ratio upper and lower limits and the ratio step corresponding to each target tag information to determine the ratio range information corresponding to each target tag information;

[0167] A quantity ratio obtaining unit, configured to obtain a preset quantity ratio of the number of samples corresponding to each of the balanced sampling data sets in the total training data set;

[0168] A matching combination determination unit is used to generate data matching based on the matching range information corresponding to the various target label information and the balanced sampled data sets, as well as the preset quantity proportions corresponding to each balanced sampled data set, so as to determine several data matching combinations corresponding to multiple label combinations.

[0169] In some specific embodiments, the first parameter search module 13 may specifically include:

[0170] The first batch subtask initiating unit is used to initiate, for any of the data ratio combinations, a parameter search batch subtask corresponding to the industry large model based on the current data ratio combination and different super-parameter range configuration files;

[0171] A first training set determination unit is configured to obtain the corresponding balanced sampled data set based on the current data ratio combination during the execution of each of the parameter search and batch running subtasks of the industry large model, and randomly extract data from the balanced sampled data set based on the current data ratio combination to determine a first training set;

[0172] A first model training unit is configured to perform model training on the industry large model based on the first training set and the hyper-parameter range configuration files corresponding to the respective parameter search and batch running subtasks of the industry large model to determine a corresponding trained model;

[0173] A first model prediction unit is used to perform model prediction on the trained model, and to perform evaluation based on the corresponding first prediction result and the first target evaluation data set corresponding to each of the parameter search and batch running subtasks of the industry large model, so as to determine the first parameter search subtask results corresponding to each of the parameter search and batch running subtasks of the industry large model;

[0174] A first task result determining unit, configured to determine a first parameter search task result corresponding to the current data matching combination based on each of the first parameter search subtask results;

[0175] The first screening unit is used to screen out the current optimal target super-parameter range configuration file and the first post-training model corresponding to the target super-parameter range configuration file based on the first preset screening rule and the first search task results corresponding to the various data ratio combinations.

[0176] In some specific embodiments, the second parameter search module 14 may specifically include:

[0177] A second batch subtask initiating unit is configured to initiate a parameter search batch subtask corresponding to the first trained model based on the target hyperparameter range configuration file and in combination with various data ratio combinations;

[0178] A second training set determining unit is configured to determine a corresponding second training set based on the data ratio combination corresponding to the current parameter search and batch running subtask of the first trained model and the corresponding balanced sampling data set during the execution of any of the parameter search and batch running subtasks of the first trained model;

[0179] A second model training unit is configured to perform model training on the first trained model based on the second training set and the target hyperparameter range configuration file to determine the trained first trained model;

[0180] a second model prediction unit, configured to perform model prediction on the first trained model after training, and to evaluate the model based on the second prediction result and a second target evaluation data set corresponding to the current parameter search batch subtask of the first trained model to determine a corresponding second parameter search subtask result;

[0181] A second task result determining unit, configured to determine a second parameter search task result corresponding to the target super-parameter range configuration file based on each of the second parameter search subtask results;

[0182] The second screening unit is used to screen out the current optimal target data ratio combination and the second trained model corresponding to the target data ratio combination based on the second preset screening rules and the second parameter search task results.

[0183] In some specific embodiments, the iteration termination module 15 may specifically include:

[0184] The iteration termination judgment unit is used to judge whether the current target data ratio combination is consistent with the data ratio combination corresponding to the target hyperparameter range configuration file, or to judge whether the current second trained model converges, so as to determine whether the preset Pipeline iteration termination condition is currently met.

[0185] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.

[0186] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the model training method for the large industry model disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0187] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0188] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0189] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the model training method for the industry large model executed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs capable of completing other specific tasks.

[0190] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned disclosed model training method for large industry models. The specific steps of this method can be referred to the corresponding content disclosed in the aforementioned embodiments and will not be repeated here.

[0191] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0192] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0193] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0194] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0195] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A model training method for a large industry model, characterized in that: include: Create a space for matching experiments corresponding to the industry model, and trigger information configuration operations on the determined model experiment space to obtain the configured space; Performing data matching based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information to determine multiple data matching combinations; In the configured space, by starting the pre-built Pipeline, the first parameter search batch task corresponding to the industry large model is initiated for various data ratio combinations and hyperparameter range configuration files, and the currently optimal target hyperparameter range configuration file and the corresponding first trained model are screened based on the corresponding first parameter search task results; Initiate a second parameter search batch task corresponding to the first trained model based on the target hyperparameter range configuration file and the multiple data ratio combinations, and screen out the current optimal target data ratio combination and the corresponding second trained model based on the corresponding second parameter search task results; Based on the current target data ratio combination and the second trained model, determine whether the preset Pipeline iteration termination condition is currently met, and if so, determine that the model parameter adjustment operation and training data ratio tuning operation corresponding to the industry large model have been completed.

2. The model training method for industry large models according to claim 1 is characterized in that: Also includes: If the preset Pipeline iteration termination condition is not met at present, the process jumps back to the step of initiating the first parameter search batch task corresponding to the industry large model for various data ratio combinations and hyper-parameter range configuration files.

3. The model training method for industry large models according to claim 1 is characterized in that: The step of creating a space for performing a matching experiment corresponding to the industry macro model and triggering an information configuration operation on the determined model experiment space includes: Create a space for proportioning experiments corresponding to the industry model to obtain a model experiment space; The experiment name, label information, experiment space information display list, administrator information and accessible list corresponding to the model experiment space are configured to determine the configured space.

4. The model training method for an industry large model according to any one of claims 1 to 3, characterized in that: The data matching generation based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information to determine multiple data matching combinations includes: Determining labeled training data sets corresponding to various target label information based on the multiple target label information corresponding to the configured space, and performing random balanced sampling on each labeled training data set to obtain balanced sampled data sets with consistent sample numbers; Configuring the upper and lower limits and the step length of the ratio corresponding to each target tag information to determine the ratio range information corresponding to each target tag information; Obtaining a preset ratio of the number of samples corresponding to each balanced sampling data set in the total training data set; Data ratio generation is performed based on the ratio range information corresponding to the various target label information and the balanced sampled data sets, as well as the preset quantity proportions corresponding to each balanced sampled data set, to determine several data ratio combinations corresponding to multiple label combinations.

5. The model training method for industry large models according to claim 4 is characterized in that, The first parameter search batch task corresponding to the industry large model is initiated for various data ratio combinations and super-parameter range configuration files, and the currently optimal target super-parameter range configuration file and the corresponding first trained model are screened out based on the corresponding first parameter search task results, including: For any of the data ratio combinations, based on the current data ratio combination and different hyper-parameter range configuration files, initiate the parameter search and batch running subtasks corresponding to the industry large model; In the process of running each of the parameter search and batch running subtasks of the industry large model, the corresponding balanced sampling data set is obtained based on the current data ratio combination, and data is randomly extracted from the balanced sampling data set based on the current data ratio combination to determine the first training set; Based on the first training set and the super-parameter range configuration files corresponding to each of the parameter search and batch running subtasks of the industry large model, the industry large model is trained to determine a corresponding trained model; Performing model prediction on the trained model, and performing evaluation based on the corresponding first prediction result and the first target evaluation data set corresponding to each of the parameter search and batch running subtasks of the industry large model, to determine the first parameter search subtask results corresponding to each of the parameter search and batch running subtasks of the industry large model; Determine a first parameter search task result corresponding to the current data matching combination based on each of the first parameter search subtask results; Based on the first preset screening rule and the first parameter search task results corresponding to the various data ratio combinations, the current optimal target super-parameter range configuration file and the first trained model corresponding to the target super-parameter range configuration file are screened out.

6. The model training method for industry large models according to claim 1 is characterized in that: The initiating a second parameter search batch running task corresponding to the first trained model based on the target hyperparameter range configuration file and the multiple data ratio combinations, and screening out the current optimal target data ratio combination and the corresponding second trained model based on the corresponding second parameter search task results, includes: Based on the target hyperparameter range configuration file, and in combination with various data ratio combinations, initiate parameter search and batch running subtasks corresponding to the first trained model; During the execution of any of the parameter search and batch running subtasks of the first trained model, determining a corresponding second training set based on the data ratio combination corresponding to the current parameter search and batch running subtask of the first trained model and the corresponding balanced sampling data set; Performing model training on the first trained model based on the second training set and the target hyperparameter range configuration file to determine the trained first trained model; Performing model prediction on the first trained model after training, and evaluating based on a second prediction result and a second target evaluation data set corresponding to the current parameter search batch subtask of the first trained model to determine a corresponding second parameter search subtask result; Determine a second parameter search task result corresponding to the target super-parameter range configuration file based on each of the second parameter search subtask results; Based on the second preset screening rule and the second parameter search task result, the current optimal target data ratio combination and the second trained model corresponding to the target data ratio combination are screened out.

7. The model training method for industry large models according to claim 1 is characterized in that: The determining whether a preset Pipeline iteration termination condition is currently satisfied based on the current target data ratio combination and the second trained model includes: Determine whether the current target data ratio combination is consistent with the data ratio combination corresponding to the target hyperparameter range configuration file, or determine whether the current second trained model converges to determine whether the preset Pipeline iteration termination condition is currently met.

8. A model training device for large industry models, characterized in that: include: The experimental space creation module is used to create a space corresponding to the industry large model for conducting ratio experiments, and trigger information configuration operations on the determined model experimental space to obtain the configured space; a data ratio generation module, configured to generate data ratios based on the multiple target label information corresponding to the configured space and the labeled training data sets corresponding to the multiple target label information, so as to determine multiple data ratio combinations; A first parameter search module is configured to initiate a first parameter search batch task corresponding to the industry large model for various data ratio combinations and hyperparameter range configuration files within the configured space by starting a pre-built pipeline, and to screen out the currently optimal target hyperparameter range configuration file and the corresponding first trained model based on the corresponding first parameter search task results; A second parameter search module is used to initiate a second parameter search batch task corresponding to the first trained model based on the target super-parameter range configuration file and the multiple data ratio combinations, and screen out the current optimal target data ratio combination and the corresponding second trained model based on the corresponding second parameter search task results; The iteration termination module is used to determine whether the preset Pipeline iteration termination condition is currently met based on the current target data ratio combination and the second trained model, and if so, determine whether the model parameter adjustment operation and training data ratio tuning operation corresponding to the industry large model have been completed.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the model training method for a large industry model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that Used to store a computer program, which, when executed by a processor, implements the model training method for industry large models as described in any one of claims 1 to 7.