Model training and business execution method and device, storage medium and equipment
By dynamically updating the data sampling strategy according to the loss value during each round of model training, the problem that fixed data sampling strategies in the prior art are difficult to adapt to model parameter changes is solved, and the efficiency and performance of model training are improved.
Patent Information
- Application Number
- CN202510194257.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
The existing data sampling methods are based on fixed strategies and are difficult to adapt to the dynamic changes of model parameters, resulting in inefficient and poor performance of model training.
During each round of model training, the data sampling strategy is dynamically updated according to the loss value of the previous round to adapt to the changes in model parameters.
By dynamically updating the data sampling strategy, it can more effectively match the actual situation of the model in each round of training, and improve the efficiency and performance of model training.
Smart Images

Figure CN120124775A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular, to a method, apparatus, storage medium, and device for model training and service execution. Background Art
[0002] With the development of artificial intelligence technology, deep learning models play an increasingly important role in various services such as image recognition, natural language processing, anomaly detection, and privacy protection. During the training process of the model, the quality, quantity, and distribution of sample data have an important impact on the performance of the model. Therefore, the data sampling technology has emerged. This technology samples different datasets by specifying the sampling probability under a data sampling strategy, thereby optimizing the distribution of training data.
[0003] However, existing data sampling methods usually sample sample data in different datasets based on a fixed data sampling strategy. However, as the model parameters are continuously updated during the training process, it is difficult for similar sample data obtained based on the same data sampling strategy as in previous training rounds to further improve the performance of the model, resulting in low training efficiency of the model and even affecting the accuracy of the model.
[0004] Therefore, how to avoid the impact of the fixed data sampling strategy on the model training efficiency and performance is an urgent problem to be solved. Summary of the Invention
[0005] This specification provides a method, apparatus, storage medium, and device for model training and service execution to update the data sampling strategy based on the loss value obtained during each round of model training.
[0006] This specification adopts the following technical solutions:
[0007] This specification provides a model training method, including:
[0008] Receiving a training request for a service model;
[0009] For each round of training of the service model except the first round according to the training request, based on the loss value obtained during the previous round of training of the service model, determining the data sampling strategy corresponding to this round of training;
[0010] Sampling sample data in different service datasets based on the data sampling strategy corresponding to this round of training to obtain the sample data for this round of training;
[0011] Inputting the sample data into the service model to determine the prediction result for the sample data through the service model;
[0012] According to the prediction result, determine the loss value of the business model in this round of training, and train the business model according to the loss value of the business model in this round of training. Moreover, according to the loss value of the business model in this round of training, determine the data sampling strategy to be used in the next round of training of the business model.
[0013] Optionally, determining the data sampling strategy to be used in the next round of training of the business model according to the loss value of the business model in this round of training specifically includes:
[0014] According to the loss value of the business model in this round of training and the probability of sampling each business data set under the data sampling strategy corresponding to this round of training, determine the reward value of the data sampling strategy corresponding to this round of training, where the greater the loss value of the business model in this round of training, the greater the reward value of the data sampling strategy corresponding to this round of training;
[0015] According to the reward value of the data sampling strategy corresponding to this round of training, determine the data sampling strategy to be used in the next round of training of the business model.
[0016] Optionally, determining the data sampling strategy to be used in the next round of training of the business model according to the reward value of the data sampling strategy corresponding to this round of training specifically includes:
[0017] According to the number of business data sets and the corresponding training round number for the next round of training of the business model, determine the exploration rate for the data sampling strategy for the next round of training of the business model;
[0018] According to the reward value of the data sampling strategy corresponding to this round of training, the probability of sampling each business data set under the data sampling strategy corresponding to this round of training, and the exploration rate for the data sampling strategy for the next round of training of the business model, determine the data sampling strategy to be used in the next round of training of the business model.
[0019] Optionally, determining the data sampling strategy to be used in the next round of training of the business model according to the reward value of the data sampling strategy corresponding to this round of training and the exploration rate for the data sampling strategy for the next round of training of the business model specifically includes:
[0020] According to the number of business data sets and the corresponding number of training rounds for this training of the business model, determine the exploration rate for the data sampling strategy for this round of training of the business model;
[0021] Determine the data sampling strategy to be used for the next round of training of the service model based on the reward value of the data sampling strategy corresponding to this round of training, the probability of sampling each service data set under the data sampling strategy corresponding to this round of training, the exploration rate for the data sampling strategy when performing this round of training on the service model, and the exploration rate for the data sampling strategy when performing the next round of training on the service model.
[0022] Optionally, when the number of training rounds for training the service model exceeds the specified number of training rounds, the exploration rate of the data sampling strategy decreases as the number of training rounds increases.
[0023] Optionally, determine the reward value of the data sampling strategy corresponding to this round of training according to the loss value of the service model in this round of training and the probability of sampling each service data set under the data sampling strategy corresponding to this round of training. Specifically, it includes:
[0024] Determine the reward value of the data sampling strategy corresponding to this round of training according to the loss value of the service model in this round of training, the decay coefficient preset for the probability of sampling each service data set under the data sampling strategy corresponding to this round of training, and the historical reward values corresponding to each historical data sampling strategy determined before this round of training.
[0025] Optionally, determine the reward value of the data sampling strategy corresponding to this round of training according to the loss value of the service model in this round of training, the decay coefficient preset for the probability of sampling each service data set under the data sampling strategy corresponding to this round of training, and the historical reward values corresponding to each historical data sampling strategy determined before this round of training. Specifically, it includes:
[0026] Determine the weight corresponding to the historical reward value according to the decay coefficient as the first weight, and determine the weight corresponding to the reward value determined based on the loss value of the service model in this round of training and the probability of sampling each service data set under the data sampling strategy corresponding to this round of training as the second weight, where there is a negative correlation between the first weight and the second weight;
[0027] Weight the historical reward value by the first weight to obtain a first weighted reward value, and weight the reward value determined based on the loss value of the service model in this round of training and the probability of sampling each service data set under the data sampling strategy corresponding to this round of training by the second weight to obtain a second weighted reward value;
[0028] Determine the reward value of the data sampling strategy corresponding to this round of training according to the first weighted reward value and the second weighted reward value.
[0029] This specification provides a service execution method, including:
[0030] Obtain service data;
[0031] Input the service data into a pre-trained service model to determine a prediction result for the service data through the service model, and execute a service according to the prediction result, where the service model is trained by the above model training method.
[0032] This specification provides a model training device, including:
[0033] A receiving module, configured to receive a training request for a service model;
[0034] A determining module, configured to, for each round of training of the service model except the first round of training according to the training request, determine a data sampling strategy corresponding to this round of training based on the loss value obtained when training the service model in the previous round;
[0035] A sampling module, configured to sample sample data in different service data sets based on the data sampling strategy corresponding to this round of training to obtain the sample data for this round of training;
[0036] An input module, configured to input the sample data into the service model to determine a prediction result for the sample data through the service model;
[0037] A training module, configured to determine the loss value of the service model in this round of training according to the prediction result, train the service model according to the loss value of the service model in this round of training, and determine the data sampling strategy to be used when performing the next round of training on the service model according to the loss value of the service model in this round of training.
[0038] This specification provides a service execution device, including:
[0039] An obtaining module, configured to obtain service data;
[0040] An execution module, configured to input the service data into a pre-trained service model to determine a prediction result for the service data through the service model, and execute a service according to the prediction result, where the service model is trained by the above model training method.
[0041] This specification provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above model training and service execution methods are implemented.
[0042] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned model training and service execution methods are implemented.
[0043] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0044] In the model training method provided in this specification, for each round of training of the service model except the first round, based on the loss value obtained when training the service model in the previous round, determine the data sampling strategy corresponding to this round of training; based on the data sampling strategy of this round of training, sample the sample data in different service data sets to obtain the sample data of this round of training; input the sample data into the service model to determine the prediction result for the sample data; according to the prediction result, determine the loss value of the service model and train the service model, and, according to the loss value of the service model in this round of training, determine the data sampling strategy to be used when performing the next round of training on the service model.
[0045] As can be seen from the above method, in the process of model training in this solution, the data sampling strategy can be updated based on the loss value determined in each round of training, so as to obtain the data sampling strategy for the next round of training the model and perform data sampling. Compared with the current method of performing data sampling based on a fixed data sampling strategy, this solution can continuously update the data sampling strategy during the model training process, so that the data sampling strategy determined each time can match the actual training situation of the model in each round. The training samples collected based on this sampling strategy can effectively improve its performance when training the model in each round, further improving the overall efficiency of model training. Description of the Drawings
[0046] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The schematic embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation to this specification. In the attached
[0047] In the figure:
[0048] Figure 1 It is a schematic flow chart of a model training method provided in this specification;
[0049] Figure 2 It is a schematic flow chart of the update of a data sampling strategy provided in this specification;
[0050] Figure 3 It is a schematic flow chart of a service execution method provided in this specification;
[0051] Figure 4 Schematic diagram of a model training device provided in this specification;
[0052] Figure 5 Schematic diagram of a service execution device provided in this specification;
[0053] Figure 6 For a corresponding one provided in this specification Figure 1 or Figure 3 Schematic diagram of an electronic device. Detailed implementation manners
[0054] To make the objectives, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of this specification. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0055] In traditional data mixing algorithms, the sampling probability of each data group is fixed before training. Although this method is simple and easy to implement, it lacks flexibility and is difficult to adjust the data sampling strategy according to the progress of model training.
[0056] Taking the DoReMi data ratio algorithm as an example, this algorithm improves the training efficiency of the model by automatically determining the data mixing ratio. Its core idea is: to optimize the weights corresponding to different data sets by comparing the information gain between a proxy model and a reference model, and then to find the best data sampling strategy by maximizing the gain information of the proxy model relative to the reference model.
[0057] However, the DoReMi algorithm needs to train multiple models to determine the optimal domain weights, which results in a high computational cost and low efficiency. Moreover, the sampling weights determined by the DoReMi algorithm have poor transferability between different model architectures or tokenizers. This means that to adapt to a new model architecture or tokenizer, it is necessary to retrain the "reference" and "proxy" models to determine the optimal weights. In addition, like the DoReMi and The Pile data sets, the weights are fixed during training and cannot adapt to the dynamic changes during training. Furthermore, the DoReMi algorithm requires additional calculation steps and considerations when calculating the domain weights, which seriously increases the training cost of the model.
[0058] The following details the technical solutions provided in each embodiment of this specification in conjunction with the drawings.
[0059] Figure 1A flowchart showing a model training method provided in this specification, including the following steps:
[0060] S100: Receive a training request for a service model;
[0061] S102: Based on the training request, for each round of training of the service model except the first round of training, determine the data sampling strategy corresponding to this round of training based on the loss value obtained when training the service model in the previous round.
[0062] To address the impact of a fixed data sampling strategy on model training efficiency and performance, this specification provides a model training method that updates the data sampling strategy used in the next round of training based on the loss value generated during each round of training of the model, thereby continuously adjusting the data sampling strategy to adapt to the dynamic changes of model parameters.
[0063] In this specification, the execution entity for implementing a model training and service execution method can be a specified device such as a server. For the sake of convenience in description, hereinafter, only the server will be used as an example of the execution entity to illustrate a model training and service execution method provided in this specification.
[0064] Among them, the server can receive a training request for a service model. In this specification, there can be multiple types of such service models, such as an image recognition model, a natural language processing model, an anomaly recognition model, a privacy protection model, etc. This specification does not make specific limitations in this regard.
[0065] After receiving the training request, the server can perform several rounds of training on the service model. For the first round of training, the server can sample the sample data for this round of training in different service datasets based on a preset initial sampling strategy.
[0066] For each round of training t of the service model except the first round of training, the server can determine the data sampling strategy π corresponding to round t based on the training request and the loss value obtained when training the service model in the previous round (t - 1) t (For the determination process of this sampling strategy, please refer to the following text. This specification will not elaborate too much here).
[0067] Among them, the above data sampling strategy includes the probability of sampling sample data in each business dataset. In practical applications, the data sources of different business datasets (such as public datasets, the Internet, social media, manual annotation, crowdsourcing platforms, Internet of Things devices, etc.), data types (such as images, text, audio, etc.), data formats (such as Comma-Separated Values (CSV), Excel, JavaScript Object Notation (JSON), etc.), and at least one of the business fields involved (such as finance, aerospace, medical, etc.) are different.
[0068] For the business models of different businesses, the types of sample data collected are also different. Among them, for an image recognition model, its corresponding sample data is mainly image data, while for a natural language processing model, its corresponding sample data can include image data, text data, and audio data.
[0069] S104: Based on the data sampling strategy corresponding to this round of training, sample data is sampled in different business datasets to obtain the sample data for this round of training.
[0070] Determine the data sampling strategy π corresponding to the t-th round of training t After that, the server can sample the sample data in different business datasets according to the probabilities of sampling different business datasets under this data sampling strategy. For any business dataset D i , ……, K, where K represents the total number of business datasets. The server can sample the sample data x, y ∼ D from it based on the probability π t (D t ) of sampling this business dataset under the data sampling strategy π. Among them, x represents the data content, and y represents the label corresponding to the data content. i i
[0071] S106: Input the sample data into the business model to determine the prediction result for the sample data through the business model.
[0072] After determining the sample data for the t-th round of training, the server can input the sample data into the business model to determine the prediction result for the sample data through the business model.
[0073] In practical applications, different sample data or business models often correspond to different prediction results. Taking image data and an image recognition model as an example, the corresponding prediction results can be image classification results or image segmentation results. For a natural language processing model, the corresponding prediction results can be semantic information, keywords, sensitive words, topic information, or target text included in the input content. For an anomaly detection model, the corresponding prediction results can be outliers or abnormal objects included in the input content.
[0074] S108: According to the prediction result, determine the loss value of the business model in this round of training, and train the business model according to the loss value of the business model in this round of training. Moreover, according to the loss value of the business model in this round of training, determine the data sampling strategy to be used for the next round of training of the business model.
[0075] After determining the prediction result corresponding to the sample data, the server can determine the loss value of the business model in the t-th round of training according to the deviation between the prediction result and the actual label corresponding to the sample data.
[0076] The server can accumulate the loss values of the training round t. To determine the gradient information and update the model parameters θ of the business model based on this gradient information, thereby performing this round of training on the business model.
[0077] Meanwhile, the server can, according to the loss value of the business model in this round of training. Determine the data sampling strategy π to be used for the next round of training of the business model. t+1 。
[0078] Specifically, for any business dataset D i , the server can, according to the loss value of the business model in the t-th round of training and the probability of sampling the business dataset D t under the data acquisition strategy π i , determine the data sampling strategy π t for the dataset D i and the reward value. Among them, the greater the loss value of the business model in this round of training , the greater the reward value of the data sampling strategy corresponding to this round of training , and vice versa.
[0079] According to the reward values of the data sampling strategy π t for each business dataset, the server can obtain the overall reward value of the data sampling strategy π t .
[0080] In this specification, the server can, in the way of moving average, determine, according to the loss value of the business model in this round of training the probability of sampling each business data under the data sampling strategy corresponding to this round of training, the preset attenuation coefficient α, and the historical reward values corresponding to the historical data sampling strategies determined before this round of training to determine the reward value of the data sampling strategy corresponding to this round of training
[0081] In practical applications, the value range of the attenuation coefficient α can be [0 - 1]. The server can, according to the attenuation coefficient α, determine the weight corresponding to the historical reward value as the first weight, and according to the first weight, determine the weight corresponding to the reward value determined based on the loss value of the business model in this round of training and the probability of sampling each business data under the data sampling strategy corresponding to this round of training as the second weight. Among them, there is a negative correlation between the first weight and the second weight.
[0082] For any business data set D i , the server can weight the historical reward value through the first weight to obtain the first weighted reward value, and weight the reward value determined based on the loss value of the business model in this round of training and the probability of sampling each business data D i under the data sampling strategy corresponding to this round of training through the second weight to obtain the second weighted reward value, and then determine, according to the first weighted reward value and the second weighted reward value, the reward value of the sampling probability for the data set D t under π i . This reward value can be expressed as:
[0083]
[0084] Among them, the first weight is equal to α, and (1 - α) is the second weight, represents the reward value determined based on the loss value of the business model in this round of training and the probability of sampling the business data set D t under the data sampling strategy π i corresponding to this round of training. It can be seen from this formula that there is a negative correlation between the first weight and the second weight, and the reward value has a positive correlation with the loss value .
[0085] Thus, the server can determine the overall reward value of the data sampling strategy π t based on the reward values of the sampling probabilities for each data set under the data sampling strategy corresponding to the t - th round of training.
[0086] It should be noted that the server may also not consider the historical reward value, but only determine the probability of sampling each piece of service data based on the loss value of the t-th round of training of the service model and the data sampling strategy π t in the following manner In this case, the value of the attenuation coefficient α is 0.
[0087] After determining the reward value corresponding to the data sampling strategy, the server can, based on this reward value, determine the data sampling strategy π to be used in the next round of training of the service model t+1 。
[0088] Specifically, the server can determine the exploration rate E for the data sampling strategy in the next round (t + 1) of training of the service model according to the quantity of the service data set and the corresponding number of training rounds in the next round of training of the service model t+1 In this specification, this exploration rate is used to represent the probability of selecting a data sampling strategy other than the optimal data sampling strategy as the data sampling strategy corresponding to the next round of training. The optimal data sampling strategy can be the data sampling strategy that maximizes the reward value.
[0089] Among them, for each round of training t, the exploration rate E corresponding to this round of training t can be expressed as:
[0090]
[0091] Among them, when the number of training rounds for training the service model exceeds the specified number of training rounds, at this time, the exploration rate of the data sampling strategy decreases as the number of training rounds increases.
[0092] Based on the above formula, the server can further determine E t+1 , and then the server can, according to the reward value of the data sampling strategy corresponding to the t-th round of training the probability of sampling each service data set under the data sampling strategy π corresponding to the t-th round of training t the exploration rate E of the data sampling strategy for the t-th round of training of the service model t and the exploration rate E of the data sampling strategy for the next round of training of the service model t+1 , determine the data sampling strategy π to be used in the next round of training of the service model t+1 。For any service data set (D i ), the probability π t+1 (D i ) of sampling this service data set under the next-round data sampling strategy can be expressed as:
[0093]
[0094] Among them, represents the set of reward values obtained from several rounds of training on the business model before. The server can determine the data sampling strategy π used in the next round of training according to the probability of sampling each business dataset under the next-round data sampling strategy. t+1 .
[0095] After that, the server can, based on the data sampling strategy π t+1 , sample the sample data in each business dataset when training the business model in the next round, and repeat the above steps to train the business model for several rounds until the training objective is met (such as reaching the preset number of training times or the model converges to the preset range).
[0096] It should be noted that for any training round t except the first round of training, the data sampling strategy can be determined and the model can be trained using the same method as above, which will not be elaborated in this specification. For the sake of understanding, this specification provides a schematic diagram of the update process of a data sampling strategy, as Figure 2 shown.
[0097] Figure 2 is a schematic diagram of the update process of a data sampling strategy provided in this specification.
[0098] Among them, in each round of training iteration, each dataset is sampled according to the current data sampling strategy. After the sample data is collected, it is input into the model to obtain the prediction result and calculate the loss value. Then, the model parameters are updated through this loss value, and the reward value of the current data sampling strategy is determined. Based on this reward value, the current data sampling strategy is updated to obtain the data sampling strategy used for the next round of training of the model.
[0099] After completing the training of the business model, the server can deploy it to execute subsequent actual operations through this model. For the sake of understanding, this specification provides a schematic diagram of the operation execution method of a business model trained in the above manner, as Figure 3 shown.
[0100] Figure 3 is a schematic diagram of the operation execution method of a business model provided in this specification, including the following steps:
[0101] S300: Obtain business data;
[0102] S302: Input the service data into a pre-trained service model to determine a prediction result for the service data through the service model, and execute a service according to the prediction result.
[0103] After the server receives a service execution request, it can obtain service data based on the service execution request, and then input it into the service model trained in the above manner, so as to execute the actual service based on the prediction result obtained by the service model.
[0104] In practical applications, different service models often correspond to different actual services. Taking an image recognition model as an example, the corresponding service data can be image data. After inputting the image data into the service model, an identification result for the image data can be obtained. Then, the server can execute services such as image classification, path planning, and unmanned device navigation based on the identification result.
[0105] For another example, for a natural language processing model, the server inputs text data into it, or inputs image data and identifies the text data contained therein, and then obtains the semantic information contained therein. The server can execute services such as intelligent customer service replies and sensitive information identification based on the semantic information.
[0106] It can be seen from the above method that this solution enables the model to reach a lower validation perplexity more quickly, reduces the number of training iterations required to achieve higher performance; improves the downstream service performance, and this solution improves the accuracy of the model by optimizing the data mixing ratio; reduces the computational cost of the model, and the computational overhead introduced during pre-training in this solution is low, ensuring the efficiency of the training process; improves the generalization ability of the model, and improves the generalization ability of the model to data in different domains by dynamically adjusting the data mixing ratio.
[0107] In this specification, the execution subject for implementing the test method of the code can refer to a specified device such as a server set up on the service platform. For the sake of convenience of description, this specification only takes the server as the execution subject as an example to illustrate a code test method provided by this specification.
[0108] The above is one or more implementation models of training and service execution methods in this specification. Based on the same idea, this specification also provides corresponding training and service execution devices, as shown in Figure 4 、 Figure 5 shown.
[0109] Figure 4 The following is a schematic diagram of a model training device provided by this specification, including:
[0110] A receiving module 400, configured to receive a training request for a service model;
[0111] A determination module 402, configured to, according to the training request, for each round of training of the service model except the first round of training, determine a data sampling strategy corresponding to this round of training based on the loss value obtained when training the service model in the previous round;
[0112] A sampling module 404, configured to sample sample data in different service data sets based on the data sampling strategy corresponding to this round of training to obtain the sample data for this round of training;
[0113] An input module 406, configured to input the sample data into the service model to determine a prediction result for the sample data through the service model;
[0114] A training module 408, configured to determine the loss value of the service model in this round of training according to the prediction result, train the service model according to the loss value of the service model in this round of training, and determine, according to the loss value of the service model in this round of training, a data sampling strategy to be used when performing the next round of training on the service model.
[0115] Optionally, the training module 408 is specifically configured to determine a reward value of the data sampling strategy corresponding to this round of training according to the loss value of the service model in this round of training and the probability of sampling each service data set under the data sampling strategy corresponding to this round of training, where the greater the loss value of the service model in this round of training, the greater the reward value of the data sampling strategy corresponding to this round of training; determine a data sampling strategy to be used when performing the next round of training on the service model according to the reward value of the data sampling strategy corresponding to this round of training.
[0116] Optionally, the training module 408 is specifically configured to determine an exploration rate for the data sampling strategy when performing the next round of training on the service model according to the number of service data sets and the training round corresponding to the next round of training of the service model; determine a data sampling strategy to be used when performing the next round of training on the service model according to the reward value of the data sampling strategy corresponding to this round of training, the probability of sampling each service data set under the data sampling strategy corresponding to this round of training, and the exploration rate for the data sampling strategy when performing the next round of training on the service model.
[0117] Optionally, the training module 408 is specifically configured to determine an exploration rate for the data sampling strategy for this round of training the service model according to the number of service data sets and the number of training rounds corresponding to training the service model; determine the data sampling strategy to be used for the next round of training the service model according to the reward value of the data sampling strategy corresponding to this round of training, the probability of sampling each service data set under the data sampling strategy corresponding to this round of training, the exploration rate for the data sampling strategy for this round of training the service model, and the exploration rate for the data sampling strategy for the next round of training the service model.
[0118] Optionally, when the number of training rounds for training the service model exceeds the specified number of training rounds, the exploration rate of the data sampling strategy decreases as the number of training rounds increases.
[0119] Optionally, the training module 408 is specifically configured to determine the reward value of the data sampling strategy corresponding to this round of training according to the loss value of the service model in this round of training, the preset attenuation coefficient of the probability of sampling each service data set under the data sampling strategy corresponding to this round of training, and the historical reward values corresponding to the respective historical data sampling strategies determined before this round of training.
[0120] Optionally, the training module 408 is specifically configured to determine the weight corresponding to the historical reward value as the first weight according to the attenuation coefficient, and determine the weight corresponding to the reward value determined based on the loss value of the service model in this round of training and the probability of sampling each service data set under the data sampling strategy corresponding to this round of training as the second weight, where there is a negative correlation between the first weight and the second weight; weight the historical reward value by the first weight to obtain a first weighted reward value, and weight the reward value determined based on the loss value of the service model in this round of training and the probability of sampling each service data set under the data sampling strategy corresponding to this round of training by the second weight to obtain a second weighted reward value; determine the reward value of the data sampling strategy corresponding to this round of training according to the first weighted reward value and the second weighted reward value.
[0121] Figure 5 The figure is a schematic diagram of a service execution device provided in this specification, including:
[0122] An acquisition module 500, configured to acquire service data;
[0123] An execution module 502, configured to input the service data into a pre-trained service model to determine a prediction result for the service data through the service model, and execute a service according to the prediction result.
[0124] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-mentioned Figure 1 model training method provided or Figure 3 business execution method provided.
[0125] This specification also provides Figure 6 a schematic structural diagram of an electronic device corresponding to Figure 1 or Figure 3 . As Figure 4 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned Figure 1 model training method or Figure 3 business execution method. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0126] In the 1990s, it was obvious to distinguish whether an improvement to a technology was an improvement in hardware (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method flows into the hardware circuits. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by a user's programming of the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0127] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0128] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0129] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0130] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0131] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks.
[0132] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks.
[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks.
[0134] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0135] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0136] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0137] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0138] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0140] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.
[0141] The above is only the embodiment of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A model training method, comprising: receiving a training request for a business model; According to the training request, for each round of training of the business model except the first round of training, based on the loss value obtained when the business model was trained in the previous round, determine the data sampling strategy corresponding to the round of training; Based on the data sampling strategy corresponding to this round of training, sample data is sampled from different business data sets to obtain sample data for this round of training; Inputting the sample data into the business model to determine a prediction result for the sample data through the business model; Based on the prediction results, determine the loss value of the business model in this round of training, and train the business model based on the loss value of the business model in this round of training; and, based on the loss value of the business model in this round of training, determine the data sampling strategy to be used in the next round of training of the business model.
2. The method according to claim 1, determining the data sampling strategy used for the next round of training of the business model according to the loss value of the business model in the current round of training, specifically comprising: Determine the reward value of the data sampling strategy corresponding to this round of training according to the loss value of the business model in this round of training and the probability of sampling each business data set under the data sampling strategy corresponding to this round of training, wherein the greater the loss value of the business model in this round of training, the greater the reward value of the data sampling strategy corresponding to this round of training; The data sampling strategy used in the next round of training for the business model is determined according to the reward value of the data sampling strategy corresponding to the current round of training.
3. The method according to claim 2, determining the data sampling strategy used for the next round of training of the business model according to the reward value of the data sampling strategy corresponding to the training round, specifically comprising: Determining an exploration rate for a data sampling strategy when performing a next round of training on the business model according to the number of business data sets and a training round corresponding to the next round of training on the business model; The data sampling strategy used for the next round of training of the business model is determined according to the reward value of the data sampling strategy corresponding to this round of training, the probability of sampling each business data set under the data sampling strategy corresponding to this round of training, and the exploration rate of the data sampling strategy when the business model is trained for the next round.
4. The method according to claim 3, determining the data sampling strategy used in the next round of training for the business model according to the reward value of the data sampling strategy corresponding to the round of training and the exploration rate of the data sampling strategy in the next round of training for the business model, specifically comprising: Determining, according to the number of the business data sets and the number of training rounds corresponding to the training of the business model, an exploration rate for a data sampling strategy when the business model is trained in this round; The data sampling strategy used for the next round of training of the business model is determined according to the reward value of the data sampling strategy corresponding to this round of training, the probability of sampling each business data set under the data sampling strategy corresponding to this round of training, the exploration rate of the data sampling strategy when the business model is trained in this round, and the exploration rate of the data sampling strategy when the business model is trained in the next round, 5. The method as claimed in claim 3, when the number of training rounds for training the business model exceeds a specified number of training rounds, the exploration rate of the data sampling strategy decreases as the number of training rounds increases.
6. The method according to claim 2, determining the reward value of the data sampling strategy corresponding to the round of training according to the loss value of the business model in the round of training and the probability of sampling each business data set under the data sampling strategy corresponding to the round of training, specifically comprises: The reward value of the data sampling strategy corresponding to this round of training is determined based on the loss value of the business model in this round of training, the preset attenuation coefficient of the probability of sampling each business data set under the data sampling strategy corresponding to this round of training, and the historical reward values corresponding to each historical data sampling strategy determined before this round of training.
7. The method according to claim 6, determining the reward value of the data sampling strategy corresponding to the round of training according to the loss value of the business model in the round of training, the preset attenuation coefficient of the probability of sampling each business data set under the data sampling strategy corresponding to the round of training, and the historical reward value corresponding to each historical data sampling strategy determined before the round of training, specifically includes: According to the attenuation coefficient, determine the weight corresponding to the historical reward value as the first weight, and according to the first weight, determine the weight corresponding to the reward value determined based on the loss value of the business model in this round of training and the probability of sampling each business data set under the data sampling strategy corresponding to this round of training as the second weight, wherein the first weight and the second weight are negatively correlated; The historical reward value is weighted by the first weight to obtain a first weighted reward value, and the reward value determined based on the loss value of the business model in the round of training and the probability of sampling each business data set under the data sampling strategy corresponding to the round of training is weighted by the second weight to obtain a second weighted reward value; A reward value for the data sampling strategy corresponding to the round of training is determined according to the first weighted reward value and the second weighted reward value.
8. A business execution method, comprising: Get business data; The business data is input into a pre-trained business model to determine a prediction result for the business data through the business model, and the business is executed according to the prediction result, wherein the business model is trained by the method described in any one of claims 1 to 7 above.
9. A model training device, comprising: A receiving module, used for receiving a training request for a business model; A determination module, configured to determine, according to the training request, for each round of training of the business model except the first round of training, a data sampling strategy corresponding to the training round based on the loss value obtained when the business model was trained in the previous round; A sampling module is used to sample sample data from different business data sets based on the data sampling strategy corresponding to this round of training to obtain sample data for this round of training; An input module, used for inputting the sample data into the business model to determine a prediction result for the sample data through the business model; A training module is used to determine the loss value of the business model in this round of training based on the prediction results, train the business model based on the loss value of the business model in this round of training, and determine the data sampling strategy used for the next round of training of the business model based on the loss value of the business model in this round of training.
10. A service execution device, comprising: Acquisition module, used to obtain business data; An execution module is used to input the business data into a pre-trained business model to determine a prediction result for the business data through the business model, and execute the business according to the prediction result, wherein the business model is trained by the method described in any one of claims 1 to 7 above.
11. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 8.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the program.