A model training method and device

By processing sparse feature parameters in model training, determining whether they meet the feature update conditions, and updating them to parameters of the next training cycle, the resource overhead and efficiency problems caused by the increase in training times are solved, and the training efficiency and effect of the model are improved.

CN114841271BActive Publication Date: 2025-05-23BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210498837.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-05-23
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

During the regular training of the model, as the number of training increases, more and more sparse features that have a smaller impact on training will be, resulting in relatively large resource overhead for model training and low training efficiency.

Method used

By obtaining the benchmark training model of the current training cycle, using the training data of the current training cycle to train the benchmark training model and its sparse feature parameters and the training parameters included, we judge whether each sparse feature parameter meets the feature update conditions based on the training results, and use the update conditions as the sparse feature parameter of the next training cycle.

Benefits of technology

The problem of greater resource overhead and lower training efficiency with the increase in training times is effectively overcome, and the efficiency and training effect of the training model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841271B_ABST
    Figure CN114841271B_ABST
Patent Text Reader

Abstract

The present invention discloses a model training method and device, and relates to the field of artificial intelligence. A specific implementation of the method includes: obtaining a benchmark training model of the current training cycle, using the determined multiple sparse feature parameters of the current training cycle combined with the training data of the current training cycle to train the benchmark training model, determining whether each sparse feature parameter meets the feature update condition according to the training result, and using the sparse feature parameters that meet the update condition as the sparse feature parameters of the next training cycle; by processing the sparse feature parameters in the model training, the problem of large resource overhead and low training efficiency caused by the increasing number of sparse features that have little impact on training as the number of training times increases is overcome; the efficiency of the training model and the training effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a model training method and device. Background Art

[0002] In prediction scenarios such as search, advertising, and recommendation (for example, predicting click-through rate), some prediction models need to be trained regularly to maintain the real-time performance of the models and improve the prediction effect.

[0003] At present, during the regular model training process, some sparse features introduced by the training model will be trained, and the sparse features introduced in the previous training will always exist in the subsequent training. As the number of training times increases, the number of sparse features with little impact on the training will increase, which leads to a relatively large resource overhead for model training and low training efficiency. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a model training method and device, which can obtain a benchmark training model of a current training cycle, use the training data of the current training cycle to train the benchmark training model and multiple sparse feature parameters contained in the benchmark training model, determine whether each sparse feature parameter meets the feature update condition according to the training result, and use the sparse feature parameters that meet the update condition as the sparse feature parameters of the training model for the next training cycle; by processing sparse feature parameters in model training, the problem of large resource overhead and low training efficiency caused by the increasing number of sparse features that have little impact on training as the number of training times increases is overcome; the efficiency of the training model and the training effect are improved.

[0005] To achieve the above-mentioned purpose, according to one aspect of an embodiment of the present invention, a model training method is provided, characterized in that it includes: obtaining a benchmark training model of a current training cycle; the benchmark training model has multiple sparse feature parameters corresponding to the current training cycle; based on the training data of the current training cycle, training the benchmark training model and training the sparse feature parameters of the benchmark training model; for each of the sparse feature parameters, judging whether the sparse feature parameter meets a preset feature update condition according to the training result, and if so, using the sparse feature parameter as the sparse feature parameter of the next training cycle; otherwise, deleting the sparse feature parameter.

[0006] Optionally, the judgment of whether the sparse feature parameters in the training results meet the preset feature update conditions includes: determining a generation timestamp corresponding to the sparse feature parameters; calculating a time difference between a current timestamp and the generation timestamp; when the time difference is less than or equal to a set time threshold, determining that the sparse feature parameters meet the preset feature update conditions; when the time difference is greater than the set time threshold, determining that the sparse feature parameters do not meet the preset feature update conditions.

[0007] Optionally, the model training method also includes: for each of the sparse feature parameters, performing the following operations: determining an initial identifier corresponding to the sparse feature parameter; when the initial identifier indicates that it is not initialized, setting a set feature parameter for the sparse feature parameter; when the initial identifier indicates that it is initialized, determining a parameter value of the sparse feature parameter; training the benchmark training model based on the training data of the current training cycle and the multiple sparse feature parameters, includes: training the benchmark training model based on the training data of the current training cycle and the set feature parameters or parameter values ​​of each of the sparse feature parameters.

[0008] Optionally, the determination of whether the sparse feature parameters in the training results meet preset feature update conditions includes: determining whether the set feature parameters or parameter values ​​of the sparse feature parameters in the training results are updated; if so, determining that the sparse feature parameters meet the preset feature update conditions when the initial identifier of the sparse feature parameters indicates that they have been initialized.

[0009] Optionally, after determining that the set feature parameters or parameter values ​​of the sparse feature parameters in the training results have been updated, the model training method further includes: performing an increment operation on the indicator value indicating the access situation preset for the sparse feature parameter; when the initial identification indication of the sparse feature parameter is uninitialized, determining whether the increment result of the indicator value indicating the access situation of the sparse feature parameter is greater than a set threshold, and if so, determining that the sparse feature parameter meets the preset feature update condition.

[0010] Optionally, the model training method, after determining that the increment result of the index value indicating the access situation of the sparse feature parameter is greater than a set threshold, further includes: updating the initial flag of the sparse feature parameter to initialized.

[0011] Optionally, the model training method trains the benchmark training model and the sparse feature parameters of the benchmark training model based on the training data of the current training cycle, including: iteratively training the benchmark training model and the sparse feature parameters of the benchmark training model based on the time sequence of real data generated within multiple set time ranges and the real data generated within multiple set time ranges.

[0012] Optionally, the model training method iteratively trains the benchmark training model and the sparse feature parameters of the benchmark training model, including: for each iteration cycle, performing the following operations: selecting real data generated within a target set time range that matches the current iteration cycle according to the time sequence of real data generated within multiple set time ranges; training the model trained in the previous iteration cycle and the sparse feature parameters of the model based on the real data generated within the target set time range that matches the current iteration cycle.

[0013] To achieve the above object, according to a second aspect of an embodiment of the present invention, a model training device is provided, characterized in that it includes: a parameter acquisition module, a model training module and an output parameter module; wherein,

[0014] The parameter acquisition module is used to acquire a benchmark training model of a current training cycle; the benchmark training model has a plurality of sparse feature parameters corresponding to the current training cycle;

[0015] The training model module is used to train the benchmark training model and the sparse feature parameters of the benchmark training model based on the training data of the current training cycle;

[0016] The output parameter module is used to determine whether each of the sparse feature parameters meets the preset feature update condition according to the training result. If yes, the sparse feature parameter is used as the sparse feature parameter of the next training cycle; otherwise, the sparse feature parameter is deleted.

[0017] Optionally, the model training device is used to determine whether the sparse feature parameters in the training results meet preset feature update conditions, including: determining a generation timestamp corresponding to the sparse feature parameters; calculating the time difference between the current timestamp and the generation timestamp; when the time difference is less than or equal to a set time threshold, determining that the sparse feature parameters meet the preset feature update conditions; when the time difference is greater than the set time threshold, determining that the sparse feature parameters do not meet the preset feature update conditions.

[0018] Optionally, the model training device is also used to perform the following operations for each of the sparse feature parameters: determining an initial identifier corresponding to the sparse feature parameter; when the initial identifier indicates that it is not initialized, setting a set feature parameter for the sparse feature parameter; when the initial identifier indicates that it is initialized, determining a parameter value of the sparse feature parameter; the training of the benchmark training model based on the training data of the current training cycle and the multiple sparse feature parameters includes: training the benchmark training model based on the training data of the current training cycle and the set feature parameters or parameter values ​​of each of the sparse feature parameters.

[0019] Optionally, the model training device is used to determine whether the sparse feature parameters in the training results meet preset feature update conditions, including: determining whether the set feature parameters or parameter values ​​of the sparse feature parameters in the training results are updated; if so, when the initial identification of the sparse feature parameters indicates that they have been initialized, determining that the sparse feature parameters meet the preset feature update conditions.

[0020] Optionally, the model training device is used to determine that after the set feature parameters or parameter values ​​of the sparse feature parameters in the training results have been updated, further includes: performing an increment operation on the indicator value indicating the access situation preset for the sparse feature parameter; when the initial identification indication of the sparse feature parameter is uninitialized, determining whether the increment result of the indicator value indicating the access situation of the sparse feature parameter is greater than a set threshold, and if so, determining that the sparse feature parameter meets the preset feature update condition.

[0021] Optionally, the model training device is used to update the initial identifier of the sparse feature parameter to initialized after determining that the increase result of the index value indicating the access situation of the sparse feature parameter is greater than a set threshold.

[0022] Optionally, the model training device is used to train the benchmark training model and the sparse feature parameters of the benchmark training model based on the training data of the current training cycle, including: iteratively training the benchmark training model and the sparse feature parameters of the benchmark training model based on the time sequence of real data generated within multiple set time ranges and the real data generated within multiple set time ranges.

[0023] Optionally, the model training device is used to iteratively train the benchmark training model and the sparse feature parameters of the benchmark training model, including: for each iteration cycle, performing the following operations: selecting real data generated within a target set time range that matches the current iteration cycle according to the time sequence of real data generated within multiple set time ranges; training the model trained in the previous iteration cycle and the sparse feature parameters of the model based on the real data generated within the target set time range that matches the current iteration cycle.

[0024] To achieve the above-mentioned purpose, according to the third aspect of an embodiment of the present invention, there is provided an electronic device for model training, characterized in that it includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement any method described in the above-mentioned model training method.

[0025] To achieve the above-mentioned purpose, according to the fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, characterized in that when the program is executed by a processor, a method as described in any one of the above-mentioned model training methods is implemented.

[0026] An embodiment of the above invention has the following advantages or beneficial effects: it is able to obtain a benchmark training model for the current training cycle, use the determined multiple sparse feature parameters of the current training cycle in combination with the training data of the current training cycle to train the benchmark training model, determine whether each sparse feature parameter meets the feature update condition based on the training result, and use the sparse feature parameters that meet the update condition as the sparse feature parameters of the next training cycle; by processing sparse feature parameters in model training, the problem of large resource overhead and low training efficiency caused by the increasing number of sparse features that have little impact on training as the number of training times increases is overcome; the efficiency of the training model and the training effect are improved.

[0027] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific implementation examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.

[0029] Figure 1 It is a flowchart of a model training method provided by an embodiment of the present invention;

[0030] Figure 2 It is a flowchart of a method for determining feature update in model training provided by an embodiment of the present invention;

[0031] Figure 3 It is a schematic diagram of a model training process provided by an embodiment of the present invention;

[0032] Figure 4 is a structural schematic diagram of a model training device provided by an embodiment of the present invention;

[0033] Figure 5 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;

[0034] Figure 6 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0035] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0036] like Figure 1 As shown, an embodiment of the present invention provides a model training method, which may include the following steps:

[0037] Step S101: Acquire a benchmark training model of a current training cycle; the benchmark training model has a plurality of sparse feature parameters corresponding to the current training cycle.

[0038] Specifically, in prediction scenarios related to CTR (Click-Through-Rate) such as search, advertising, and recommendation, prediction models (for example, DeepFM, WDL, MMoE and other models) are usually used for prediction. In order to make the model effective, it is often necessary to train the set model regularly. For example, according to the application scenario, the real data generated by the corresponding application every day (for example, the user characteristics generated by the recommended application and the real data of the recommended objects, etc.) is determined as training data to train the set model.

[0039] Get the benchmark training model of the current training cycle, where the current training cycle can be a set time range, for example: the time range of the training cycle is one day, two days, etc. When the time range of the training cycle is one day, the current training cycle can be indicated as the day before the current date. Get the benchmark training model of the current training cycle is to get the model trained the day before the current date, where the model contains multiple sparse feature parameters at the training point of the current training cycle; the benchmark training model is associated with any set model, and the benchmark training model can be any one of the training models generated in N iterative trainings.

[0040] Further, the multiple sparse feature parameters corresponding to the current training cycle of the benchmark training model are obtained; specifically, the parameters contained in the benchmark training model are divided into sparse feature parameters and dense feature parameters, and the dense feature parameters are also dense parameters. Usually, the number of dense feature parameters will not increase with the increase of training times and the scale of training data, that is, the parameter quantity of dense feature parameters is relatively fixed; while the sparse feature parameters are also sparse parameters, and the number and scale of their feature parameters will increase with the increase of training times and the scale of training data; usually, sparse feature parameters are usually fixed-length floating-point numerical arrays, where the features are obtained from the training data, for example, for model 1, the training data comes from the e-commerce mall, and the feature parameters may be user gender, user age, etc. It can be understood that the benchmark training model trained in each training cycle contains multiple sparse feature parameters associated with the model of the training cycle, and the one or more sparse feature parameters contained in the benchmark training models corresponding to the two training cycles are different; that is, the multiple sparse feature parameters corresponding to the current training cycle of the benchmark training model are obtained.

[0041] Furthermore, multiple sparse feature parameters of the current training cycle can be obtained from the parameter server (i.e., parameter pulling); parameter pulling is to first obtain one or more features of the training data, for example, the features are gender and age, and their values ​​are "male" and "20 years old", and then the feature values ​​"male" and "20 years old" are converted into feature values ​​of digital type (usually, the conversion can be performed using a hash function), and the converted feature values ​​of digital type are used as keywords to initiate a request to the parameter server, thereby obtaining the corresponding sparse feature parameters (e.g., float array) and completing the parameter pulling operation.

[0042] Step S102: Based on the training data of the current training cycle, train the benchmark training model and train the sparse feature parameters of the benchmark training model.

[0043] Specifically, the training data of the current training cycle is generally within a set time range, which can be one day, two days, N days, etc. It is understandable that the set time can be a fixed time interval or a non-fixed time interval; preferably, in order to make the benchmark training model effective, the latest real data generated every half day, every day or every two days can be used as training data; the benchmark training model is trained with the training data, and the sparse feature parameters of the benchmark training model are trained to obtain the benchmark training model of the next cycle and the sparse feature parameters contained in the benchmark training model of the next cycle. It is understandable that in addition to training the sparse feature parameters of the benchmark training model, the training of the benchmark training model also requires the training of model components such as dense parameters and loss functions contained in the model.

[0044] Furthermore, based on the time sequence of the real data generated within a plurality of set time ranges and the real data generated within a plurality of set time ranges, the benchmark training model and the sparse feature parameters of the benchmark training model are iteratively trained.

[0045] Among them, the iterative training of the benchmark training model and the sparse feature parameters of the benchmark training model include: for each iteration cycle, performing the following operations: selecting the real data generated within the target set time range that matches the current iteration cycle according to the time sequence of the real data generated within multiple set time ranges; based on the real data generated within the target set time range that matches the current iteration cycle, training the model trained in the previous iteration cycle and the sparse feature parameters of the model.

[0046] like Figure 3As shown, exemplarily, the benchmark training model 0 indicated by 301 is the initial model of the model. The method for determining the initial model can be to train the historical data corresponding to the application scenario as training data (i.e., initial cold start training), determine the output sparse feature parameters of the benchmark training model 0, and use the output sparse feature parameters of the benchmark training model 0 as the sparse feature parameters of the benchmark training model 1 corresponding to the next training cycle indicated by 302; the current iteration cycle takes the cycle for obtaining the benchmark training model 1 as an example, selects the real data (e.g., real data 1) generated within the target setting time range matching the current iteration cycle (e.g., the current iteration cycle is one day, and the target setting time range is the day before the current date), and based on the real data 1, the benchmark training model 0 is trained and the sparse feature parameters of the benchmark training model 0 to obtain the benchmark training model 1, that is, the model trained in the previous iteration cycle (benchmark training model 0) is trained to obtain a new model as the benchmark training model 1; for example: the real data 1 generated on January 1 is used to train the benchmark training model 0 on January 2 to generate the benchmark training model 1, wherein January 1 is the target setting time range matching the current iteration cycle.

[0047] Furthermore, the update of the sparse feature parameters determined in the process of training the benchmark training model 1 is used as the sparse feature parameters of the benchmark training model 1 for training in the next iteration cycle, and the benchmark training model 2 is generated based on the training results of the benchmark training model 1; that is, the benchmark training model is iteratively trained; it can be seen that the set model is iteratively trained using time sequence and real training data, and the iterative training makes the set model timely, thereby improving the accuracy of model prediction.

[0048] Step S103: for each of the sparse feature parameters, determine whether the sparse feature parameter meets the preset feature update condition according to the training result; if yes, use the sparse feature parameter as the sparse feature parameter of the next training cycle; otherwise, delete the sparse feature parameter.

[0049] Specifically, before training the benchmark training model, each sparse feature parameter contained in the benchmark training model is obtained, wherein an initial identifier is set for each sparse feature parameter; it can be understood that after the training model is completed in an iteration cycle, one or more newly added sparse feature parameters can be generated, and the initial identifier of the newly added sparse feature parameters is set to uninitialized.

[0050] Furthermore, for each of the sparse feature parameters, the following operations are performed: determining an initial identifier corresponding to the sparse feature parameter; for example: using is_init to represent the initial identifier; is_init being true represents initialized, and is_init being false represents uninitialized; when the initial identifier indicates uninitialized, setting feature parameters for the sparse feature parameter; when the initial identifier indicates initialized, determining a parameter value for the sparse feature parameter; training the benchmark training model based on the training data of the current training cycle and the multiple sparse feature parameters, includes: training the benchmark training model based on the training data of the current training cycle and the set feature parameters or parameter values ​​of each of the sparse feature parameters. Specifically, before training the current benchmark training model, the validity of the sparse feature parameters associated with the model is determined, for example, the initial identifier corresponding to the sparse feature parameters is determined. If the sparse feature parameters have been initialized, the parameter values ​​corresponding to the sparse feature parameters are obtained to train the current benchmark training model. If the sparse feature parameters have not been initialized, the feature parameters are set for the sparse feature parameters, wherein the set feature parameters can be the default values ​​of the sparse feature parameters, thereby training the benchmark training model based on the training data of the current training cycle and the multiple sparse feature parameters.

[0051] Further, it is determined whether the sparse feature parameters in the training results meet the preset feature update conditions; wherein, after the training of the benchmark training model of an iteration cycle is completed, before the output sparse feature parameters are determined based on the sparse feature parameters, one or more sparse feature parameters need to be updated; the judgment is made based on the results of the training model combined with the preset feature update conditions, that is, it is determined whether the set feature parameters or parameter values ​​of the sparse feature parameters in the training results have been updated. Specifically, during the training process, the sparse feature parameters are first pulled (Pull) to obtain the feature parameters corresponding to the current training data, and then the model is forward propagated (Forward) and backward propagated (Backward) to obtain the gradient of the feature parameters, and then the sparse feature parameters are updated (Push) The gradient update (Push) operation is performed, and the access status of the parameter is represented as being accessed once.

[0052] Furthermore, the judgment of whether the sparse feature parameters in the training results meet the preset feature update conditions includes: determining a generation timestamp corresponding to the sparse feature parameters; calculating the time difference between the current timestamp and the generation timestamp; when the time difference is less than or equal to a set time threshold, determining that the sparse feature parameters meet the preset feature influence conditions; when the time difference is greater than the set time threshold, determining that the sparse feature parameters do not meet the preset feature update conditions. Specifically, for example, for sparse feature parameter A, the timestamp stored for parameter A in the first iteration cycle is time 1; in subsequent multiple iterations of training, when parameter A is updated, timestamp 2 corresponding to the update is saved for parameter A; the time difference between timestamp 2 and timestamp 1 is judged, and if the time difference is greater than a set time threshold (for example, 7 days, 15 days, etc.); parameter A is determined to be an expired feature, so that parameter A can be removed, that is, parameter A will not be used as the sparse feature parameter used in the next iteration model training, that is, the sparse feature parameter is deleted for the next training cycle; expired sparse feature parameters with less impact are screened out by time difference, thereby overcoming the problem that there will be more and more sparse features with less impact on training, which will result in a relatively large resource overhead for model training and cause low training efficiency.

[0053] Further, it is determined whether the set feature parameters or parameter values ​​of the sparse feature parameters in the training results have been updated. If so, when the initial identification of the sparse feature parameters indicates that they have been initialized, it is determined that the sparse feature parameters meet the preset feature update conditions. That is, when it is determined that the parameter value corresponding to the sparse feature parameter needs to be updated, the initial identification of the sparse feature parameter is determined. If the initial identification indicates that it has been initialized, the parameter value of the corresponding sparse feature parameter is directly updated, otherwise, step S201-step S207 are executed.

[0054] Furthermore, if the preset feature update condition is met, the sparse feature parameters are used as the sparse feature parameters of the next training cycle; otherwise, the sparse feature parameters are deleted for the next training cycle. By judging the real-time and effectiveness of the sparse feature parameters determined for the next training cycle, the problem of increasing sparse features with little impact on training, which results in a relatively large resource overhead for model training and low training efficiency, is overcome.

[0055] like Figure 2 As shown, an embodiment of the present invention provides a method for determining feature update in model training, and the method may include the following steps:

[0056] Step S201: determining whether the set feature parameters or parameter values ​​of the sparse feature parameters in the training results are updated.

[0057] Step S202: performing a value-added operation on the index value indicating the access situation preset by the sparse feature parameter.

[0058] Specifically, when it is determined that the set feature parameters or parameter values ​​of the sparse feature parameters in the training results are updated, an increment operation is performed on the indicator value indicating the access situation preset for the sparse feature parameters; wherein, an initial identifier (is_init) and a preset indicator value indicating the access situation (for example: show) are set for each sparse feature parameter; when it is determined that the sparse feature parameter (the parameter value) needs to be updated, the indicator value indicating the access situation is increased by 1, for example: show=show+1; that is, an increment operation is performed on the indicator value indicating the access situation preset for the sparse feature parameter.

[0059] Step S203: Determine whether the initial flag of the sparse feature parameter indicates that it is not initialized. If yes, execute step S204; otherwise, execute step S205.

[0060] Step S204: Determine whether the index value increment result of the sparse feature parameter indicating the access situation is greater than a set threshold value. If yes, execute step S206; otherwise, execute step S207.

[0061] Specifically, when the initial identification indicates that it is not initialized, determine whether the index value indicating the access situation is greater than a set threshold (for example, the threshold is set to Q times); if it is greater, determine that the sparse feature parameter meets the preset feature update condition; otherwise, it does not meet.

[0062] Step S205: Determine whether the sparse feature parameters meet a preset feature update condition.

[0063] Step S206: updating the initial flag of the sparse feature parameter to initialized, and determining whether the sparse feature parameter meets a preset feature update condition.

[0064] Step S207: Determine whether the sparse feature parameters satisfy a preset feature update condition.

[0065] The description of step S201-step S207 is that after determining that the set feature parameters or parameter values ​​of the sparse feature parameters in the training results have been updated, it further includes: performing an increment operation on the indicator value indicating the access situation preset for the sparse feature parameters; when the initial identification indication of the sparse feature parameters is uninitialized, determining whether the increment result of the indicator value indicating the access situation of the sparse feature parameters is greater than a set threshold, and if so, determining that the sparse feature parameters meet the preset feature update conditions.

[0066] And, after determining that the increment result of the index value indicating the access situation of the sparse feature parameter is greater than the set threshold, it also includes: updating the initial flag of the sparse feature parameter to initialized.

[0067] It can be understood that determining that the sparse feature parameter does not satisfy the preset feature update condition means determining not to update the sparse feature parameter, and not updating the initial flag of the sparse feature parameter to be initialized, that is, the initial flag of the sparse feature parameter remains in a false state; and the sparse feature parameter whose initial flag remains false (uninitialized) will not be used as a benchmark training model for outputting sparse feature parameters to train the next iteration cycle, that is, the sparse feature parameter will be deleted.

[0068] It can be seen that by setting an initial identifier and an indicator value indicating the access situation for the sparse feature parameters, and judging the state of the initial identifier and comparing the indicator value indicating the access situation with the set threshold, the feature parameters that do not meet the update conditions can be filtered, that is, the sparse feature parameters are deleted for the next training cycle; this overcomes the problem that in the current periodic model training process, as the number of training times increases, there will be more and more sparse features that have little impact on training, which leads to a relatively large resource overhead for model training and causes low training efficiency.

[0069] like Figure 4 As shown, the embodiment of the present invention provides a model training device 400, including: a parameter acquisition module 401, a model training module 402 and an output parameter module 403; wherein,

[0070] The parameter acquisition module 401 is used to acquire a reference training model of a current training cycle; the reference training model has a plurality of sparse feature parameters corresponding to the current training cycle;

[0071] The training model module 402 is used to train the benchmark training model and the sparse feature parameters of the benchmark training model based on the training data of the current training cycle;

[0072] The output parameter module 403 is used to determine whether each of the sparse feature parameters meets the preset feature update condition according to the training result. If yes, the sparse feature parameter is used as the sparse feature parameter of the next training cycle; otherwise, the sparse feature parameter is deleted.

[0073] An embodiment of the present invention also provides an electronic device for model training, comprising: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in any of the above embodiments.

[0074] An embodiment of the present invention further provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method provided in any of the above embodiments is implemented.

[0075] Figure 5 An exemplary system architecture 500 to which the model training method or model training device according to the embodiments of the present invention can be applied is shown.

[0076] As Figure 5 shown, the system architecture 500 may include terminal devices 501, 502, 503, a network 504, and a server 505. The network 504 is used to provide a medium for communication links between the terminal devices 501, 502, 503 and the server 505. The network 504 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0077] Users can use the terminal devices 501, 502, 503 to interact with the server 505 through the network 504 to receive or send messages, etc. Various client applications may be installed on the terminal devices 501, 502, 503, such as an e-commerce client application, a web browser application, a search application, an instant messaging tool, and an email client, etc.

[0078] The terminal devices 501, 502, 503 may be various electronic devices with a display screen and supporting various client applications, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0079] The server 505 may be a server providing various services, such as a background management server that provides support for the client applications used by the users with the terminal devices 501, 502, 503. The background management server may process requests for predicting data using the prediction model and feed back the data corresponding to the prediction results to the terminal devices.

[0080] It should be noted that the model training method provided by the embodiments of the present invention is generally executed by the server 505, and correspondingly, the model training device is generally set in the server 505.

[0081] It should be understood that Figure 5 the numbers of the terminal devices, the network, and the server in

[0082] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 6 Next, reference is made to Figure 6The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0083] like Figure 6 As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0084] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage section 608 as needed.

[0085] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-mentioned functions defined in the system of the present invention are executed.

[0086] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0087] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0088] The modules and / or units involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules and / or units can also be provided in a processor. For example, it can be described as: a processor includes a parameter acquisition module, a model training module, and a parameter output module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the parameter acquisition module can also be described as "a module for acquiring the baseline training model of the current training cycle and the determined multiple sparse feature parameters of the current training cycle".

[0089] As another aspect, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: acquiring the baseline training model of the current training cycle and the determined multiple sparse feature parameters of the current training cycle; training the baseline training model based on the training data of the current training cycle and the multiple sparse feature parameters; for each of the sparse feature parameters, determining whether the sparse feature parameter in the training result meets a preset feature update condition. If so, using the sparse feature parameter as the sparse feature parameter for the next training cycle; otherwise, deleting the sparse feature parameter.

[0090] The embodiments of the present invention can acquire the baseline training model of the current training cycle, train the baseline training model by using the determined multiple sparse feature parameters of the current training cycle in combination with the training data of the current training cycle, determine whether each sparse feature parameter meets the feature update condition according to the training result, and use the sparse feature parameter that meets the update condition as the sparse feature parameter for the next training cycle. By processing the sparse feature parameters during model training, it overcomes the problem that as the number of training times increases, there will be more and more sparse features that have less impact on training, resulting in large resource consumption and low training efficiency; and improves the efficiency and effect of training the model.

[0091] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A model training method, It is characterized in that include: Get the benchmark training model for the current training cycle; The benchmark training model has a plurality of sparse feature parameters corresponding to the current training cycle; Based on the training data of the current training cycle, training the benchmark training model and training the sparse feature parameters of the benchmark training model; For each of the sparse feature parameters, judging whether the sparse feature parameter meets a preset feature update condition according to the training result, and if so, taking the sparse feature parameter as the sparse feature parameter of the next training cycle; Otherwise, delete the sparse feature parameters; Among them, judging whether the sparse feature parameters meet the preset feature update conditions based on the training results includes: determining the generation timestamp corresponding to the sparse feature parameters; calculating the time difference between the current timestamp and the generation timestamp; when the time difference is less than or equal to the set time threshold, determining that the sparse feature parameters meet the preset feature update conditions; when the time difference is greater than the set time threshold, determining that the sparse feature parameters do not meet the preset feature update conditions.

2. The method according to claim 1, It is characterized in that Also includes: For each of the sparse feature parameters, perform the following operations: Determining an initial identifier corresponding to the sparse feature parameter; In a case where the initial identification indicates that the feature parameter is not initialized, setting feature parameters for the sparse feature parameter setting; When the initial identifier indicates that the parameter value of the sparse feature parameter has been initialized, determining the parameter value of the sparse feature parameter; The step of training the benchmark training model based on the training data of the current training cycle and the plurality of sparse feature parameters includes: The benchmark training model is trained based on the training data of the current training cycle and the set feature parameters or parameter values ​​of each of the sparse feature parameters.

3. The method according to claim 2, It is characterized in that The step of judging whether the sparse feature parameters meet a preset feature update condition according to the training result includes: Determine whether the set feature parameters or parameter values ​​of the sparse feature parameters in the training results are updated. If so, when the initial identification of the sparse feature parameters indicates that they are initialized, determine that the sparse feature parameters meet the preset feature update conditions.

4. The method according to claim 3, It is characterized in that After determining that the set feature parameter or parameter value of the sparse feature parameter in the training result is updated, the method further includes: Performing a value-added operation on the index value indicating the access situation preset by the sparse feature parameter; In the case where the initial identification indication of the sparse feature parameter is uninitialized, It is determined whether the increment result of the index value indicating the access situation of the sparse feature parameter is greater than a set threshold value. If so, it is determined that the sparse feature parameter meets a preset feature update condition.

5. The method according to claim 4, It is characterized in that After determining that the increment result of the index value indicating the access situation of the sparse feature parameter is greater than the set threshold, the method further includes: The initial flag of the sparse feature parameter is updated to be initialized.

6. The method according to claim 1, It is characterized in that Training the benchmark training model based on the training data of the current training cycle, and training the sparse feature parameters of the benchmark training model, including: The benchmark training model and the sparse feature parameters of the benchmark training model are iteratively trained based on the time sequence of the real data generated within a plurality of set time ranges and the real data generated within a plurality of set time ranges.

7. The method according to claim 6, It is characterized in that Iteratively training the benchmark training model and the sparse feature parameters of the benchmark training model, including: For each iteration cycle, perform the following operations: According to the time sequence of the real data generated within the multiple set time ranges, select the real data generated within the target set time range that matches the current iteration cycle; Based on real data generated within a target set time range matching the current iteration cycle, the model trained in the previous iteration cycle and the sparse feature parameters of the model are trained.

8. A model training device, It is characterized in that include: Obtain parameter module, training model module and output parameter module; among them, The parameter acquisition module is used to acquire a benchmark training model of a current training cycle; the benchmark training model has a plurality of sparse feature parameters corresponding to the current training cycle; The training model module is used to train the benchmark training model and the sparse feature parameters of the benchmark training model based on the training data of the current training cycle; The output parameter module is used to determine, for each of the sparse feature parameters, whether the sparse feature parameter meets the preset feature update condition according to the training result. If so, the sparse feature parameter is used as the sparse feature parameter of the next training cycle; otherwise, the sparse feature parameter is deleted; wherein, the determination of whether the sparse feature parameter meets the preset feature update condition according to the training result includes: determining the generation timestamp corresponding to the sparse feature parameter; calculating the time difference between the current timestamp and the generation timestamp; when the time difference is less than or equal to the set time threshold, determining that the sparse feature parameter meets the preset feature update condition; when the time difference is greater than the set time threshold, determining that the sparse feature parameter does not meet the preset feature update condition.

9. An electronic device, It is characterized in that include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer readable medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for updating model parameters based on federated learning

    CN111931950A

  • Structured sparse parameter processing method, device and equipment and storage medium

    CN112508190A