Iterative Training Method, Iterative Training Device and Electronic Device for Hyperparameters

By using normal distribution parameters and target excitation functions in the recommendation system to adjust the hyperparameters, the problem of slow iteration speed is solved, and the iteration efficiency and the overall performance of the recommendation system are improved.

CN116361564BActive Publication Date: 2025-05-30MICRO DREAM TECHTRONIC NETWORK TECH CHINACO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310280280.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-21
Publication Date
2025-05-30
Estimated Expiration
2043-03-21

AI Technical Summary

Technical Problem

The iteration speed of hyperparameters in the prior art is slow, resulting in inefficient adjustment of hyperparameters in the sorting and fusion model of recommendation systems.

Method used

During each iterative training of hyperparameters, the normal distribution parameters corresponding to at least two material estimated scores are obtained, the initial hyperparameters are determined, and the material mixing result and excitation function value are determined by sorting the fusion model and the target excitation function, and the hyperparameters are adjusted.

Benefits of technology

It improves the speed and efficiency of hyperparameter iteration, reduces the dependence on the historical parameter adjustment experience of business personnel, and improves the overall performance of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116361564B_ABST
    Figure CN116361564B_ABST
Patent Text Reader

Abstract

The present application discloses an iterative training method, an iterative training device and an electronic device for hyperparameters. The method includes: obtaining at least two sets of first normal distribution parameters corresponding to at least two material estimated scores, and determining at least two sets of initial hyperparameters; based on the sorting fusion model using each set of initial hyperparameters respectively and at least two material estimated scores, determining at least two material mixed arrangement results, and respectively obtaining target data corresponding to each material mixed arrangement result; based on the target data corresponding to each material mixed arrangement result and the target incentive function, determining at least two first incentive function values; based on the K sets of initial hyperparameters corresponding to the first K incentive function values, determining at least two sets of second normal distribution parameters corresponding to at least two material estimated scores; in the case of determining that the iterative training of the current hyperparameters satisfies the convergence condition, taking a set of initial hyperparameters corresponding to the maximum first incentive function value obtained in the current iteration as the target hyperparameters of the sorting fusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of hyperparameter search, and particularly to an iterative training method for hyperparameters, an iterative training device, and an electronic device. Background Art

[0002] As one of the means to extract valuable content that meets the user's interests from a vast amount of information, recommendation systems have flourished in the era of the Internet information explosion. Among them, the recommendation system based on social relationships uses the user's social network in the recommendation system to perform more accurate content recommendations. The central idea is that users will be equally interested in the content that the people they follow are interested in. Therefore, there are various types of materials in the material pool of the recommendation system, such as the materials of the followed people, the materials that the followed people are interested in, and the materials in the followed interest fields. When facing the recommendation of various types of materials, it is necessary to establish a ranking fusion model based on unified value to better balance the influence of various types of materials and make the overall recommendation system achieve the optimal effect.

[0003] In the related art, in the ranking fusion model, the hyperparameters are adjusted manually based on artificial experience. After manually adjusting the hyperparameters, the online core business data indicators are observed, and then the hyperparameters are iteratively adjusted according to historical experience. There is a problem of slow iterative speed of hyperparameters. Summary of the Invention

[0004] This application discloses an iterative training method for hyperparameters, an iterative training device, and an electronic device to solve the problem of slow iterative speed of hyperparameters existing in the related art.

[0005] To solve the above problems, this application adopts the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides an iterative training method for hyperparameters, including: in each iterative training process of hyperparameters, obtaining at least two sets of first normal distribution parameters corresponding to at least two material prediction scores one by one, and determining at least two sets of initial hyperparameters based on the normal distribution corresponding to each set of first normal distribution parameters; determining at least two material mixing results corresponding to the at least two sets of initial hyperparameters one by one based on the sorting fusion model using each set of initial hyperparameters in the at least two sets of initial hyperparameters and the at least two material prediction scores, and respectively obtaining the target data corresponding to each material mixing result; determining at least two first incentive function values based on the target data corresponding to each material mixing result and the target incentive function; determining at least two sets of second normal distribution parameters corresponding to the at least two material prediction scores one by one based on the K sets of initial hyperparameters corresponding to the top K first incentive function values arranged from largest to smallest in terms of numerical value, where K is greater than 0 and less than the total number of the first incentive function values; in the case where it is determined that the iterative training of hyperparameters in the current iteration meets the convergence condition, taking a set of initial hyperparameters corresponding to the largest first incentive function value obtained in the current iteration as the target hyperparameters used by the sorting fusion model; where the convergence condition includes: the volatility between at least two first incentive function values determined in the iterative training process of hyperparameters in the current iteration and at least two first incentive function values determined in the iterative training process of hyperparameters in the previous iteration is less than a first threshold, or the volatility between the at least two sets of second normal distribution parameters and the at least two sets of first normal distribution parameters is less than a second threshold.

[0007] Second aspect, an embodiment of the present application provides an iterative training device for hyperparameters, including: a first acquisition module, configured to, in each iterative training process of hyperparameters, acquire at least two sets of first normal distribution parameters corresponding to at least two material prediction scores one by one, and determine at least two sets of initial hyperparameters based on the normal distribution corresponding to each set of first normal distribution parameters; a second acquisition module, configured to determine at least two material mixing results corresponding to the at least two sets of initial hyperparameters one by one based on the sorting fusion model using each set of initial hyperparameters in the at least two sets of initial hyperparameters and the at least two material prediction scores, and respectively acquire target data corresponding to each material mixing result; a first determination module, configured to determine at least two first incentive function values based on the target data corresponding to each material mixing result and a target incentive function; a second determination module, configured to determine at least two sets of second normal distribution parameters corresponding to the at least two material prediction scores one by one based on K sets of initial hyperparameters corresponding to the first K incentive function values arranged from large to small in value, where K is greater than 0 and less than the total number of the first incentive function values; a judgment module, configured to, when determining that the iterative training of hyperparameters in the current iteration meets the convergence condition, use a set of initial hyperparameters corresponding to the maximum first incentive function value obtained in the current iteration as the target hyperparameters used by the sorting fusion model; where the convergence condition includes: the volatility between at least two first incentive function values determined in the iterative training process of hyperparameters in the current iteration and at least two first incentive function values determined in the iterative training process of hyperparameters in the previous iteration is less than a first threshold, or the volatility between the at least two sets of second normal distribution parameters and the at least two sets of first normal distribution parameters is less than a second threshold.

[0008] Third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, it implements the steps of the method as described in the first aspect.

[0009] Fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, it implements the steps of the method as described in the first aspect.

[0010] Fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the method as described in the first aspect.

[0011] Sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium, and the program product is executed by at least one processor to implement the method as described in the first aspect.

[0012] The embodiment of the present application provides an iterative training method for hyperparameters. During each iterative training process of hyperparameters, at least two sets of first normal distribution parameters corresponding to at least two material prediction scores are obtained. Based on the normal distribution corresponding to each set of first normal distribution parameters, at least two sets of initial hyperparameters are determined. Based on the sorting fusion model using each set of initial hyperparameters in at least two sets of initial hyperparameters and at least two material prediction scores, at least two material mixed arrangement results corresponding to at least two sets of initial hyperparameters are determined, and the target data corresponding to each material mixed arrangement result is obtained respectively. Based on the target data corresponding to each material mixed arrangement result and the target incentive function, at least two first incentive function values are determined. Then, based on the K sets of initial hyperparameters corresponding to the top K first incentive function values arranged from large to small in terms of numerical value, at least two sets of second normal distribution parameters corresponding to at least two material prediction scores are determined. When it is determined that the iterative training of hyperparameters in the current iteration meets the convergence condition, a set of initial hyperparameters corresponding to the maximum first incentive function value obtained in the current iteration is used as the target hyperparameters for the sorting fusion model. In the present application, the convergence condition includes: the volatility between at least two first incentive function values determined in the iterative training process of hyperparameters in the current iteration and at least two first incentive function values determined in the iterative training process of hyperparameters in the previous iteration is less than the first threshold, or the volatility between at least two sets of second normal distribution parameters and at least two sets of first normal distribution parameters is less than the second threshold. Through the iterative training method for hyperparameters disclosed in the present application, the problem of slow iterative speed in the related art of performing hyperparameter iteration based on artificial experience can be solved, the iterative efficiency can be improved, and at the same time, business personnel do not need to have relevant historical parameter tuning experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a schematic flowchart of an iterative training method for hyperparameters disclosed in the embodiment of the present application;

[0014] Figure 2 It is a schematic diagram of a multi-type material mixed arrangement recommendation system disclosed in the embodiment of the present application;

[0015] Figure 3 It is a schematic structural diagram of an iterative training device for hyperparameters disclosed in the embodiment of the present application;

[0016] Figure 4 It is a schematic structural diagram of an electronic device disclosed in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0018] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0019] Next, in conjunction with the accompanying drawings, a hyperparameter iterative training method, an iterative training device, and an electronic device provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0020] The embodiments of the present application provide a hyperparameter iterative training method. Figure 1 It is a schematic flowchart of a hyperparameter iterative training method disclosed in the embodiments of the present application. As Figure 1 shown, the method includes the following steps.

[0021] S110. In each hyperparameter iterative training process, obtain at least two sets of first normal distribution parameters corresponding to at least two material prediction scores, and determine at least two sets of initial hyperparameters based on the normal distribution corresponding to each set of first normal distribution parameters.

[0022] It should be noted that each material prediction score corresponds to a set of first normal distribution parameters, and each material prediction score corresponds to at least two initial hyperparameters. The initial hyperparameters are obtained by sampling based on the normal distribution corresponding to a set of first normal distribution parameters corresponding to the material prediction score.

[0023] For different material estimated scores, the same number of initial hyperparameters are sampled. That is, for each material estimated score, at least two initial hyperparameters are determined based on the normal distribution corresponding to a set of first normal distribution parameters corresponding to the material estimated score. One initial hyperparameter is taken from each of the at least two initial hyperparameters corresponding to each material estimated score to form a set of initial hyperparameters (i.e., an initial hyperparameter sequence). Exemplarily, material estimated score 1 corresponds to three initial hyperparameters 2.7, 1.8, and 2.3, and material estimated score 2 corresponds to three initial hyperparameters 2.1, 2.6, and 1.2. Then three sets of initial hyperparameters are determined. The first set of initial hyperparameters includes 2.7 and 2.1, the second set of initial hyperparameters includes 1.8 and 2.6, and the third set of initial hyperparameters includes 2.3 and 1.2. The number of initial hyperparameters in each set of initial hyperparameters is equal to the number of material estimated scores.

[0024] In one implementation, a certain type of material can be input into the estimation model to obtain the material estimated score output by the estimation model. Exemplarily, as Figure 2 shown, the way to obtain the material estimated score is as follows: The first type of material is input into estimation models 1, 2, and 3 respectively to obtain the material estimated score 1 output by estimation model 1, the material estimated score 2 output by estimation model 2, and the material estimated score 3 output by estimation model 3. The second type of material is input into estimation models 4, 5, and 6 respectively to obtain the material estimated score 4 output by estimation model 4, the material estimated score 5 output by estimation model 5, and the material estimated score 6 output by estimation model 6. The third type of material is input into estimation models 7, 8, and 9 respectively to obtain the material estimated score 7 output by estimation model 7, the material estimated score 8 output by estimation model 8, and the material estimated score 9 output by estimation model 9. It should be noted that a certain type of material can be input into one estimation model to obtain a material estimated score corresponding to this type of material, or a certain type of material can be input into multiple estimation models respectively to obtain multiple material estimated scores corresponding to this type of material. The present application does not make specific limitations on this.

[0025] S120. Based on the sorting fusion model using each set of initial hyperparameters in the at least two sets of initial hyperparameters respectively and the at least two material estimated scores, determine at least two material mixing results corresponding one-to-one to the at least two sets of initial hyperparameters, and respectively obtain the target data corresponding to each material mixing result.

[0026] It should be noted that the sorting fusion model is part of the multi-type material mixing recommendation system. Exemplarily, the sorting fusion model can be Figure 2 the 210 part shown in

[0027] In the present application, based on a sorting fusion model using a set of initial hyperparameters and at least two material prediction scores, a material mixing arrangement result can be determined, that is, a set of initial hyperparameters corresponds to a material mixing arrangement result.

[0028] In one implementation, determining a material mixing arrangement result based on a sorting fusion model using a set of initial hyperparameters and at least two material prediction scores may include: inputting at least two material prediction scores into the sorting fusion model using a set of initial hyperparameters to obtain fusion domain scores of at least two material types output by the sorting fusion model, where at least two material prediction scores correspond to at least two material types, and each material type corresponds to a fusion domain score; determining a material mixing arrangement result based on the fusion domain scores of at least two material types.

[0029] Exemplarily, the first type of material corresponds to material prediction scores 1, 2, and 3, the second type of material corresponds to material prediction scores 4, 5, and 6, and a certain set of initial hyperparameters includes initial hyperparameter 1 corresponding to material prediction score 1, initial hyperparameter 2 corresponding to material prediction score 2, initial hyperparameter 3 corresponding to material prediction score 3, initial hyperparameter 4 corresponding to material prediction score 4, initial hyperparameter 5 corresponding to material prediction score 5, and initial hyperparameter 6 corresponding to material prediction score 6. By inputting material prediction scores 1, 2, 3, 4, 5, and 6 into the sorting fusion model using initial hyperparameters 1, 2, 3, 4, 5, and 6, the first fusion domain score corresponding to the first type of material and the second fusion domain score corresponding to the second type of material output by the sorting fusion model are obtained, and then a material mixing arrangement result is determined based on the first fusion domain score and the second fusion domain score. In one implementation, the first fusion domain score = material prediction score 1 * initial hyperparameter 1 + material prediction score 2 * initial hyperparameter 2 + material prediction score 3 * initial hyperparameter 3. The determination method of the second fusion domain score is similar to that of the first fusion domain score and will not be elaborated here.

[0030] After determining the mixing arrangement results of at least two materials, the target data corresponding to each material mixing arrangement result is obtained respectively, and the target data is the response data based on the material mixing arrangement result.

[0031] S130. Determine at least two first incentive function values based on the target data corresponding to each material mixing arrangement result and the target incentive function.

[0032] It should be noted that the target incentive function used in each iteration training process is the same incentive function. According to the target data corresponding to each material mixing arrangement result, a first incentive function value is determined.

[0033] S140. Determine at least two sets of second normal distribution parameters corresponding one-to-one to the at least two material prediction scores based on K sets of initial hyperparameters corresponding to the top K first activation function values arranged from large to small in value, where K is greater than 0 and less than the total number of the first activation function values.

[0034] That is to say, arrange at least two first activation function values in descending order, and determine at least two sets of second normal distribution parameters based on the K sets of initial hyperparameters corresponding to the top K first activation function values, where one material prediction score corresponds to one set of second normal distribution parameters.

[0035] S150. In the case where it is determined that the iterative training of the current hyperparameters meets the convergence condition, use a set of initial hyperparameters corresponding to the maximum first activation function value obtained in the current iteration as the target hyperparameters used by the sorting fusion model.

[0036] Among them, the convergence condition includes: the volatility between at least two first activation function values determined in the iterative training process of the current hyperparameters and at least two first activation function values determined in the iterative training process of the previous hyperparameters is less than a first threshold, or the volatility between the at least two sets of second normal distribution parameters and the at least two sets of first normal distribution parameters is less than a second threshold.

[0037] In this application, if the volatility between at least two first activation function values determined in the iterative training process of the current hyperparameters and at least two first activation function values determined in the iterative training process of the previous hyperparameters is less than the first threshold, or the volatility between at least two sets of second normal distribution parameters and at least two sets of first normal distribution parameters is less than the second threshold, it is determined that the iterative training of the current hyperparameters meets the convergence condition, and a set of initial hyperparameters corresponding to the maximum first activation function value obtained in the current iteration is used as the target hyperparameters of the sorting fusion model; if the volatility between at least two first activation function values determined in the iterative training process of the current hyperparameters and at least two first activation function values determined in the iterative training process of the previous hyperparameters is greater than or equal to the first threshold, and the volatility between at least two sets of second normal distribution parameters and at least two sets of first normal distribution parameters is greater than or equal to the second threshold, it is determined that the iterative training of the current hyperparameters does not meet the convergence condition, and return to execute the iterative training of the next hyperparameters. It should be noted that the sorting fusion model using the target hyperparameters can better balance the influence of various types of materials, and thus make the determined material mixed arrangement result better.

[0038] In one implementation manner, the convergence condition may include: the volatility between the first K first excitation function values determined in the current iteration training process of the hyperparameters and the first K first excitation function values determined in the previous iteration training process of the hyperparameters is less than a first threshold, or the volatility between at least two sets of second normal distribution parameters and at least two sets of first normal distribution parameters is less than a second threshold.

[0039] The embodiments of the present application provide an iterative training method for hyperparameters. In each iteration training process of the hyperparameters, at least two sets of first normal distribution parameters corresponding to at least two material estimation scores are obtained. Based on the normal distribution corresponding to each set of first normal distribution parameters, at least two sets of initial hyperparameters are determined. Based on the sorting fusion model using each set of initial hyperparameters in at least two sets of initial hyperparameters and at least two material estimation scores respectively, at least two material mixing results corresponding to at least two sets of initial hyperparameters are determined, and the target data corresponding to each material mixing result is obtained respectively. Based on the target data corresponding to each material mixing result and the target excitation function, at least two first excitation function values are determined. Then, based on the K sets of initial hyperparameters corresponding to the first K first excitation function values arranged from large to small in terms of values, at least two sets of second normal distribution parameters corresponding to at least two material estimation scores are determined. When it is determined that the current iteration training of the hyperparameters meets the convergence condition, a set of initial hyperparameters corresponding to the maximum first excitation function value obtained in the current iteration is used as the target hyperparameter for the sorting fusion model. In the present application, the convergence condition includes: the volatility between at least two first excitation function values determined in the current iteration training process of the hyperparameters and at least two first excitation function values determined in the previous iteration training process of the hyperparameters is less than a first threshold, or the volatility between at least two sets of second normal distribution parameters and at least two sets of first normal distribution parameters is less than a second threshold. Through the iterative training method for hyperparameters disclosed in the present application, the problem of slow iteration speed in the related art of performing hyperparameter iteration based on manual experience can be solved, the iteration efficiency can be improved, and at the same time, business personnel do not need to have relevant historical tuning experience.

[0040] In addition, through the iterative training method for hyperparameters disclosed in the present application, the parameter coverage range of sampling the initial hyperparameters based on the normal distribution is relatively large, which can improve the convergence of parameter iteration and the possibility of finding the optimal target hyperparameter.

[0041] In the embodiments of the present application, the obtaining of at least two sets of first normal distribution parameters corresponding to at least two material estimated scores may include: obtaining at least two sets of second normal distribution parameters corresponding to the at least two material estimated scores determined in the iterative training process of the previous hyperparameters; determining at least two sets of first normal distribution parameters in the iterative training process of the current hyperparameters based on the at least two sets of second normal distribution parameters in the iterative training process of the previous hyperparameters; wherein, in the iterative training process of the initial hyperparameters, the at least two sets of first normal distribution parameters are pre-configured parameters. Exemplarily, in the iterative training process of the initial hyperparameters, for each material estimated score, a set of first normal distribution parameters can be pre-configured. For example, the first mean in the pre-configured set of first normal distribution parameters is μ 0 , and the first variance is σ 0 , and then the corresponding normal distribution can be determined based on the pre-configured first normal distribution parameters.

[0042] In one implementation, each set of first normal distribution parameters includes a first mean and a first variance, and each set of second normal distribution parameters includes a second mean and a second variance. The determining of at least two sets of first normal distribution parameters in the iterative training process of the current hyperparameters based on the at least two sets of second normal distribution parameters in the iterative training process of the previous hyperparameters may include: using the at least two second means of the previous time as the at least two first means in the iterative training process of the current hyperparameters, and adding noise perturbations to the at least two second variances of the previous time as the at least two first variances in the iterative training process of the current hyperparameters. Taking one material estimated score corresponding to one set of first normal distribution parameters and one set of second normal distribution parameters as an example, if the second mean of the previous time is μ i , and the second variance is σ i , where i is an integer greater than 0, then use μ i as the first mean in the iterative training process of the current hyperparameters, that is, the first mean μ i+1 in the iterative training process of the current hyperparameters = μ i , and add the noise perturbation σ i to σ noise as the first variance in the iterative training process of the current hyperparameters, that is, the first variance σ i+1 in the iterative training process of the current hyperparameters = σ i + σ noise . By adding noise perturbations to the second variance of the previous time as the first variance of the current time, overfitting can be prevented. It should be noted that the noise perturbation σ noise is a value randomly generated by the system within a preset range.

[0043] In the embodiments of the present application, the at least two material estimated scores correspond to at least two material types, each material estimated score corresponds to one material type, and each material type corresponds to at least one material estimated score. The target data may include at least one of the following: per capita interaction volume, interaction rate, click-through rate, per capita reading duration, and the exposure proportion of each material type in the at least two material types. In the present application, the material types may include followed person materials, materials of interest to the followed person, materials in the followed interest fields, etc. Alternatively, the material types may also include followed materials, community materials, non-followed materials, etc. The present application does not make specific limitations thereto.

[0044] In the present application, before determining at least two first incentive function values based on the target data corresponding to each material mixing result and the target incentive function, it may further include: determining the target incentive function based on the target data. In the present application, the incentive function is used to evaluate the pros and cons of each group of hyperparameters. Exemplarily, when the target data includes per capita interaction volume, interaction rate, click-through rate, per capita reading duration, exposure proportion of material type 1, exposure proportion of material type 2, and exposure proportion of material type 3, the determined target incentive function reard = (per capita interaction volume + interaction rate + click-through rate + per capita reading duration) * (difference between the exposure proportion of material type 1 and the first target proportion) * (difference between the exposure proportion of material type 2 and the second target proportion) * (difference between the exposure proportion of material type 3 and the third target proportion), where the first target proportion, the second target proportion, and the third target proportion are pre-configured parameters, and the first target proportion, the second target proportion, and the third target proportion may be the same or different. The above target incentive function includes a benefit term and a constraint term. Among them, the benefit term is the main optimization target of the model, such as reading duration, interaction rate, etc., and the constraint term is used to ensure some other constraints and ensure that the exposure amounts of different types of content reach a certain proportion.

[0045] In addition, for different business scenarios, the target incentive function can be set according to the business objectives corresponding to the business scenarios, and based on the setting of the target incentive function, each parameter iteration is interpretable, greatly improving the interpretability in the hyperparameter search process.

[0046] For the hyperparameter iterative training method provided by the embodiments of the present application, the execution subject may be a hyperparameter iterative training device. In the embodiments of the present application, taking the hyperparameter iterative training device executing the hyperparameter iterative training method as an example, the hyperparameter iterative training device provided by the embodiments of the present application is described.

[0047] Figure 3 It is a structural schematic diagram of a hyperparameter iterative training device disclosed in the embodiments of the present application. As Figure 3As shown in the figure, the iterative training device 300 for hyperparameters includes: a first acquisition module 310, a second acquisition module 320, a first determination module 330, a second determination module 340, and a judgment module 350.

[0048] In this application, the first acquisition module 310 is configured to, in each iterative training process of hyperparameters, acquire at least two sets of first normal distribution parameters corresponding to at least two material prediction scores one by one, and determine at least two sets of initial hyperparameters based on the normal distribution corresponding to each set of first normal distribution parameters; the second acquisition module 320 is configured to, based on the sorting fusion model using each set of initial hyperparameters in the at least two sets of initial hyperparameters and the at least two material prediction scores, determine at least two material mixing results corresponding to the at least two sets of initial hyperparameters one by one, and respectively acquire the target data corresponding to each material mixing result; the first determination module 330 is configured to determine at least two first incentive function values based on the target data corresponding to each material mixing result and the target incentive function; the second determination module 340 is configured to determine at least two sets of second normal distribution parameters corresponding to the at least two material prediction scores one by one based on the K sets of initial hyperparameters corresponding to the top K first incentive function values arranged from largest to smallest in terms of numerical value, where K is greater than 0 and less than the total number of the first incentive function values; the judgment module 350 is configured to, when it is determined that the iterative training of hyperparameters in the current iteration meets the convergence condition, use a set of initial hyperparameters corresponding to the maximum first incentive function value obtained in the current iteration as the target hyperparameters used by the sorting fusion model; where the convergence condition includes: the volatility between at least two first incentive function values determined in the iterative training process of hyperparameters in the current iteration and at least two first incentive function values determined in the iterative training process of hyperparameters in the previous iteration is less than the first threshold, or the volatility between the at least two sets of second normal distribution parameters and the at least two sets of first normal distribution parameters is less than the second threshold.

[0049] In one implementation, the first acquisition module 310 acquiring at least two sets of first normal distribution parameters corresponding to at least two material prediction scores one by one includes: acquiring at least two sets of second normal distribution parameters corresponding to the at least two material prediction scores one by one determined in the iterative training process of hyperparameters in the previous iteration; determining at least two sets of first normal distribution parameters in the iterative training process of hyperparameters in the current iteration based on the at least two sets of second normal distribution parameters in the iterative training process of hyperparameters in the previous iteration; where, in the initial iterative training process of hyperparameters, the at least two sets of first normal distribution parameters are pre-configured parameters.

[0050] In one implementation, each set of first normal distribution parameters includes a first mean and a first variance, and each set of second normal distribution parameters includes a second mean and a second variance. The first obtaining module 310 determines at least two sets of first normal distribution parameters in the current iteration training process of the hyperparameters based on at least two sets of second normal distribution parameters in the previous iteration training process of the hyperparameters, including: using at least two previous second means as at least two first means in the current iteration training process of the hyperparameters, and adding noise perturbations to at least two previous second variances as at least two first variances in the current iteration training process of the hyperparameters.

[0051] In one implementation, the at least two material prediction scores correspond to at least two material types, each material prediction score corresponds to one material type, and each material type corresponds to at least one material prediction score. The target data includes at least one of the following: per capita interaction volume, interaction rate, click-through rate, per capita reading duration, and the exposure ratio of each material type in the at least two material types.

[0052] In one implementation, the first determination module 330 is further configured to: before determining at least two first incentive function values based on the target data corresponding to each material mixing result and the target incentive function, determine the target incentive function based on the target data.

[0053] The hyperparameter iterative training device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal.

[0054] The hyperparameter iterative training device provided in the embodiments of the present application can implement each process implemented by the hyperparameter iterative training method embodiments. To avoid repetition, it will not be elaborated here.

[0055] Optionally, as Figure 4 shown, the embodiments of the present application further provide an electronic device 400, including a processor 401 and a memory 402. A program or instruction that can run on the processor 401 is stored on the memory 402. When the program or instruction is executed by the processor 401, each step of the above hyperparameter iterative training method embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0056] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0057] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned iterative training method embodiment of the hyperparameters is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0058] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disk or optical disc, etc.

[0059] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement each process of the above-mentioned iterative training method embodiment of the hyperparameters, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0060] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system or system-on-chip, etc.

[0061] The embodiments of the present application provide a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned iterative training method embodiment of the hyperparameters, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0062] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0063] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0064] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. An iterative training method for hyperparameters, characterized in that, applied to a recommendation system, includes: In each iterative training process of hyperparameters, obtain at least two groups of first normal distribution parameters corresponding one-to-one to at least two material estimated scores, and based on the normal distribution corresponding to each group of first normal distribution parameters, determine at least two groups of initial hyperparameters, wherein the at least two material estimated scores correspond to at least two material types, each material estimated score corresponds to one material type, and each material type corresponds to at least one material estimated score; Based on the sorting fusion model using each group of initial hyperparameters in the at least two groups of initial hyperparameters and the at least two material estimated scores, determine at least two material mixed arrangement results corresponding one-to-one to the at least two groups of initial hyperparameters, and respectively obtain the target data corresponding to each material mixed arrangement result, wherein the target data includes at least one of the following: per capita interaction volume, interaction rate, click-through rate, per capita reading duration, and the exposure ratio of each material type in the at least two material types; Based on the target data corresponding to each material mixed arrangement result and the target incentive function, determine at least two first incentive function values; Based on the K groups of initial hyperparameters corresponding to the first K incentive function values arranged from large to small in value, determine at least two groups of second normal distribution parameters corresponding one-to-one to the at least two material estimated scores, wherein K is greater than 0 and less than the total number of the first incentive function values; In the case where it is determined that the iterative training of hyperparameters in the current iteration meets the convergence condition, use a group of initial hyperparameters corresponding to the maximum first incentive function value obtained in the current iteration as the target hyperparameters used by the sorting fusion model; wherein, the convergence condition includes: the volatility between at least two first incentive function values determined in the iterative training process of hyperparameters in the current iteration and at least two first incentive function values determined in the iterative training process of hyperparameters in the previous iteration is less than the first threshold, or the volatility between the at least two groups of second normal distribution parameters and the at least two groups of first normal distribution parameters is less than the second threshold.

2. The iterative training method according to claim 1, characterized in that, The obtaining of at least two groups of first normal distribution parameters corresponding one-to-one to at least two material estimated scores includes: Obtain at least two groups of second normal distribution parameters corresponding one-to-one to the at least two material estimated scores determined in the iterative training process of hyperparameters in the previous iteration; Based on the at least two groups of second normal distribution parameters in the iterative training process of hyperparameters in the previous iteration, determine at least two groups of first normal distribution parameters in the iterative training process of hyperparameters in the current iteration; wherein, in the iterative training process of hyperparameters for the first time, the at least two groups of first normal distribution parameters are pre-configured parameters.

3. The iterative training method according to claim 2, characterized in that, Each set of first normal distribution parameters includes a first mean and a first variance, and each set of second normal distribution parameters includes a second mean and a second variance. The determining of at least two sets of first normal distribution parameters in the iterative training process of the current hyperparameters based on at least two sets of second normal distribution parameters in the iterative training process of the previous hyperparameters includes: Using at least two second means of the previous time as at least two first means in the iterative training process of the current hyperparameters, and adding noise perturbations to at least two second variances of the previous time to obtain at least two first variances in the iterative training process of the current hyperparameters.

4. The iterative training method according to claim 1, characterized in that, before determining at least two first incentive function values based on the target data corresponding to each material mixing result and the target incentive function, further comprising: determining a target incentive function based on the target data.

5. An iterative training device for hyperparameters, characterized in that, applied to a recommendation system, comprising: a first acquisition module, configured to, in each iterative training process of hyperparameters, acquire at least two sets of first normal distribution parameters corresponding to at least two material prediction scores one by one, and determine at least two sets of initial hyperparameters based on the normal distribution corresponding to each set of first normal distribution parameters, wherein the at least two material prediction scores correspond to at least two material types, each material prediction score corresponds to one material type, and each material type corresponds to at least one material prediction score; a second acquisition module, configured to determine at least two material mixing results corresponding to the at least two sets of initial hyperparameters one by one based on the sorting fusion model using each set of initial hyperparameters in the at least two sets of initial hyperparameters and the at least two material prediction scores, and respectively acquire target data corresponding to each material mixing result, wherein the target data includes at least one of the following: per capita interaction volume, interaction rate, click-through rate, per capita reading duration, and the exposure ratio of each material type in the at least two material types; a first determination module, configured to determine at least two first incentive function values based on the target data corresponding to each material mixing result and the target incentive function; a second determination module, configured to determine at least two sets of second normal distribution parameters corresponding to the at least two material prediction scores one by one based on K sets of initial hyperparameters corresponding to the first K incentive function values arranged from large to small in value, where K is greater than 0 and less than the total number of the first incentive function values; a judgment module, configured to, when determining that the iterative training of the current hyperparameters meets the convergence condition, use a set of initial hyperparameters corresponding to the maximum first incentive function value obtained in the current time as the target hyperparameters used by the sorting fusion model; wherein the convergence condition includes: the volatility between at least two first incentive function values determined in the iterative training process of the current hyperparameters and at least two first incentive function values determined in the iterative training process of the previous hyperparameters is less than a first threshold, or the volatility between the at least two sets of second normal distribution parameters and the at least two sets of first normal distribution parameters is less than a second threshold.

6. The iterative training device according to claim 5, wherein, the first acquisition module acquires at least two sets of first normal distribution parameters corresponding one-to-one to at least two material prediction scores, including: acquiring at least two sets of second normal distribution parameters corresponding one-to-one to the at least two material prediction scores determined during the iterative training process of the previous hyperparameters; determining at least two sets of first normal distribution parameters during the iterative training process of the current hyperparameters based on the at least two sets of second normal distribution parameters during the iterative training process of the previous hyperparameters; wherein, during the iterative training process of the initial hyperparameters, the at least two sets of first normal distribution parameters are pre-configured parameters.

7. The iterative training device according to claim 6, wherein, each set of first normal distribution parameters includes a first mean and a first variance, each set of second normal distribution parameters includes a second mean and a second variance, and the first acquisition module determines at least two sets of first normal distribution parameters during the iterative training process of the current hyperparameters based on the at least two sets of second normal distribution parameters during the iterative training process of the previous hyperparameters, including: using the at least two second means of the previous time as the at least two first means during the iterative training process of the current hyperparameters, and using the at least two second variances of the previous time after adding noise perturbations as the at least two first variances during the iterative training process of the current hyperparameters.

8. An electronic device, wherein, it includes a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the hyperparameter iterative training method according to any one of claims 1-4 are implemented.

9. A readable storage medium, wherein, a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the hyperparameter iterative training method according to any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • Model training method, related system and storage medium

    CN113407820A

  • Hyper-parameter determination method and device of deep learning model, equipment and storage medium

    CN114418091A