Hyper-parameter processing method and related device

By obtaining and processing multiple sets containing hyperparameter value combinations and feedback values, and calculating and selecting the best combination, the problem of low computing resource consumption and optimization efficiency in traditional methods is solved, and more efficient hyperparameter tuning is achieved.

CN120012880APending Publication Date: 2025-05-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311547607.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In machine learning, when traditional Bayesian parameter optimizers are used for hyperparameter tuning, they consume a lot of computing resources, and as the explored hyperparameter combination increases, the optimization efficiency decreases.

Method used

By obtaining a plurality of first sets, each set contains a feedback value and a hyperparameter value combination, the second set is determined based on these combinations, and the evaluation score of each combination is calculated, and the target value combination is selected for model tuning.

Benefits of technology

This method saves computing resources when adjusting hyperparameters, improves optimization efficiency, and can find the best hyperparameter combination more quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012880A_ABST
    Figure CN120012880A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a hyper-parameter processing method and a related device, which can be applied to scenes such as artificial intelligence, cloud technology and smart traffic and are used for optimizing and adjusting hyper-parameters, saving computing resources and improving optimization efficiency. The method comprises the steps that N first sets are acquired, each first set comprises a first feedback value and a first value combination, the first value combination comprises parameter values of a plurality of hyper-parameters, and each first feedback value is used for indicating the characterization capacity of a neural network model under the corresponding first value combination; when N meets a preset condition, P second sets are determined based on each first value combination and each first feedback value; based on each first value combination, each first feedback value, each second value combination and each second feedback value, calculating an evaluation score corresponding to the second value combination; and selecting a target value combination based on the evaluation scores of all the second value combinations so as to carry out model optimization processing on the neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method for processing hyperparameters and related devices. Background Art

[0002] Machine learning is one of the hottest research directions in the field of artificial intelligence. In machine learning, there are usually two types of parameters involved, hyperparameters and normal parameters. Among them, hyperparameters are operating parameters set before starting learning, rather than normal parameters obtained through model training. Hyperparameters define high-level concepts about machine learning models, such as complexity or learning ability, and the setting of hyperparameters directly affects the performance of the model. Therefore, hyperparameters need to be repeatedly tested and adjusted to achieve the best performance.

[0003] In the traditional scheme of adjusting hyperparameters, it is usually implemented based on Bayesian parameter optimizer. However, in the process of hyperparameter optimization by Bayesian parameter optimizer, more computing resources are needed to complete the distribution fitting process of parameter value combinations, and as the number of parameter value combinations of the explored hyperparameters increases, not only more computing resources are needed, but also the optimization efficiency is poor. Summary of the invention

[0004] The embodiments of the present application provide a method for processing hyperparameters and related devices, which can be used to save computing resources and improve optimization efficiency during the process of adjusting hyperparameters.

[0005] In the first aspect, an embodiment of the present application provides a method for processing hyperparameters. The method includes obtaining N first sets, each of which includes a first feedback value and a first value combination, the first value combination includes parameter values ​​of multiple hyperparameters, each of which is used to indicate the representation capability of the neural network model under the corresponding first value combination, and N is an integer greater than or equal to 2; when the N satisfies the preset condition, P second sets are determined based on each of the first value combination and each of the first feedback values, each of which includes a second feedback value and a second value combination of multiple hyperparameters, and the second feedback value is used to indicate the predictive representation capability of the neural network model under the corresponding second value combination, and P is an integer greater than or equal to 2; based on each of the first value combination and each of the first feedback values, and each of the second value combination and each of the second feedback values, the evaluation score corresponding to the second value combination is calculated; based on the evaluation scores of all the second value combinations, a target value combination is selected from the P second value combinations for model tuning of the neural network model.

[0006] In a second aspect, an embodiment of the present application provides a hyperparameter processing device. The hyperparameter processing device includes an acquisition unit and a processing unit. The acquisition unit is used to acquire N first sets, each of which includes a first feedback value and a first value combination, the first value combination includes parameter values ​​of multiple hyperparameters, each of which is used to indicate the representation ability of the neural network model under the corresponding first value combination, and N is an integer greater than or equal to 2; the processing unit is used to determine P second sets based on each first value combination and each first feedback value when N meets a preset condition, each of which includes a second feedback value and a second value combination of multiple hyperparameters, and the second feedback value is used to indicate the predictive representation ability of the neural network model under the corresponding second value combination, and P is an integer greater than or equal to 2; the processing unit is used to calculate the evaluation score corresponding to the second value combination based on each first value combination and each first feedback value, and each second value combination and each second feedback value; the processing unit is used to select a target value combination from the P second value combinations based on the evaluation scores of all the second value combinations, so as to perform model tuning processing on the neural network model.

[0007] In some optional embodiments, the processing unit is used to: when the N is greater than the first threshold, calculate the value difference between the third set and each of the fourth sets in the Q fourth sets, and obtain the distance vector between the third set and each of the fourth sets, wherein the Q fourth sets are Q first sets selected from the N first sets, and the third set is selected from the Q fourth sets based on the first feedback value in each of the fourth sets, 1≤Q<N, and Q is an integer; calculate the expected gradient descent change based on the distance vector between all the third sets and the fourth sets; perform norm solving processing on the parameter range of each of the hyperparameters, and the largest first feedback value and the smallest first feedback value in the Q fourth sets to obtain the first diagonal vector modulus; determine the second set based on the third set, the expected gradient descent change, the first diagonal vector modulus, the preset action rate and the preset attenuation coefficient.

[0008] In some other optional embodiments, the processing unit is used to: perform modulus processing on each of the Q first distance vectors to obtain the modulus of the corresponding first distance vector, each first distance vector is a distance vector between the third set and the corresponding fourth set; solve the quotient between each first distance vector and the modulus of the corresponding first distance vector to determine the unit vector corresponding to the first distance vector; sum the unit vectors of the Q first distance vectors to obtain the expected gradient descent change.

[0009] In other optional embodiments, the processing unit is used to: calculate the unit vector of the expected gradient descent change based on the expected gradient descent change; multiply the unit vector of the expected gradient descent change, the preset action rate, the preset attenuation coefficient and the first diagonal vector modulus to obtain an expected gradient value; and sum the third set and the expected gradient value to obtain a second set.

[0010] In other optional embodiments, the processing unit is used to: calculate the parameter interval length corresponding to the hyperparameter based on the parameter range of each of the hyperparameters, and perform norm solution processing on the parameter interval lengths of all the hyperparameters to obtain a first value; calculate the difference between the maximum first feedback value and the minimum first feedback value corresponding to Q of the fourth sets to obtain a second value; perform norm solution processing on the first value and the second value to obtain a first diagonal vector modulus.

[0011] In other optional embodiments, the processing unit is used to: perform modulus processing on the expected gradient descent change to obtain the modulus of the expected gradient descent change; solve the quotient between the expected gradient descent change and the modulus of the expected gradient descent change to obtain the unit vector of the expected gradient descent change.

[0012] In some other optional embodiments, the processing unit is used to: sum the first value combination in the third set and the expected gradient value to obtain a second value combination; sum the first feedback value in the third set and the expected gradient value to obtain a second feedback value; and obtain a second set based on the second value combination and the second feedback value.

[0013] In other optional embodiments, the processing unit is used to: calculate the gradient evaluation score of the third value combination based on each of the first feedback value and the third feedback value in the Q fourth sets, the third feedback value being the second feedback value in any one of the second sets, and the third value combination being the second value combination in the second set corresponding to the third feedback value; calculate the distance evaluation score of the third value combination based on the third value combination, each first value combination in the fourth sets, the parameter range of each of the hyperparameters, and the preset parameter adjustment coefficient; sum the gradient evaluation score of the third value combination and the distance evaluation score of the third value combination to obtain the evaluation score of the third value combination.

[0014] In some other optional embodiments, the processing unit is used to: respectively calculate the first difference value between each of the first feedback values ​​and the third feedback value in the Q fourth sets, and select a target difference value from the Q first difference values, wherein the target difference value is the minimum value among the Q first difference values; calculate the second difference value between the largest first feedback value and the smallest first feedback value in the Q fourth sets; calculate the quotient between the target difference value and the second difference value to obtain a gradient evaluation score of a third value combination.

[0015] In other optional embodiments, the processing unit is used to: calculate the parameter interval length corresponding to the hyperparameter based on the parameter range of each of the hyperparameters, and perform norm solution processing on the parameter interval lengths of all the hyperparameters to obtain a second diagonal vector modulus; respectively calculate the first similarity distance between the third value combination and each first value combination in the fourth set, and determine the minimum first similarity distance from the Q first similarity distances; calculate the distance evaluation score of the third value combination based on the second diagonal vector modulus, the minimum first similarity distance and the preset adjustment coefficient.

[0016] In other optional embodiments, the processing unit is used to: when N is less than or equal to a first threshold, process each first value combination and each first feedback value in the N first sets through a random generator to generate P second sets; calculate the parameter interval length corresponding to the hyperparameter based on the parameter range of each hyperparameter, and perform norm solution processing on the parameter interval lengths of all the hyperparameters to obtain a second diagonal vector modulus; respectively calculate the second similarity distance between the fourth value combination and each of the first value combinations, and determine the minimum second similarity distance from the second similarity distances, the fourth value combination being any second value combination in the second set; calculate the distance evaluation score based on the second diagonal vector modulus, the minimum second similarity distance and the preset adjustment coefficient to obtain an evaluation score corresponding to the fourth value combination.

[0017] In some other optional implementations, the processing unit is used to select the second value combination corresponding to the maximum evaluation score as the target value combination.

[0018] In other optional embodiments, the processing unit is also used for: after selecting the second value combination corresponding to the largest evaluation score as the target value combination, when the parameter value of each of the hyperparameters in the target value combination does not satisfy the preset default value of the corresponding hyperparameter, calculating the target default value of the corresponding hyperparameter based on the parameter range and preset adjustment precision of each hyperparameter; and adjusting the parameter value of each of the hyperparameters to the target default value of the corresponding hyperparameter.

[0019] In other optional embodiments, the processing unit is also used to: after selecting a target value combination from P second value combinations based on the evaluation scores of all the second value combinations, respectively calculate the third similarity distance between the target value combination and each first value combination in the N first sets; when each of the third similarity distances is less than a second threshold, perform model tuning processing on the neural network model based on the target value combination.

[0020] The third aspect of the embodiment of the present application provides a hyperparameter processing device, including: a memory, an input / output (I / O) interface and a memory. The memory is used to store program instructions. The processor is used to execute the program instructions in the memory to perform the hyperparameter processing method corresponding to the implementation method of the first aspect above.

[0021] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer is enabled to execute a method corresponding to the implementation method of the first aspect described above.

[0022] The fifth aspect of the embodiments of the present application provides a computer program product containing instructions, which, when executed on a computer or a processor, enables the computer or the processor to execute the method corresponding to the implementation method of the first aspect.

[0023] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0024] In the embodiment of the present application, N first sets are obtained, and each first set includes a first feedback value and a first value combination, where N is an integer greater than or equal to 2. The described first value combination includes parameter values ​​of multiple hyperparameters, and each first feedback value can indicate the characterization capability of the neural network model under the corresponding first value combination. After obtaining N first sets, by judging whether N meets the preset conditions, and then when N meets the preset conditions, P second sets are determined based on each first value combination and each first feedback value, where P is an integer greater than or equal to 2. In each second set, a second feedback value and a second value combination of multiple hyperparameters are included. The described second feedback value is used to indicate the predictive characterization capability of the neural network model under the corresponding second value combination. In this way, the evaluation score of the corresponding second value combination is calculated based on each first value combination, each first feedback value, and each second value combination and each second feedback value, and then the target value combination is selected from the P second value combinations based on the evaluation scores of all second value combinations, so that the target value combination is used to perform model tuning processing on the neural network model. That is to say, in the embodiment of the present application, when the number of first value combinations of known feedback values ​​meets the preset conditions, the candidate second value combinations and the corresponding second feedback values ​​are predicted directly based on the known first feedback values ​​and the corresponding first value combinations, and then after evaluating the evaluation scores of each second value combination, the target value combination is selected from the candidate second value combinations based on the evaluation scores. Compared with the traditional Bayesian parameter tuning method, the present application does not need to use more computing resources to complete the selection of parameter value combinations required for distribution fitting, which not only saves computing resources, but also selects the optimal target value combination from the perspective of the evaluation scores of the candidate second value combinations, greatly improving the optimization efficiency of the subsequent model optimization process. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1 A schematic diagram of a model processing framework provided in an embodiment of the present application is shown;

[0027] Figure 2 A flow chart of a method for processing hyperparameters provided in an embodiment of the present application is shown;

[0028] Figure 3 A schematic diagram of a process for determining a second set provided in an embodiment of the present application is shown;

[0029] Figure 4 A schematic diagram of the calculation process of the evaluation score provided in the embodiment of the present application is shown;

[0030] Figure 5 The overall process diagram of the method for processing hyperparameters provided in the embodiment of the present application is shown;

[0031] Figure 6 A schematic diagram showing an application scenario of the method for hyperparameter processing provided in an embodiment of the present application;

[0032] Figure 7 A schematic diagram of the functional module structure of the hyperparameter processing device provided in an embodiment of the present application is shown;

[0033] Figure 8 A schematic diagram of the hardware structure of the hyperparameter processing device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0034] The embodiments of the present application provide a method for processing hyperparameters and related devices, which can be used to save computing resources and improve optimization efficiency during the process of adjusting hyperparameters.

[0035] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0036] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0037] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the implementation of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] With the research and progress of artificial intelligence (AI) technology, AI technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, driverless cars, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, AI technology will be applied in more fields and play an increasingly important role.

[0039] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, pre-trained model technology, big data processing technology, operation / interaction systems, mechatronics and other technologies. Among them, the pre-trained model is also called a large model or a basic model. After fine-tuning, it can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech technology, natural language processing technology, and machine learning / deep learning.

[0040] In the field of artificial intelligence, different neural network models such as prediction models and pre-trained models are usually relied on to process pending tasks. In different neural network models, different hyperparameters are also set so that the corresponding neural network models can have different characterization capabilities. The described characterization capabilities can be understood as the model prediction capabilities or model generalization capabilities of the neural network model. It should be noted that the neural network model described in the embodiments of the present application may include but is not limited to linear regression models, logistic regression models, ridge regression models, lasso regression models, support vector machines, regression tree models, prediction models, pre-trained models, etc., which are not specifically limited in this application.

[0041] Take the field of machine learning in artificial intelligence as an example. In this field of machine learning, the hyperparameters mentioned above are understood as parameters used to control the neural network model. They are different from the parameters automatically learned by training the neural network model through training data, which usually need to be configured manually. Whether the hyperparameters can be set reasonably will directly affect the model performance of the entire neural network model.

[0042] In traditional hyperparameter tuning schemes, it is usually implemented based on Bayesian parameter optimizer. However, this traditional tuning scheme requires more computing resources to complete the fitting process, and as the number of parameter value combinations of the explored hyperparameters increases, the computing resources consumed also gradually increase, and the optimization efficiency is low.

[0043] Therefore, in order to solve the above-mentioned technical problems, the embodiment of the present application provides a method for processing hyperparameters. The method for processing hyperparameters can be applied to scenarios such as cloud technology, artificial intelligence, smart transportation, assisted driving, big data, etc., and the present application does not make specific limitations.

[0044] Exemplarily, the method for hyperparameter processing provided in the present application can be applied to hyperparameter processing devices with data processing capabilities, such as terminal devices, servers, question-and-answer robots, etc. Among them, terminal devices may include but are not limited to smart phones, desktop computers, laptops, tablet computers, smart speakers, vehicle-mounted devices, smart watches, wearable smart devices, intelligent voice interaction devices, smart home appliances, aircraft, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (context delivery networks, CDN), and big data and artificial intelligence platforms, etc., and this application does not make specific restrictions. In addition, the terminal devices and servers mentioned can be directly connected or indirectly connected by wired communication or wireless communication, etc., and this application does not make specific restrictions.

[0045] Figure 1 A schematic diagram of a model processing framework provided in an embodiment of the present application is shown.

[0046] like Figure 1 As shown, in the model processing framework, business data is first obtained and business indicators of the business data are counted. After the business indicators are obtained, the hyperparameters in the neural network model are used to control the process of preprocessing and feature extraction of the business indicators. Furthermore, a neural network model such as a multi-layer perception mechanism with hyperparameters set is also used to predict the task revenue obtained after the task to be processed goes online. Finally, the prediction ability of the neural network model under the current hyperparameters is evaluated, and the prediction evaluation score is fed back and output to adjust the configuration of the hyperparameters through the corresponding prediction evaluation score.

[0047] Need to explain, for Figure 1 The prediction evaluation score fed back may be the performance on the total cross-validation validation set in the modeling phase; or, it may be the performance feedback on the test set; or, it may be the effect feedback obtained after the model is put on the system and a product such as a virtual game is put on the actual market, which is not limited in the specific embodiments of the present application.

[0048] Regarding the above Figure 1How the present application implements the process of tuning the hyperparameters is shown in the figure. The following introduces a method for hyperparameter processing provided in an embodiment of the present application in conjunction with the accompanying drawings. Figure 2 A flow chart of the method for processing hyperparameters provided in an embodiment of the present application is shown. Figure 2 As shown, the method for processing hyperparameters may include the following steps:

[0049] 201. Obtain N first sets, each of which includes a first feedback value and a first value combination, the first value combination includes parameter values ​​of multiple hyperparameters, each first feedback value is used to indicate the representation ability of the neural network model under the corresponding first value combination, and N is an integer greater than or equal to 2.

[0050] In this example, the hyperparameters may include some parameters for controlling the neural network model to preprocess the input business data, and may also include some parameters related to the multi-layer perceptron structure and fitting ability control processing. Alternatively, the hyperparameters mentioned may also include, but are not limited to, learning rate, regularization coefficient, network structure parameters (such as the number of nodes or activation function, etc.), data batch size, etc., which are not limited in the specific embodiments of this application.

[0051] As an illustrative description, the above parameters used to control the neural network model to pre-process the input business data include but are not limited to F parameters, V parameters or C parameters, etc., which are not limited in this application.

[0052] The F parameter mentioned above is used to indicate the filling method of missing values, which is an integer type. For example, when the F parameter value is 0, it indicates zero value filling; when the F parameter value is 1, it indicates minimum value filling; when the F parameter value is 2, it indicates mean and virtual variable value filling, etc. Exemplarily, in some examples, the preset default value of the F parameter can be configured in advance to be 0, 1 or 2, so that after the target value combination is determined later, the value of the corresponding hyperparameter in the target value combination is adjusted based on the preset default value, which plays a role in further improving the optimization efficiency and optimization effect.

[0053] It should be noted that the value of the F parameter may include other values ​​in addition to the above-mentioned 0, 1 or 2 in practical applications, which is not limited in this application. In addition, for zero value filling, minimum value filling, mean value and virtual variable value filling, in addition to using 0, 1 or 2 to represent them respectively, other values ​​or symbols may also be used in practical applications to represent them, which is not limited in this application.

[0054] The V parameter mentioned above is used to represent the version number of the selected feature combination, which is an integer type. The feature combination mentioned may be composed of multiple features. For example, features may include but are not limited to video features, text features or image features, voice features or features in other scenarios, etc. Taking virtual games as an example, the features mentioned above may include but are not limited to video features of game videos in short video platforms, game live broadcast features in live broadcast platforms, or game features of other virtual games. For example, game features may include but are not limited to revenue within the first 30 days after the virtual game goes online, the number of reservations for the virtual game, the number of comments, the number of positive reviews, etc., which are not specifically limited in this application. The video features mentioned may include but are not limited to the viewing time, the number of barrages, the number of collections, the number of bloggers, or the number of video submissions in the first T days before the video goes online, which are not specifically limited. The game live broadcast characteristics mentioned may include but are not limited to the number of live broadcast platforms, the number of anchors per day, the number of viewers, and the number of times popular anchors enter the top 1 of the daily list T days before each live broadcast platform goes online in the game live broadcast; or the number of live broadcast platforms, the number of anchors per day, the number of viewers, and the number of times popular anchors enter the top 1 of the daily list T days before all live broadcast platforms go online, which is not limited in the specific embodiments of the present application.

[0055] It should be noted that the V parameter mentioned above may include other values ​​besides the version numbers of the various feature combinations mentioned above, which are not limited in this application.

[0056] The described C parameter is used to represent the coding feature corresponding to the category attribute, which is an integer type. Exemplarily, in the process of using the C parameter to represent the coding feature corresponding to the category attribute, the corresponding category attribute can also be determined by the number of samples of the input business data, and then the samples with a small number of samples are merged by controlling the merging degree threshold of the long-tail value to determine the corresponding category attribute. In other words, when the number of samples of the input business data is less than a certain threshold, it is uniformly marked as Tiny type, and the original category value of the sample is ignored; and when the number of samples is greater than or equal to the threshold, the original category attribute of the sample is used. For example, assuming that the total number of samples of the input business data is 100, and the threshold is 50, if 90 of the category attributes are attribute A, and each of the remaining 10 samples belongs to a different category, the category attribute of the remaining 10 samples can be defined as attribute B. For example, taking the game scene as an example, the C parameter includes but is not limited to values ​​such as C1 value to C5 value, which is not specifically limited in this application. Among them, the C1 value is used to indicate the developer to which the game sample belongs; the C2 value is used to indicate the game publisher; the C3 value and C4 value represent the game play type and game type respectively; and the C5 value represents the game commercial revenue model.

[0057] It should be noted that, in addition to the C1 to C5 values ​​mentioned above, the C parameters mentioned above may also include other values ​​in practical applications, which are not limited in this application.

[0058] In addition, some parameters mentioned above involving the multi-layer perceptron structure and the fitting ability control processing include but are not limited to the number of nodes in each layer of the perceptron structure, or the number of nodes and the regularization term coefficient. For example, taking the multi-layer perceptron structure including a 3-layer structure as an example, where the 3rd layer is the output layer, the corresponding hyperparameters may include but are not limited to N1, N2, L1 to L3. Among them, N1 represents the number of nodes of the aforementioned feature data input into the multi-layer perceptron structure flowing to the 1st layer of the perceptron; N2 represents the number of nodes of the 2nd layer of the perceptron; L1 represents the L1 regularization term coefficient and L2 regularization term coefficient of the node weight value and offset of the 1st layer of the perceptron; L2 represents the L1 regularization term coefficient and L2 regularization term coefficient of the node weight value and offset of the 2nd layer of the perceptron; L3 represents the L1 regularization term coefficient and L2 regularization term coefficient of the node weight value and offset of the 3rd layer of the perceptron, and the specific application is not limited.

[0059] It should be noted that the hyperparameters mentioned in this application may be other hyperparameters in practical applications in addition to the F parameters, V parameters, C parameters, N1, N2, L1, L2 and L3 parameters mentioned above, which are not specifically limited in this application.

[0060] After setting the reasonable parameter range of the default values ​​of each hyperparameter mentioned above in advance, different hyperparameter parameter values ​​can be used to combine, so that the neural network model can iteratively predict the task to be processed under different hyperparameter value combinations to feedback the characterization capabilities under different hyperparameter value combinations. Thus, different hyperparameter value combinations and corresponding feedback values ​​can be constructed to obtain corresponding sets. For example, after N iterations, N first sets can be constructed using the value combinations corresponding to all hyperparameters used in each iteration and the feedback values ​​obtained after the iteration, where N is an integer greater than or equal to 2. That is, in each first set, a first feedback value and a first value combination are included. The first value combination in each first set includes parameter values ​​of multiple hyperparameters. The first feedback value in each first set can reflect the characterization capability of the neural network model under the corresponding first value combination.

[0061] As a schematic description, a vector can be used to represent the N first sets mentioned above, namely G i =[p i (x 1i ,x 2i ,...,x ki ),v i ], i∈[1,N], i is an integer. Among them, Gi represents the i-th first set among N first sets, p i (x 1i ,x 2i ,...,x ki ) represents the first value combination in the i-th first set, and k represents the first value combination p i The number of hyperparameters included in x, k ≥ 1, and k is an integer, ki Then it means that the first value combination p i The parameter value of the kth hyperparameter in . In addition, for v i , then it is understood as the first set G i The first feedback value in indicates that the neural network model is in the first value set p i The ability to represent below.

[0062] For example, when i=1, then G1=[p1(x 11 ,x 21 ,...,x k1 ),v1] represents the first set corresponding to the first iteration when N iterations are performed. Among them, p1(x 11 ,x 21 ,...,x k1 ) represents the first value combination in the first set G1, x 11 represents the parameter value of the first hyperparameter in the first value combination p1, and so on. k1 Indicates the parameter value of the kth hyperparameter in the first value combination p1. In addition, v1 in the first set G1 indicates the first feedback value of the neural network model under the first value combination p1.

[0063] It should be noted that, for the case where i takes other values, the corresponding first set, first value combination and first feedback value can all be understood by referring to the content corresponding to the aforementioned i=1, and will not be elaborated here.

[0064] Optionally, in the multidimensional parameter space, each first value combination p i It can also be expressed as an exploration point, that is, the coordinate position in the multidimensional parameter space. In this way, the value combinations of all hyperparameters in different iterations can be quickly located in the multidimensional parameter space.

[0065] Exemplarily, after constructing the N first sets, the N first sets can be stored in a database or other storage medium. In this way, in the subsequent hyperparameter tuning process, the N first sets can be first obtained from the database or other storage medium, for example.

[0066] 202. When N meets the preset conditions, P second sets are determined based on each first value combination and each first feedback value, each second set includes a second feedback value and a second value combination of multiple hyperparameters, each second feedback value is used to indicate the predictive representation ability of the neural network model under the corresponding second value combination, and P is an integer greater than or equal to 2.

[0067] In this example, when N satisfies different preset conditions, each first combination and each first feedback value can be processed in different ways to determine P second sets, where P is an integer and P≥2. In each second set, a second feedback value and a second value combination are included. Among them, the second value combination in each second set is used to reflect the expected parameter values ​​of multiple hyperparameters. In addition, the second feedback value in each second set is used to indicate the predictive characterization capability of the neural network model under the corresponding second value combination. Exemplarily, the second set mentioned above can also be referred to as a candidate set, that is, it is understood to be a candidate set obtained by prediction. The set name of the second set is not limited in the specific embodiments of the present application.

[0068] As a schematic description, a vector can be used to represent the P second sets mentioned above, namely S c =[p' c (x 1c ,x 2c ,...,x kc ),v' c ]=(p' c ,v' c ), c∈[1,P], c is an integer. Among them, S c represents the cth second set among P second sets, p' c (x 1c ,x 2c ,...,x kc ) represents the second value combination in the cth second set, x kc Then it means that the second value combination p' c The parameter value of the kth hyperparameter in . In addition, for v' c , then it can be understood as the second set S c The second feedback value in indicates that the neural network model is in the second value set p' c The ability to represent below.

[0069] For example, when c = 2, then S2 = [p'2(x 12 , x 22 ,...,x k2 ), v'2] represents the second set determined when the second iteration process is performed on all the first value combinations and all the first feedback values.12 , x 22 ,...,x k2 ) represents the second value combination in the second set S2, x 12 represents the parameter value of the first hyperparameter in the second value combination p'2, and so on. k2 Indicates the parameter value of the kth hyperparameter in the second value combination p'2. In addition, v'2 in the second set S2 indicates the second feedback value of the neural network model under the second value combination p'2.

[0070] The above-mentioned N satisfies different preset conditions, mainly including the case where N is greater than the first threshold, or the case where N is less than or equal to the first threshold. It should be noted that the above-mentioned first threshold can be determined according to user needs, for example, including but not limited to 5, 10, 15.5, etc., and this application does not limit it.

[0071] For example, when N is greater than the first threshold, the virtual gradient descent method shown in the subsequent method 1 can be used to determine the P second sets. When N is less than or equal to the first threshold, the random generation method shown in the subsequent method 2 can be used to determine the P second sets. For details, please refer to the contents described in the following methods 1 and 2, that is:

[0072] Method 1: When N is greater than the first threshold

[0073] Exemplarily, when N is greater than the first threshold, a virtual gradient descent method may be used to determine the P second sets. Specifically, Figure 3 FIG. 2 shows a schematic diagram of a process for determining a second set provided in an embodiment of the present application. Figure 3 As shown, in this embodiment, at least the following steps 301 to 304 are executed, namely:

[0074] 301. When N is greater than the first threshold, calculate the value difference between the third set and each fourth set in the Q fourth sets, and obtain the distance vector between the third set and each fourth set, wherein the Q fourth sets are Q first sets selected from the N first sets, and the third set is selected from the Q fourth sets based on the first feedback value in each fourth set, 1≤Q<N, and Q is an integer.

[0075] In this example, when N is greater than the first threshold, Q fourth sets may be selected from the N first sets obtained in step 201. For example, a roulette selector may be used to select Q fourth sets from the N first sets.

[0076] The Q described has a value range of 1≤Q<N, and Q is an integer. As a schematic description, the Q fourth sets can be represented by a vector: G" q =[p" q (x 1q ,x 2q ,...,x kq ),v" q ]=(p" q ,v" q ), q∈[1,Q], q is an integer. Among them, G" q represents the qth fourth set among Q fourth sets, p" q (x 1q ,x 2q ,...,x kq ) represents the qth fourth set G" q The first value combination in v" q Then it means the qth fourth set G" q The first feedback value in indicates that the neural network model is in the first value set p" q The ability to represent below.

[0077] It should be noted that the Q fourth sets selected in the present application are still the first sets in essence. In the present application, different names are used to clearly describe and calculate the correlation between sets, and this limitation is not made in practical applications. In addition, the method of selecting Q fourth sets from N first sets can be a random selection method; or; based on the first feedback value in each first set, the first set corresponding to the first feedback value being less than a preset selection threshold is selected as the fourth set; or, the first set corresponding to the first feedback value being greater than or equal to a preset selection threshold is selected as the fourth set, etc. In the embodiment of the present application, the method of selecting Q fourth sets from N first sets is not limited.

[0078] After selecting Q fourth sets, it is also possible to select each fourth set G" in the Q fourth sets. q =[p" q (x 1q ,x 2q ,...,x kq ),v" q ]=(p" q ,v" q ) in the first feedback value v" q , select the third set from these Q fourth sets. As an illustrative description, v"1 to v" q The fourth set corresponding to the minimum value in is selected as the third set. For example, the third set is represented by G m =[pm (x 1m , x 2m ,...,x km ), v m ]=(p m ,v m ). Among them, v m For v"1 to v" q The minimum value among them, m∈[1,q], m is an integer, p m (x 1m , x 2m ,...,x km ) is represented by v m The corresponding first value combination.

[0079] For example, assuming Q = 4, the corresponding four fourth sets are G"1 = (p"1, v"1), G"2 = (p"2, v"2), G"3 = (p"3, v"3), G"4 = (p"4, v"4). If the values ​​of v"1 to v"4 are 0.5, 0.6, 0.7, and 0.8 respectively, the fourth set G"1 = (p"1, v"1) corresponding to the minimum value v"1 = 0.5 can be determined as the third set G m , and v"1=0.5 in G"1=(p"1,v"1) is determined as v m , and p"1 in G"1=(p"1,v"1) is determined as p m .

[0080] In some examples, the third set G mentioned above m The first value combination p in m (x 1m , x 2m ,...,x km ), sometimes also called the first subcombination, then the corresponding v m Then represents the first feedback value corresponding to the first sub-combination. Similarly, each fourth set G" mentioned above q The first value combination p" q , which is sometimes also called the second subcombination, then the corresponding v" q It represents the first feedback value corresponding to the second sub-combination.

[0081] Thus, after obtaining the third set and Q fourth sets, the value differences between the third set and each fourth set are calculated to obtain the distance vector between the third set and each fourth set. For example, the distance vector between the third set and each fourth set satisfies the formula arx q .in,

[0082] arx q =Gm -G" q =(p m ,v m )-(p" q ,v" q )

[0083] =[x 1m , x 2m ,...,x km , v m ]-[x 1q , x 2q ,...,x kq ,v" q ].

[0084] It should be noted that in the above formula arx q In this case, q≠m. And, arx q Denotes the third set G m and the qth fourth set G" q The distance vector between them. Through this distance vector arx q , we can know the qth fourth set G" q The first value combination p" q , points to the third set G in the multidimensional parameter space m The first value combination p in m direction and distance.

[0085] 302. Calculate an expected gradient descent change based on all distance vectors between the third set and the fourth set.

[0086] In this example, according to the content of step 301 above, the distance vector between the third set and each fourth set can be calculated, that is, Q first distance vectors are calculated. It should be noted that each first distance vector is understood as the distance vector between the third set and any fourth set. In this way, after calculating the distance vectors between the third set and each fourth set respectively, the expected gradient descent change can be calculated based on the distance vectors between all the third sets and the fourth sets. Through the expected gradient descent change, it is possible to know the direction of the first value combination in the fourth set to which it is necessary to go, and perform gradient attenuation according to the calculated expected gradient. It should be noted that the expected gradient descent change provided in the present application is a vector with size and direction.

[0087] As an illustrative description, in the process of calculating the expected gradient descent change, each of the Q first distance vectors described above can be modulo processed to obtain the modulus of the corresponding first distance vector, for example ||arx q||. In this way, the quotient between each first distance vector and the modulus of the corresponding first distance vector is solved to determine the unit vector corresponding to the first distance vector, for example After obtaining the unit vectors of the Q first distance vectors, the unit vectors of the Q first distance vectors are summed to obtain the expected gradient descent change, for example: in, and Specifically, the names of the expected gradient descent changes are not limited in the embodiments of the present application, and other names may be used to represent them in practical applications.

[0088] 303. Perform norm solving processing on the parameter range of each of the hyperparameters, and the largest first feedback value and the smallest first feedback value in the Q fourth sets to obtain a first diagonal vector norm.

[0089] In this example, in addition to executing step 302 to calculate the expected gradient descent change, in the process of determining the second set, it is also necessary to calculate the diagonal vector modulus of the entire k-dimensional parameter space and the 1-dimensional feedback dimension (i.e., k+1-dimensional space). By way of example, the parameter range of each hyperparameter, and the largest first feedback value and the smallest feedback value in the Q fourth sets can be norm-solved to obtain the first diagonal vector modulus. In other words, the first diagonal vector modulus is understood as the diagonal vector modulus of the parameter range of the k+1-dimensional space composed of the parameter range of the entire k-dimensional parameter space and the known range of the 1-dimensional feedback dimension.

[0090] As a schematic description, in the process of calculating the first diagonal vector modulus, the parameter interval length of the corresponding hyperparameter can be calculated based on the parameter range of each hyperparameter, for example (R j -L j ). Among them, R j represents the maximum default value of the jth hyperparameter among the k hyperparameters, L j Indicates the minimum default value of the jth hyperparameter. For example, taking the aforementioned F parameter as a hyperparameter of any one dimension in the value combination, its maximum default value is 2 and its minimum default value is 0.

[0091] In this way, after calculating the parameter interval lengths of all hyperparameters, the parameter interval lengths of all hyperparameters are processed by norm solution to obtain the first value, for example Here, n is understood as the value in the n-order norm algorithm. For example, when n=1, it represents the 1-norm, when n=2, it represents the 2-norm, etc., and this application does not make specific limitations. In addition, the minimum first feedback value and the maximum first feedback value can be determined from the Q fourth sets, and then the difference between the maximum first feedback value and the minimum first feedback value is calculated to obtain the second value, for example (max v" q -min v" q ).

[0092] In this way, after obtaining the first value and the second value, the norm of the first value and the second value is solved to obtain the first diagonal vector norm. For example: Where ||A|| represents the magnitude of the first diagonal vector.

[0093] 304. Determine a second set based on the third set, the expected gradient descent change, the first diagonal vector magnitude, the preset action rate, and the preset attenuation coefficient.

[0094] In this example, the preset action rate is sometimes also referred to as the exploration step size. For example, the preset action rate can be represented by the symbol O r The value of 0 can include but is not limited to 0.05, etc., which can be determined according to user needs and is not limited in this application. The preset attenuation coefficient described can be understood as the overall attenuation of the iteration steps as the iteration steps move. For example, the preset attenuation coefficient can be represented by the symbol 0 a Indicates that its value may include but is not limited to 0.96, etc., which may depend on user needs and is not limited in this application.

[0095] Thus, in the process of determining the second set, the expected gradient descent change amount can be modulo-calculated to obtain the modulus of the expected gradient descent change amount, for example: The quotient between the expected gradient descent change and the modulus of the expected gradient descent change is further solved, and the unit vector of the expected gradient descent change is calculated, such as After calculating the unit vector of the expected gradient descent change, the unit vector of the expected gradient descent change, the preset action rate, the preset attenuation coefficient, and the first diagonal vector modulus are multiplied to obtain the expected gradient value. For example, the expected gradient value satisfies the formula Grad, where It should be noted that Grad represents the expected gradient value, O r Indicates the default action rate, O a represents the preset attenuation coefficient. In this way, the third set and the expected gradient value are summed to obtain the second set, for example

[0096] As an illustrative description, the first value combination in the third set can be summed with the expected gradient value to obtain a second value combination, for example, p' c =p m +Grad; Similarly, the first feedback value in the third set is summed with the expected gradient value to obtain the second feedback value, for example, v' c =v m +Grad. Thus, based on the second value combination and the second feedback value, a second set is constructed, that is, (p' c ,v' c )=(p m +Grad,v m +Grad).

[0097] Repeat the above steps 301 to 304 for P times, and then obtain P second sets. It should be noted that for each second set, the details can be understood by referring to the contents described in the above step 202.

[0098] By using the virtual gradient descent method in the above method 1 to determine the most preferred candidate second value combination in each iteration process, it is possible to automatically adjust the values ​​of hyperparameters through iteration, greatly improving the optimization efficiency of subsequent parameter tuning.

[0099] Method 2: When N is less than or equal to the first threshold

[0100] Exemplarily, when N is less than or equal to the first threshold, each first value combination and each first feedback value in the N first sets can be processed with the help of a random generator to generate P second sets. It should be noted that the expression form of the P second sets described here can be understood with reference to the P second sets in the aforementioned method 1, and will not be repeated here. In addition, in addition to the above-mentioned methods 1 and 2 for generating P second sets, other methods may also be included in practical applications, which are not specifically limited in this application.

[0101] In the above manner, when it is known that the number of first value combinations and first feedback values ​​in the N first sets is small, by randomly generating P second sets, the exploration process of complex candidate value combinations can be simplified and the optimization efficiency can be improved.

[0102] 203. Based on each first value combination and each first feedback value, and each second value combination and each second feedback value, calculate an evaluation score corresponding to the second value combination.

[0103] In this example, the evaluation score described can be used to evaluate the model feedback effect brought about by the neural network model under the corresponding parameter value combination; or, the evaluation score can also be understood as the score of the representation ability of the neural network model. Therefore, after determining all the first value combinations, the first feedback values, and all the second value combinations, the second feedback values, the evaluation score corresponding to the second value combination can also be calculated based on each first value combination and each first feedback value, and each second value combination and each second feedback value.

[0104] From the above-mentioned method 1 and method 2 in step 202, it can be seen that the difference between the two methods is that in method 1, not only the parameter value but also the gradient needs to be considered; while in method 2, only the parameter value needs to be considered. Therefore, for the situation shown in the above-mentioned method 1, in the process of calculating the evaluation score of each second value combination, it is necessary to comprehensively consider the evaluation of the parameter value and the evaluation of the gradient, while for the situation shown in the above-mentioned method 2, it is only necessary to consider the evaluation of the parameter value. This application will be described separately from different embodiments, as follows:

[0105] Method 1: When N is greater than the first threshold

[0106] For example, for the situation shown in the method 1 in the aforementioned step 202, in the process of calculating the evaluation score of each second value combination, the present application takes any second value combination in the second set (that is, the third value combination shown later) as an example, Figure 4 The figure shows a schematic diagram of the calculation process of the evaluation score provided in the embodiment of the present application. Figure 4 As shown, the calculation process of the evaluation score at least includes executing the following steps 401 to 403 to calculate the evaluation score of the third value combination. Please refer to the following steps for specific understanding, namely:

[0107] 401: Calculate a gradient evaluation score of a third value combination based on the third feedback value and each first feedback value in the Q fourth sets, where the third feedback value is a second feedback value in any one of the second sets, and the third value combination is a second value combination in the second set corresponding to the third feedback value.

[0108] In this example, the third feedback value corresponds to the third value combination. For example, any second set among the P second sets, for example (p' c ,v' c ) as an example, the second value combination p' c As the third value combination, the third feedback value is v' c In this way, based on the third feedback value and the selected Q fourth sets (for example (p" q ,v" q)), calculate the gradient evaluation score of the third value combination.

[0109] It should be noted that the Q fourth sets described here can refer to the aforementioned Figure 3 The contents of the fourth set described in step 301 can be understood and will not be described in detail here.

[0110] As an illustrative description, in the process of calculating the gradient evaluation score of the third value combination, the first difference value between each first feedback value and the third feedback value in the Q fourth sets can be calculated respectively, thereby obtaining Q first difference values. Further, a target difference value is selected from the Q first difference values, for example, the smallest first difference value is selected from the Q first difference values ​​as the target difference value, for example, the target difference value is min(v" q -v' c ), where v" q -v' c is the first difference between the first feedback value and the third feedback value in any fourth set. In addition, it is also necessary to determine the maximum first feedback value and the minimum first feedback value from the Q fourth sets, and then calculate the second difference between the maximum first feedback value and the minimum first feedback value in the Q fourth sets, for example, the second difference value is maxv" q -minv" q Finally, the quotient between the target difference value and the second difference value is calculated to obtain the gradient evaluation score of the third value combination, for example Among them, SC g Represents the gradient evaluation score of the third value combination.

[0111] For example, when Q=3, the corresponding three fourth sets are: G"1=(p"1,v"1)=[p"1(x 11 ,x 21 ,...,x k1 ),0.5)], G"2=(p"2,v"2)=[p"2(x 12 ,x 22 ,...,x k2 ),0.6)], G"3=(p"3,v"3)=[p"3(x 13 ,x 23 ,...,x k3 ),0.7)]. If the fourth second set among the P second sets (ie when c=4) is taken as an example, the third value combination is p'4=p'4(x 14 ,x 24 ,...,x k4), the corresponding third feedback value is v'4=0.4. At this time, the first difference values ​​between v"1, v"2, v"3 and v'4 can be calculated, that is, 0.1, 0.2, and 0.3, respectively, and then the smallest first difference value (that is, 0.1) is selected as the target difference value. In addition, the largest first feedback value (that is, 0.7) and the smallest first feedback value (that is, 0.5) can be selected from the calculations of v1, v2, and v3, and the calculated second difference value is 0.7-0.5=0.2. Further, the third value combination p'4=p'4(x 14 ,x 24 ,...,x k4 ) has a gradient evaluation score of

[0112] It should be noted that the calculation process of the gradient evaluation scores of other second value combinations can be understood by referring to the gradient evaluation scores of the third value combination, and will not be described in detail here.

[0113] 402. Calculate a distance evaluation score based on the third value combination, the first value combination in each fourth set, the parameter range of each hyperparameter, and the preset parameter adjustment coefficient.

[0114] In this example, for the third value combination, the distance evaluation score of the third value combination can also be calculated based on the third value combination, the first value combination in each fourth set, the parameter range of each hyperparameter, and the preset parameter adjustment coefficient. It should be noted that the third value combination described here can be understood with reference to the content described in the aforementioned step 401, and will not be repeated here.

[0115] As an illustrative description, in the process of calculating the distance evaluation score of the third value combination, it is also necessary to calculate the diagonal vector norm of the entire k-dimensional parameter space, for example, based on the parameter range of each hyperparameter, the parameter interval length of the corresponding hyperparameter is calculated, and the parameter interval lengths of all hyperparameters are normed to obtain the second diagonal vector norm, for example in, represents the second diagonal vector magnitude, R j -L j Represents the parameter interval length of the jth hyperparameter. It should be noted that how to calculate the parameter interval length of each hyperparameter here can be understood by referring to the aforementioned calculation of the parameter interval length of the hyperparameter in the first diagonal vector norm, which will not be repeated here.

[0116] In addition, it is also necessary to calculate the first similarity distance between the third value combination and the first value combination in each fourth set, for example ||p' c -p" q||. It should be noted that how to calculate the first similarity distance here can be to use a similarity algorithm such as a cosine similarity algorithm and a Euclidean distance to calculate the similarity distance between the third value combination and the first value combination in each fourth set. The specific embodiment of the present application will not be described in detail. In this way, after obtaining Q first similarity distances, the smallest first similarity distance is selected from the Q first similarity distances, for example, min||p' c -p" q ||.

[0117] Finally, based on the calculated second diagonal vector modulus, the minimum first similarity distance and the preset adjustment coefficient, the distance evaluation score of the third value combination is calculated. For example, the distance evaluation score of the third value combination satisfies the formula SC d ,in w d represents the preset adjustment coefficient, and N is the number of the first set.

[0118] For example, taking the example shown in step 401 above as an example, assuming that p'4=p'4(x 14 ,x 24 ,...,x k4 ) and p"1(x 11 ,x 21 ,...,x k1 )、p"2(x 12 ,x 22 ,...,x k2 )、p"3(x 13 ,x 23 ,...,x k3 ) are 0.4, 0.45, and 0.6 respectively, from which it can be determined that the minimum first similarity distance is 0.4. If the second diagonal vector modulus N = 10, w d =0.5, then the third value combination p'4=p'4(x 14 ,x 24 ,...,x k4 ) has a distance evaluation score of

[0119] It should be noted that the calculation process of the distance evaluation scores of other second value combinations can be understood by referring to the distance evaluation score of the third value combination, and will not be elaborated here.

[0120] 403. Sum the gradient evaluation score of the third value combination and the distance evaluation score of the third value combination to obtain an evaluation score of the third value combination.

[0121] In this example, for the third value combination, after calculating the gradient evaluation score and the distance evaluation score of the third value combination, the sum of the gradient evaluation score and the distance evaluation score of the third value combination is calculated to obtain the evaluation score of the third value combination.

[0122] For example, taking the example shown in step 401 above as an example, the third value combination p'4=p'4(x 14 ,x 24 ,...,x k4 ) has a gradient evaluation score and a distance evaluation score of 0.5 and 3.33 respectively. After summing, the third value combination p'4=p'4(x 14 ,x 24 ,...,x k4 ) has an evaluation score of 3.83.

[0123] In summary, the calculation process of the evaluation scores of other second value combinations can be understood by referring to the evaluation scores of the third value combination, and will not be elaborated here.

[0124] Method 2: When N is less than or equal to the first threshold

[0125] For example, when N is less than or equal to the first threshold, only the parameter value needs to be considered in determining the P second sets, and the expected gradient is not considered. Therefore, in this case, the evaluation score of each second value combination can also be reflected only from the distance evaluation score between the value combinations.

[0126] Taking any second value combination in the second set (i.e., the fourth value combination mentioned later) as an example, in the process of calculating the distance evaluation score of the fourth value combination, it is necessary to calculate the diagonal vector modulus of the entire k-dimensional parameter space, for example, based on the parameter range of each hyperparameter, the parameter interval length of the corresponding hyperparameter is calculated, and the parameter interval lengths of all hyperparameters are normed to obtain the second diagonal vector modulus, for example in, represents the second diagonal vector magnitude, R j -L j Represents the parameter interval length of the jth hyperparameter. It should be noted that how to calculate the parameter interval length of each hyperparameter here can be understood by referring to the aforementioned calculation of the parameter interval length of the hyperparameter in the first diagonal vector norm, which will not be repeated here.

[0127] In addition, it is also necessary to calculate the second similarity distance between the fourth value combination and the first value combination in each first set, for example ||p' c -p" q||. It should be noted that how to calculate the second similarity distance here can be to use a distance similarity algorithm such as a cosine similarity algorithm and a Euclidean distance to calculate the similarity distance between the fourth value combination and the first value combination in each first set. The specific embodiment of the present application will not be described in detail. In this way, after obtaining Q second similarity distances, the smallest second similarity distance is selected from the Q second similarity distances, for example, min||p' c -p" q ||.

[0128] Finally, based on the calculated second diagonal vector modulus, the minimum second similarity distance and the preset adjustment coefficient, the distance evaluation score of the fourth value combination is calculated. For example, the distance evaluation score of the fourth value combination satisfies the formula SC d ,in w d represents the preset adjustment coefficient, and N is the number of the first set.

[0129] It should be noted that the calculation process of the distance evaluation scores of other second value combinations can be understood by referring to the distance evaluation scores of the fourth value combination, and will not be described in detail here. Figure 4 The distance evaluation score of the third value combination described in step 402 can be understood through the content, which will not be repeated here.

[0130] It should be noted that, in practical applications, other methods besides the above-mentioned method 1 and method 2 may be used to calculate the evaluation score of each second value combination, which is not specifically limited in this application.

[0131] 204. Select a target value combination from the P second value combinations based on the evaluation scores of all the second value combinations, so as to perform model tuning processing on the neural network model.

[0132] In this example, after calculating the evaluation score of each second value combination in the P second sets, a target value combination can be selected from the evaluation scores of the P second value combinations. For example, the second value combination corresponding to the largest evaluation score can be selected as the target value combination. For example, in the case of P=3, the corresponding second value combinations p'1 to p'3 have evaluation scores of 50 points, 60 points, and 80 points, respectively, so that the target selection combination can be determined to be the value combination corresponding to 80 points, that is, the second value combination p'3. In this way, after determining the target value combination, the target value combination can be used to perform model tuning on the neural network model.

[0133] Optionally, in other examples, maximizing the evaluation score can be used as an example, and an optimization problem can be constructed for the evaluation score, such as min-A. Where A is the evaluation score of each second value combination. For example, in the case of P=3, the evaluation scores of the corresponding second value combinations p'1 to p'3 are 50 points, 60 points, and 80 points, respectively, so min-A=-80 can be determined, and then the target selection combination is determined to be p'3.

[0134] Optionally, in the above Figure 2 , Figure 3 or Figure 4 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, after determining the target value combination, it is also possible to determine whether the parameter value of each hyperparameter in the target value combination satisfies the preset default value of the corresponding hyperparameter. Figure 2 The contents described in step 201 in the above description can be understood, and will not be repeated here. In this way, when it is determined that the parameter value of any one or more hyperparameters does not meet the preset default value of the corresponding hyperparameter, the target default value of the corresponding hyperparameter can also be calculated based on the parameter range and preset adjustment accuracy of each hyperparameter by, for example, a POI Transformer converter. Further, after calculating the target default value of each hyperparameter, the parameter value of each hyperparameter is adjusted to the target default value of the corresponding hyperparameter.

[0135] The target default value of the hyperparameter satisfies the following formula: j >0, Among them, x ji represents the target default value of the jth hyperparameter, L j represents the minimum default value of the jth hyperparameter, R j represents the maximum default value of the jth hyperparameter, D j represents the preset adjustment accuracy of the jth hyperparameter, j∈[1, k], j is an integer. Or, when D j = 0, the default value of the hyperparameter satisfies: ji ∈[L j ,R j ].

[0136] It should be noted that the preset adjustment accuracy D mentioned above j Can be any positive real number, not limited to 10 m The precision is equal to the power of m, and the value of m may include but is not limited to 0.01, 0.1, 1, etc., which is not limited in this application.

[0137] Optionally, in the above Figure 2 , Figure 3 or Figure 4 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, after selecting the target value combination from P second value combinations according to the evaluation scores of all second value combinations, the third similarity distance between the target value combination and each first value combination in the N first sets can be calculated respectively. Further, it is determined whether each third similarity distance is less than the second threshold. It should be noted that each third similarity distance is less than the second threshold, indicating that the target value combination is not the same as or similar to each first value combination in the existing N sets. Based on this, when each third similarity distance is less than the second threshold, the neural network model is subjected to model tuning processing based on the target value combination.

[0138] On the contrary, if any of the third similarity distances is greater than the second threshold, it means that the target value combination already exists in the N sets. At this time, using the target value combination to perform model tuning on the neural network model will not bring about a significant performance change. At this time, you only need to use the original value combination for processing.

[0139] It should be noted that the second threshold mentioned above can be determined according to user needs, for example, including but not limited to 50, 10.5, etc., and this application does not limit it.

[0140] The embodiment of the present application mainly includes but is not limited to satisfying one or more of the following conditions as to when to stop the model tuning process of the neural network model. For example, condition 1: the current iteration step number is greater than or equal to the preset total iteration step number; condition 2: the evaluation score of the second value combination is greater than 0.6; condition 3: the decrease rate of the feedback value of M consecutive iterations is less than a preset threshold, where M is an integer greater than or equal to 2.

[0141] It should be noted that in practical applications, other conditions for stopping the adjustment of hyperparameters may also be included, which are not specifically limited in the embodiments of the present application.

[0142] Figure 5 The overall flow chart of the method for processing hyperparameters provided in the embodiment of the present application is shown. Figure 5 As shown, first obtain N first sets, and determine whether N is greater than a first threshold. Each first set described here can refer to the aforementioned Figure 2 The contents described in step 201 are understood.

[0143] When it is determined that N is greater than the first threshold, Q fourth sets are selected from the N first sets, and a third set is selected from the Q fourth sets based on the first feedback value in each fourth set. In this way, the value differences between the third set and each fourth set in the Q fourth sets are calculated to obtain the distance vectors between the third set and each fourth set. Furthermore, based on the distance vectors between all third sets and fourth sets, the expected gradient descent change is calculated. Similarly, it is also necessary to perform norm solving processing on the parameter range of each hyperparameter, and the largest first feedback value and the smallest first feedback value in the Q fourth sets to obtain the first diagonal vector modulus. In this way, the second set is determined based on the third set, the expected gradient descent change, the first diagonal vector modulus, the preset action rate, and the preset attenuation coefficient. Finally, the iteration is repeated P times to determine P second sets. For details, please refer to the aforementioned Figure 3 The contents described in step 301 to step 304 can be understood for reference only and will not be elaborated here.

[0144] In this way, when N is greater than the first threshold, the evaluation score of the second value combination in each second set (i.e., the sum of the gradient evaluation score and the distance evaluation score) can also be calculated, and then the second value combination corresponding to the largest evaluation score can be selected as the target value combination.

[0145] On the contrary, when N is less than or equal to the first threshold, P second sets are randomly generated based on the random generator, and the evaluation score (i.e., distance evaluation score) of the second value combination in each second set is calculated, and then the second value combination corresponding to the largest evaluation score is selected as the target value combination.

[0146] In this way, after determining the target value combination, it is possible to determine whether the parameter value of each hyperparameter in the target value combination satisfies the preset default value of the corresponding hyperparameter. In the case where the parameter value of each hyperparameter in the target value combination does not satisfy the preset default value of the corresponding hyperparameter, the target default value of the corresponding hyperparameter is calculated based on the parameter range and preset adjustment accuracy of each hyperparameter, and the parameter value of each hyperparameter is adjusted to the target default value of the corresponding hyperparameter. Conversely, in the case where the parameter value of each hyperparameter in the target value combination satisfies the preset default value of the corresponding hyperparameter, there is no need to adjust the parameter value, and the parameter value of the corresponding hyperparameter in the target value combination can be used.

[0147] After adjusting the parameter values ​​of the hyperparameters, the similarity distance between the target value combination and each first value combination in the N first sets can also be calculated, and when each similarity distance is less than the second threshold, the neural network model is tuned based on the target value combination.

[0148] On the contrary, when any similarity distance is greater than or equal to the second threshold, the neural network model uses the first value combination when the similarity distance is greater than or equal to the second threshold to successively execute the aforementioned Figure 1 The process of preprocessing and feature extraction of business data, the process of predicting the task benefits obtained after the task to be processed is launched, and the process of feeding back and outputting the predicted evaluation score are shown in order to enter the model tuning process.

[0149] Furthermore, during the model tuning process, it is determined whether the model tuning reaches an iteration stop condition, and if the iteration stop condition is reached, the model tuning process is stopped.

[0150] Need to explain, for Figure 5 For example, the first set, the second set, the value combination, the gradient evaluation score, etc. shown in Figures 2 to 4 You can understand the contents described in the article, and will not go into details here.

[0151] Figure 6 A schematic diagram showing an application scenario of the method for processing hyperparameters provided in an embodiment of the present application is shown. Figure 6 As shown, taking the application of the hyperparameter processing method provided in the present application in a virtual game scene as an example, after using the hyperparameter processing method of the present application, it is possible to optimize the hyperparameters in the prediction neural network model used in the first month after the launch of virtual games such as Game A and Game B, so that the model representation ability of the neural network model is greatly improved, thereby improving the prediction results of the expected revenue of the virtual game in the first month.

[0152] It should be noted that the method of hyperparameter processing provided in the embodiment of the present application can be applied to scenes such as advertising, live broadcast or short video in addition to being applied to virtual game scenes, and is not limited in the specific embodiment of the present application. In addition, in addition to predicting the relevant income and costs of the first month, it is also possible to predict the relevant income and costs of the second month and the first year of the product launch; or, it can also be used to predict the number of people active on the product, predict the number of clicks on the product, predict the amount of attention paid to the product, etc., such as predicting the number of active game players, predicting the number of online game downloads, etc., which are not limited in the specific embodiment of the present application.

[0153] The above mainly introduces the scheme provided by the embodiment of the present application from the perspective of the method. It can be understood that in order to realize the above functions, the hardware structure and / or software module corresponding to the execution of each function are included. Those skilled in the art should easily realize that, in combination with the modules and algorithm steps of each example described in the embodiment disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0154] The embodiment of the present application can divide the functional modules of the device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0155] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0156] The hyperparameter processing device in the embodiment of the present application is described in detail below. Figure 7 FIG. 1 is a schematic diagram of an embodiment of a hyperparameter processing device provided in an embodiment of the present application. Figure 7 As shown, the hyperparameter processing device may include an acquisition unit 701 and a processing unit 702.

[0157] Among them, the acquisition unit 701 is used to obtain N first sets, each of which includes a first feedback value and a first value combination, the first value combination includes parameter values ​​of multiple hyperparameters, and each first feedback value is used to indicate the representation ability of the neural network model under the corresponding first value combination, and N is an integer greater than or equal to 2. For details, please refer to the aforementioned Figure 2 The contents described in step 201 can be understood for reference only and will not be elaborated here.

[0158] Processing unit 702 is used to determine P second sets based on each first value combination and each first feedback value when N meets the preset condition, each second set includes a second feedback value and a second value combination of multiple hyperparameters, the second feedback value is used to indicate the predictive characterization ability of the neural network model under the corresponding second value combination, and P is an integer greater than or equal to 2. For details, please refer to the aforementioned Figure 2 The contents described in step 202 can be understood for reference only and will not be elaborated here.

[0159] The processing unit 702 is used to calculate the evaluation score of the corresponding second value combination based on each first value combination and each first feedback value, and each second value combination and each second feedback value. Figure 2 The contents described in step 203 can be understood for reference only and will not be described in detail here.

[0160] Processing unit 702 is used to select a target value combination from P second value combinations based on the evaluation scores of all second value combinations, so as to perform model tuning processing on the neural network model. Figure 2 The contents described in step 204 can be understood for reference only and will not be described in detail here.

[0161] In some optional embodiments, the processing unit 702 is used to: when N is greater than a first threshold, calculate the value difference between the third set and each fourth set in the Q fourth sets, and obtain the distance vector between the third set and each fourth set, wherein the Q fourth sets are Q first sets selected from the N first sets, and the third set is selected from the Q fourth sets based on the first feedback value in each fourth set, 1≤Q<N, and Q is an integer; based on the distance vector between all third sets and fourth sets, calculate the expected gradient descent change; perform norm solving processing on the parameter range of each hyperparameter, and the largest first feedback value and the smallest first feedback value in the Q fourth sets to obtain the first diagonal vector modulus; determine the second set based on the third set, the expected gradient descent change, the first diagonal vector modulus, the preset action rate, and the preset attenuation coefficient. For details, please refer to the aforementioned Figure 3 The contents described in step 301 to step 304 can be understood for reference only and will not be elaborated here.

[0162] In other optional implementations, the processing unit 702 is used to: perform modulus processing on each of the Q first distance vectors to obtain the modulus of the corresponding first distance vector, each first distance vector being a distance vector between the third set and the corresponding fourth set; solve the quotient between each first distance vector and the modulus of the corresponding first distance vector to determine the unit vector of the corresponding first distance vector; and sum the unit vectors of the Q first distance vectors to obtain the expected gradient descent change.

[0163] In other optional embodiments, the processing unit 702 is used to: calculate the unit vector of the expected gradient descent change based on the expected gradient descent change; multiply the unit vector of the expected gradient descent change, the preset action rate, the preset attenuation coefficient and the first diagonal vector modulus to obtain the expected gradient value; and sum the third set and the expected gradient value to obtain the second set.

[0164] In other optional implementations, the processing unit 702 is used to: calculate the parameter interval length of the corresponding hyperparameter based on the parameter range of each hyperparameter, and perform norm solution processing on the parameter interval lengths of all hyperparameters to obtain a first value; calculate the difference between the maximum first feedback value and the minimum first feedback value corresponding to Q fourth sets to obtain a second value; perform norm solution processing on the first value and the second value to obtain a first diagonal vector modulus.

[0165] In other optional embodiments, the processing unit 702 is used to: perform modulus processing on the expected gradient descent change to obtain the modulus of the expected gradient descent change; solve the quotient between the expected gradient descent change and the modulus of the expected gradient descent change to obtain the unit vector of the expected gradient descent change.

[0166] In other optional implementations, the processing unit 702 is used to: sum the first value combination in the third set with the expected gradient value to obtain a second value combination; sum the first feedback value in the third set with the expected gradient value to obtain a second feedback value; and obtain a second set based on the second value combination and the second feedback value.

[0167] In other optional embodiments, the processing unit 702 is used to: calculate the gradient evaluation score of the third value combination based on each first feedback value and the third feedback value in the Q fourth sets, where the third feedback value is the second feedback value in any one of the second sets, and the third value combination is the second value combination in the second set corresponding to the third feedback value; calculate the distance evaluation score of the third value combination based on the third value combination, the first value combination in each fourth set, the parameter range of each hyperparameter, and the preset parameter adjustment coefficient; sum the gradient evaluation score of the third value combination and the distance evaluation score of the third value combination to obtain the evaluation score of the third value combination.

[0168] In some other optional embodiments, the processing unit 702 is used to: respectively calculate the first difference value between each first feedback value and the third feedback value in the Q fourth sets, and select a target difference value from the Q first difference values, where the target difference value is the minimum value among the Q first difference values; calculate the second difference value between the largest first feedback value and the smallest first feedback value in the Q fourth sets; calculate the quotient between the target difference value and the second difference value to obtain a gradient evaluation score of the third value combination.

[0169] In other optional embodiments, the processing unit 702 is used to: calculate the parameter interval length of the corresponding hyperparameter based on the parameter range of each hyperparameter, and perform norm solution processing on the parameter interval lengths of all hyperparameters to obtain a second diagonal vector modulus; respectively calculate the first similarity distance between the third value combination and the first value combination in each fourth set, and determine the minimum first similarity distance from the Q first similarity distances; calculate the distance evaluation score of the third value combination based on the second diagonal vector modulus, the minimum first similarity distance and the preset adjustment coefficient.

[0170] In other optional embodiments, the processing unit 702 is used to: when N is less than or equal to the first threshold, process each first value combination and each first feedback value in the N first sets through a random generator to generate P second sets; based on the parameter range of each hyperparameter, calculate the parameter interval length of the corresponding hyperparameter, and perform norm solution processing on the parameter interval lengths of all hyperparameters to obtain the second diagonal vector modulus; respectively calculate the second similarity distance between the fourth value combination and each first value combination, and determine the minimum second similarity distance from the second similarity distances, the fourth value combination being the second value combination in any second set; based on the second diagonal vector modulus, the minimum second similarity distance and the preset adjustment coefficient, calculate the distance evaluation score to obtain the evaluation score corresponding to the fourth value combination.

[0171] In some other optional implementations, the processing unit 702 is used to select the second value combination corresponding to the maximum evaluation score as the target value combination.

[0172] In other optional embodiments, the processing unit 702 is also used to: after selecting the second value combination corresponding to the maximum evaluation score as the target value combination, when the parameter value of each hyperparameter in the target value combination does not meet the preset default value of the corresponding hyperparameter, calculate the target default value of the corresponding hyperparameter based on the parameter range of each hyperparameter and the preset adjustment precision; adjust the parameter value of each hyperparameter to the target default value of the corresponding hyperparameter.

[0173] In other optional embodiments, the processing unit 702 is also used to: after selecting a target value combination from P second value combinations based on the evaluation scores of all second value combinations, respectively calculate the third similarity distance between the target value combination and each first value combination in the N first sets; when each third similarity distance is less than the second threshold, perform model tuning processing on the neural network model based on the target value combination.

[0174] The above describes the hyperparameter processing device in the embodiment of the present application from the perspective of modular functional entities, and the following describes the hyperparameter processing device in the embodiment of the present application from the perspective of hardware processing. Figure 8 Schematic diagram of the structure of the hyperparameter processing device provided in the embodiment of the present application. The hyperparameter processing device may have relatively large differences due to different configurations or performances, including but not limited to Figure 7 The hyperparameter processing device shown in FIG.

[0175] like Figure 8 As shown, the hyperparameter processing device 300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 322 (for example, one or more processors) and memory 332, and one or more storage media 330 (for example, one or more mass storage devices) storing application programs 342 or data 344. Among them, the memory 332 and the storage medium 330 may be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the classification processing device. Further, the central processor 322 may be configured to communicate with the storage medium 330, and execute a series of instruction operations in the storage medium 330 on the hyperparameter processing device 300. Exemplarily, the central processor 322 is used to execute computer execution instructions stored in the storage medium 330, thereby realizing the method of hyperparameter processing provided in the above-mentioned embodiment of the present application.

[0176] The hyperparameter processing device 300 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input and output interfaces 358, and / or one or more operating systems 341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0177] The steps performed by the hyperparameter processing device in the above embodiment can be based on the Figure 8 The hyperparameters shown deal with the device structure.

[0178] Need to explain, Figure 8 The CPU 322 in the memory 332 can call the computer execution instructions stored in the memory 332 to make the hyperparameter processing device execute the following Figures 2 to 5 The method in the corresponding method embodiment.

[0179] Specifically, Figure 7 The function / implementation process of the processing unit 702 in Figure 8 The central processing unit 322 in the memory 332 calls the computer execution instructions stored in the memory 332 to achieve it. Figure 7 The function / implementation process of the acquisition unit 701 in can be Figure 8 It is implemented by the input and output interface 358 in.

[0180] A computer-readable storage medium is also provided in an embodiment of the present application, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the above embodiments are implemented.

[0181] A computer program product is also provided in an embodiment of the present application, including a computer program, which, when executed by a processor, implements the steps of the methods described in the above embodiments.

[0182] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0183] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0184] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0185] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0186] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0187] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0188] The computer program product includes one or more computer instructions. When loading and executing the computer execution instruction on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instruction can be stored in a computer-readable storage medium, or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer instruction can be transmitted from a website site, a computer, a server or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server, a data center, etc. that includes one or more available media integration. Available media can be magnetic media, (for example, floppy disk, hard disk, tape), optical media (for example, DVD), or semiconductor media (such as SSD)) and the like.

[0189] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for hyperparameter processing, characterized in that: include: Obtain N first sets, each of which includes a first feedback value and a first value combination, the first value combination includes parameter values ​​of multiple hyperparameters, each of the first feedback values ​​is used to indicate the representation capability of the neural network model under the corresponding first value combination, and N is an integer greater than or equal to 2; When N satisfies a preset condition, P second sets are determined based on each of the first value combinations and each of the first feedback values, each of the second sets includes a second feedback value and a second value combination of a plurality of the hyperparameters, the second feedback value is used to indicate the predictive characterization capability of the neural network model under the corresponding second value combination, and P is an integer greater than or equal to 2; Based on each of the first value combinations and each of the first feedback values, and each of the second value combinations and each of the second feedback values, calculating an evaluation score corresponding to the second value combination; Based on the evaluation scores of all the second value combinations, a target value combination is selected from P second value combinations for use in performing model tuning processing on the neural network model.

2. The method according to claim 1, characterized in that When N satisfies a preset condition, determining P second sets based on each first value combination and each first feedback value includes: When N is greater than a first threshold, calculating the value difference between the third set and each of the fourth sets in the Q fourth sets, to obtain a distance vector between the third set and each of the fourth sets, wherein the Q fourth sets are Q first sets selected from the N first sets, and the third set is selected from the Q fourth sets based on the first feedback value in each of the fourth sets, 1≤Q<N, and Q is an integer; Calculating an expected gradient descent change based on all distance vectors between the third set and the fourth set; Performing norm solving processing on the parameter range of each of the hyperparameters, and the largest first feedback value and the smallest first feedback value in the Q fourth sets to obtain a first diagonal vector norm; A second set is determined based on the third set, the expected gradient descent change, the first diagonal vector magnitude, a preset action rate, and a preset attenuation coefficient.

3. The method according to claim 2, characterized in that Calculating the expected gradient descent change based on the distance vectors between all the third sets and the fourth set, comprising: Performing modulo processing on each of the Q first distance vectors to obtain a modulus corresponding to the first distance vector, each of the first distance vectors being a distance vector between the third set and the corresponding fourth set; Solving the quotient between each of the first distance vectors and the modulus corresponding to the first distance vector, and determining a unit vector corresponding to the first distance vector; The unit vectors of the Q first distance vectors are summed to obtain the expected gradient descent change.

4. The method according to any one of claims 2 to 3, characterized in that Determining a second set based on the third set, the expected gradient descent change, the first diagonal vector magnitude, a preset action rate, and a preset attenuation coefficient includes: Based on the expected gradient descent change, calculating a unit vector of the expected gradient descent change; Performing a product process on the unit vector of the expected gradient descent change, the preset action rate, the preset attenuation coefficient, and the first diagonal vector modulus to obtain an expected gradient value; The third set and the expected gradient value are summed to obtain a second set.

5. The method according to any one of claims 2 to 3, characterized in that The parameter range of each of the hyperparameters and the maximum first feedback value and the minimum first feedback value corresponding to the Q fourth sets are subjected to norm solving processing to obtain a first diagonal vector norm, including: Based on the parameter range of each of the hyperparameters, a parameter interval length corresponding to the hyperparameter is calculated, and a norm solving process is performed on the parameter interval lengths of all the hyperparameters to obtain a first value; Calculate the difference between the maximum first feedback value and the minimum first feedback value corresponding to the Q fourth sets to obtain a second value; A norm solving process is performed on the first value and the second value to obtain a first diagonal vector norm.

6. The method according to claim 4, characterized in that Calculating a unit vector of the expected gradient descent change based on the expected gradient descent change includes: Performing a modulus process on the expected gradient descent change to obtain a modulus of the expected gradient descent change; Solve the quotient between the expected gradient descent change and the modulus of the expected gradient descent change to obtain a unit vector of the expected gradient descent change.

7. The method according to claim 4, characterized in that The third set and the expected gradient value are summed to obtain a second set, including: Summing the first value combination in the third set and the expected gradient value to obtain a second value combination; Summing the first feedback value in the third set and the expected gradient value to obtain a second feedback value; A second set is obtained based on the second value combination and the second feedback value.

8. The method according to any one of claims 2 to 3, characterized in that Calculating an evaluation score corresponding to the second value combination based on each of the first value combinations and each of the first feedback values, and each of the second value combinations and each of the second feedback values, includes: Calculate a gradient evaluation score of a third value combination based on each of the first feedback values ​​and the third feedback value in the Q fourth sets, where the third feedback value is any second feedback value in the second set, and the third value combination is the second value combination in the second set corresponding to the third feedback value; Calculate the distance evaluation score of the third value combination based on the third value combination, each first value combination in the fourth set, the parameter range of each hyperparameter, and the preset parameter adjustment coefficient; The gradient evaluation score of the third value combination and the distance evaluation score of the third value combination are summed to obtain the evaluation score of the third value combination.

9. The method according to claim 8, characterized in that Calculating a gradient evaluation score of a third value combination based on each of the first feedback value and the third feedback value in the Q fourth sets includes: respectively calculating a first difference value between each of the first feedback values ​​and the third feedback value in the Q fourth sets, and selecting a target difference value from the Q first difference values, wherein the target difference value is a minimum value among the Q first difference values; Calculating a second difference value between the largest first feedback value and the smallest first feedback value in the Q fourth sets; A quotient between the target difference value and the second difference value is calculated to obtain a gradient evaluation score of a third value combination.

10. The method according to claim 8, characterized in that Calculating a distance evaluation score of the third value combination based on the third value combination, each first value combination in the fourth set, a parameter range of each hyperparameter, and a preset parameter adjustment coefficient includes: Based on the parameter range of each of the hyperparameters, the parameter interval length corresponding to the hyperparameter is calculated, and the parameter interval lengths of all the hyperparameters are subjected to norm solving processing to obtain a second diagonal vector norm; respectively calculating first similarity distances between the third value combination and each first value combination in the fourth set, and determining a minimum first similarity distance from the Q first similarity distances; Based on the second diagonal vector modulus, the minimum first similarity distance and the preset adjustment coefficient, a distance evaluation score of the third value combination is calculated.

11. The method according to claim 1, characterized in that When N satisfies a preset condition, determining P second sets based on each of the first value combinations and each of the first feedback values ​​includes: When N is less than or equal to a first threshold, each first value combination and each first feedback value in the N first sets are processed by a random generator to generate P second sets; Calculating an evaluation score corresponding to the second value combination based on each of the first value combinations and each of the first feedback values, and each of the second value combinations and each of the second feedback values, includes: Based on the parameter range of each of the hyperparameters, the parameter interval length corresponding to the hyperparameter is calculated, and the parameter interval lengths of all the hyperparameters are subjected to norm solving processing to obtain a second diagonal vector norm; respectively calculating the second similarity distance between the fourth value combination and each of the first value combinations, and determining the minimum second similarity distance from the second similarity distances, the fourth value combination being any second value combination in the second set; Based on the second diagonal vector modulus, the minimum second similarity distance and the preset adjustment coefficient, a distance evaluation score is calculated to obtain an evaluation score corresponding to the fourth value combination.

12. The method according to any one of claims 1 to 3 and 10, characterized in that: Selecting a target value combination from P second value combinations based on the evaluation scores of all the second value combinations includes: The second value combination corresponding to the maximum evaluation score is selected as the target value combination.

13. The method according to claim 12, characterized in that After selecting the second value combination corresponding to the largest evaluation score as the target value combination, the method further includes: When the parameter value of each of the hyperparameters in the target value combination does not satisfy the preset default value of the corresponding hyperparameter, calculating the target default value of the corresponding hyperparameter based on the parameter range and preset adjustment accuracy of each hyperparameter; The parameter value of each of the hyperparameters is adjusted to the target default value corresponding to the hyperparameter.

14. The method according to claim 1, characterized in that After selecting a target value combination from P second value combinations based on the evaluation scores of all the second value combinations, the method further includes: respectively calculating a third similarity distance between the target value combination and each first value combination in the N first sets; When each of the third similarity distances is less than the second threshold, the neural network model is subjected to model tuning processing based on the target value combination.

15. A hyperparameter processing device, characterized in that: include: an acquisition unit, configured to acquire N first sets, each of which includes a first feedback value and a first value combination, the first value combination includes parameter values ​​of a plurality of hyperparameters, each of which is used to indicate a representation capability of a neural network model corresponding to the first value combination, and N is an integer greater than or equal to 2; a processing unit, configured to determine, when N satisfies a preset condition, P second sets based on each of the first value combinations and each of the first feedback values, each of the second sets including a second feedback value and a plurality of second value combinations of the hyperparameters, the second feedback value being used to indicate the predictive characterization capability of the neural network model under the corresponding second value combination, and P being an integer greater than or equal to 2; The processing unit is configured to calculate an evaluation score corresponding to each of the first value combinations and each of the first feedback values, and each of the second value combinations and each of the second feedback values; The processing unit is used to select a target value combination from P second value combinations based on the evaluation scores of all the second value combinations, so as to perform model tuning processing on the neural network model.

16. A hyperparameter processing device, characterized in that: include: An input / output interface, a processor and a memory, wherein program instructions are stored in the memory; The processor is used to execute program instructions stored in the memory to perform the method according to any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises instructions, which, when executed on a computer device, cause the computer device to perform the method according to any one of claims 1 to 14.

18. A computer program product, characterized in that The computer program product comprises instructions, which, when executed on a computer device, cause the computer device to perform the method according to any one of claims 1 to 14.