Super-parameter optimization method, super-parameter recommendation method and related device
Through clustering and hyperparameter optimization, the problem of finding the optimal solver hyperparameter is solved to improve the solution performance based on mathematical planning problems in heterogeneous data sets.
Patent Information
- Application Number
- CN202311737518.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-17
AI Technical Summary
When processing heterogeneous data sets, it is difficult for the prior art to find an optimal set of solver hyperparameters to improve the solution performance of mathematical planning problems.
By obtaining the solve characteristics of multiple mathematical planning problems, clustering these problems, and superparameter optimization is performed for each cluster cluster to obtain the target superparameter of each cluster, thereby optimizing the hyperparameter configuration of the solver.
The solution performance of mathematical planning problems is improved, and through clustering and hyperparameter optimization, hyperparameter configurations suitable for various problems can be found more effectively.
Smart Images

Figure CN120163210A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computers, and in particular, to a hyperparameter optimization method, a hyperparameter recommendation method, and related devices. Background Art
[0002] Mathematical programming (MP) problems are an important area in optimization problems, and many practical problems faced in operations research can be handled by mathematical programming problems. However, with the rapid increase in the complexity and scale of the actual problem business scenarios, some MP problems in the industrial community have become extremely large and complex, and people usually use mathematical programming solvers to solve actual MP problems. The solver can optimize the solution time of the MP problem by adjusting the hyper-parameter configuration.
[0003] Currently, for the hyper-parameter configuration of the solver, a series of black-box optimization methods can be used to output a set of optimal hyperparameters for one or more MP problems uploaded by the user. However, in many scenarios, the multiple MP problems uploaded by the user are often heterogeneous, that is, a heterogeneous dataset, which means that the optimal solver hyperparameters corresponding to these MP problems may vary greatly, and it is difficult to find a set of optimal hyperparameters through black-box optimization.
[0004] For the heterogeneous dataset uploaded by the user, how to configure suitable hyperparameters is a technical problem to be solved urgently. Summary of the Invention
[0005] Embodiments of the present application provide a hyperparameter optimization method, a hyperparameter recommendation method, and related devices, which are used to improve the solving performance of the solver when the solver solves mathematical programming problems.
[0006] The first aspect of the embodiments of the present application provides a hyperparameter optimization method:
[0007] Obtain the dataset to be solved, where the dataset includes multiple mathematical programming (MP) problems; by configuring different hyperparameters for the solver, solve multiple MP problems respectively to obtain the solution characteristics of multiple MP problems. Among them, the solution characteristic of the first MP problem indicates the change in the solution performance of the solver when configuring different hyperparameters to solve the first MP problem respectively, and the first MP problem is any one of the multiple MP problems; then, according to the solution characteristics of each MP problem, divide the multiple MP problems into M clustering clusters; where each clustering cluster includes at least one MP problem, M≥2 and M is an integer; optimize the solver for the MP problems in each clustering cluster to obtain the target hyperparameters of each clustering cluster; where the target hyperparameters corresponding to different clustering clusters are different, and each MP problem in the clustering cluster corresponds to the same target hyperparameter, and the target hyperparameter of each clustering cluster is used to optimize the hyperparameter configuration of the solver.
[0008] In this application, hyperparameters are parameters that must be determined before solving the MP problem and cannot be updated during the calculation process. The solver includes multiple configurable hyperparameters. Based on different hyperparameters, the solution performance of the solver is different. Among them, the solution performance can be understood as the feedback information obtained by the solver during the solution process, and the feedback information includes information such as solution time, solution success rate, number of linear programming (LP) iterations, and number of nodes.
[0009] Using the above method, cluster multiple MP problems according to the solution characteristics of each MP problem. Since this solution characteristic is associated with the solution performance, MP problems with similar solution performance can be clustered into the same cluster, improving the efficiency of solving MP problems. When performing hyperparameter optimization for a class of MP problems with similar solution performance, a hyperparameter configuration with better solution performance can be obtained, thereby improving the solution performance of the solver.
[0010] In some optional embodiments, optimizing the solver for the MP problems in each clustering cluster to obtain the target hyperparameters of each clustering cluster includes: performing hyperparameter optimization on the MP problems belonging to the first clustering cluster based on the black-box optimization algorithm to obtain the target hyperparameters for the solver to solve the MP problems belonging to the first clustering cluster, and the first clustering cluster is any one of the M clustering clusters.
[0011] In this application, regarding the influence between the hyperparameter configuration of the solver and the solver performance, since the internal working principle cannot be directly accessed or understood, it can be regarded as a black-box function, where the independent variable is the hyperparameter configuration and the dependent variable is the solution performance. Using the above method, the hyperparameter optimization process is automatically performed through the black-box optimization algorithm, improving the efficiency and accuracy of the optimization process.
[0012] In some alternative embodiments, by configuring different hyperparameters for the solver, multiple MP problems are solved respectively to obtain the solution characteristics of the multiple MP problems, including: for the first MP problem, determining the first information and the second information of the first MP problem, where the first information is the solution performance of the first MP problem under different hyperparameters, and the second information is the solution performance of the first MP problem under the baseline hyperparameters; wherein, the baseline hyperparameters are configured values, or the baseline hyperparameters are any one of the multiple hyperparameters; according to the first information and the second information of the first MP problem, the solution characteristics of the first MP problem are determined.
[0013] In the foregoing embodiments, since the first information includes the solution performance data of the first MP problem under various hyperparameter settings, the performance of the first MP problem under different hyperparameters can be comprehensively understood; and the performance reference under the baseline hyperparameters helps to evaluate the relative effectiveness of other hyperparameter configurations; by combining the first information and the second information, the solution characteristics of the first MP problem can more accurately predict and evaluate the impact of different hyperparameter configurations on the solution performance of the first MP problem.
[0014] In some alternative embodiments, the solution characteristics of the first MP problem are a set of differences between each first information and the corresponding second information.
[0015] In some alternative embodiments, the solution characteristics of the first MP problem are a set of first ratios, where the first ratio is the ratio of the difference between each first information and the corresponding second information to the second information.
[0016] In the foregoing embodiments, by calculating the difference between each first information and the corresponding second information, the degree of difference between the data can be understood. These differences can help to understand the "outliers" or "deviant points" in the data, so as to better grasp the characteristics of the data distribution; by calculating and analyzing the ratio of the difference to the second information, the patterns and trends in the data can be identified. This ratio can help to understand the change of the first information relative to the second information, so as to better grasp the change law of the data.
[0017] In some alternative embodiments, according to the solution characteristics of each MP problem, multiple MP problems are determined as M clustering clusters, including: according to the solution characteristics of each MP problem, comparing any two MP problems to determine the similarity between each two MP problems; then dividing the MP problems whose similarity meets the constraint conditions into the same clustering cluster, and finally obtaining M clustering clusters corresponding to the multiple MP problems.
[0018] In the foregoing embodiments, based on the similarity between different MP problems as the clustering basis, MP problems with similar characteristics can be divided into the same group, making the effect of subsequent hyperparameter optimization for the MP problems in the same clustering cluster better.
[0019] In the present application, before comparing the solution feature progress of each two MP problems, MP problems that failed to be solved based on each hyperparameter within a preset time can be deleted from multiple MP problems, because the MP problems that failed to be solved belong to bad point data, that is, they may be unsolvable MP problems. Removing bad point data can reduce the amount of calculation and improve the subsequent clustering effect.
[0020] In some optional embodiments, MP problems whose similarities satisfy constraints are divided into a cluster to obtain M clusters corresponding to multiple MP problems, including: constructing an affinity matrix based on the similarity between every two MP problems of the aforementioned multiple MP problems, wherein the multiple elements of each row in the affinity matrix are the similarities between the second MP problem and each of the multiple MP problems, the second MP problem is any one of the multiple MP problems, and the second MP problem in each row will not be repeated; and then clustering the multiple MP problems using a clustering algorithm and the affinity matrix to obtain M clusters.
[0021] In the aforementioned implementation, each element in the affinity matrix represents the similarity between two MP problems, which helps to analyze the association and similarity between different MP problems. By comparing the elements between different rows, the similarities and differences between different MP problems can be understood. When clustering multiple MP problems, a better clustering effect is obtained.
[0022] In some optional embodiments, before clustering multiple MP problems based on a clustering algorithm and an affinity matrix to obtain M clustering clusters, the method further includes: summing the elements in each row of the affinity matrix to obtain a summed value of the elements in each row of the affinity matrix, wherein the summed value can be used to indicate the degree of compatibility of the second MP problem in the row with MP problems other than the second MP problem in the multiple MP problems, and the larger the summed value, the more similar the second MP problem is to the other MP problems; deleting the second MP problem corresponding to a target row in the affinity matrix to obtain a processed affinity matrix, wherein the summed value of the target row is less than a first threshold, and the processed affinity matrix is used to cluster multiple MP problems based on a clustering algorithm to obtain M clustering clusters.
[0023] In this application, each row in the affinity matrix represents the degree of clustering of a second MP problem with other MP problems. Deleting the rows with lower sum values in the affinity matrix can be understood as deleting the MP problems with lower clustering degrees among multiple MP problems. Among them, the first threshold is not a fixed value, but a value used to distinguish "normal values" from "outliers" and define the boundary point. Exemplarily, for example, if 90% of the sum values are greater than 0.3, then the first threshold is 0.3 at this time, so as to regard the values greater than 0.3 as normal values and the values less than or equal to 0.3 as outliers. The selection of the first threshold may vary according to the data distribution.
[0024] In the foregoing embodiment, the second MP problem with a low clustering degree is used as an outlier among multiple MP problems. The screening of outliers can increase the similarity of samples within the cluster after clustering and improve the clustering effect.
[0025] In some alternative embodiments, the method further includes: training a first model based on each MP problem and the corresponding clustering cluster of each MP problem, where the first model is used to extract features of the MP problem to obtain the corresponding model representation.
[0026] In some alternative embodiments, the first model is a graph neural network GNN model. Training the first model based on each MP problem and the corresponding clustering cluster of each MP problem includes: generating a plurality of bipartite graphs according to multiple MP problems, where each bipartite graph corresponds to a graph representation of an MP problem; obtaining a bipartite graph group from the plurality of bipartite graphs, the bipartite graph group includes an anchor bipartite graph, a positive bipartite graph, and a negative bipartite graph, the anchor bipartite graph is one of the plurality of bipartite graphs, the MP problem corresponding to the positive bipartite graph belongs to the same clustering cluster as the MP problem corresponding to the anchor bipartite graph, and the MP problem corresponding to the negative bipartite graph does not belong to the same clustering cluster as the MP problem corresponding to the anchor bipartite graph; inputting the bipartite graph group into the loss function of the first model to update the first model.
[0027] In this application, the trained first model can be used to extract the model representation of the MP problem. The model representation can be understood as the vectorized features obtained by the first model after analyzing the MP problem (or called problem characteristics or structural characteristics). In addition, in addition to obtaining the model representation of the target MP problem using the first model; in practical applications, the problem model statistical features of the MP problem can also be extracted by manual extraction, and the structural features of the MP problem constraint matrix can also be extracted using a convolutional neural network.
[0028] In this application, through the trained first model, the model representation of each MP problem in the training set can be obtained, and the model representation of each MP problem and the corresponding target hyperparameters are stored in the model representation - hyperparameter database. In the above - mentioned manner, since the elements in the clustering cluster have similar solution performance information, after the trained first model obtains the similar model representation of the target MP problem, combined with the model representation - hyperparameter database, the hyperparameters corresponding to the similar model representation in the database are used as the target hyperparameters.
[0029] The second aspect of the embodiment of this application provides a hyperparameter recommendation method:
[0030] Obtain the target mathematical programming (MP) problem; perform feature extraction on the target MP problem based on the first model to obtain the model representation of the target MP problem; determine the target hyperparameters matching the target MP problem from the model representation - hyperparameter database according to the model representation of the target MP problem. The model representation - hyperparameter database stores multiple mapping relationships, and each mapping relationship indicates the corresponding relationship between the model representation and hyperparameters of an MP problem.
[0031] Using the above - mentioned method, after determining the model representation of the target MP problem, the hyperparameters corresponding to the similar historical MP problems in the model representation - hyperparameter database are used as the target hyperparameters of the target MP problem, without the need for traditional black - box optimization to find the optimal solution, realizing the online fast hyperparameter configuration of the solver. The first model here can refer to the first model described in the first aspect, and the obtained target hyperparameters are more in line with the solution performance of the target MP problem.
[0032] In some possible implementation manners, determining the target hyperparameters matching the target MP problem from the model representation - hyperparameter database according to the model representation of the target MP problem includes: determining at least one model representation of an MP problem in the model representation - hyperparameter database whose difference in feature similarity with the model representation of the target MP problem is less than a first value, and then determining one hyperparameter as the target hyperparameter of the target MP problem according to the hyperparameters corresponding to each model representation of the at least one MP problem in the model representation.
[0033] In the foregoing implementation manner, determining at least one MP problem similar to the target MP problem from multiple MP problems and then determining one hyperparameter as the target hyperparameter of the target MP problem is faster than the process of determining the target hyperparameter by traditional black - box optimization.
[0034] In some possible implementation manners, the first model is a graph neural network (GNN) model. Feature extraction is performed on the target MP problem based on the first model to obtain a model representation of the target MP problem, including: generating a corresponding bipartite graph according to the target MP problem, where the bipartite graph is a graph representation of the target MP problem; and using the GNN model to process the bipartite graph of the target MP problem to obtain the model representation of the target MP problem.
[0035] A third aspect of the present application provides a hyperparameter optimization device, including:
[0036] An acquisition module, configured to acquire a data set, where the data set includes multiple mathematical programming (MP) problems; a processing module, configured to solve the multiple MP problems respectively by configuring different hyperparameters for a solver to obtain solution features of the multiple MP problems, where the solution feature of the first MP problem indicates the change in the solution performance of the first MP problem when the solver configures different hyperparameters to solve the first MP problem, and the first MP problem is any one of the multiple MP problems; the processing module is further configured to divide the multiple MP problems into M clustering clusters according to the solution features of each MP problem; where each clustering cluster includes at least one MP problem, M≥2, and M is an integer; the processing module is further configured to perform optimization processing on the solver for the MP problems in each clustering cluster to obtain the target hyperparameters of each clustering cluster; where the target hyperparameters corresponding to different clustering clusters are different, and each MP problem in the clustering cluster corresponds to the same target hyperparameter, and the target hyperparameter of each clustering cluster is used to optimize the hyperparameter configuration of the solver.
[0037] In some optional implementation manners, the processing module is specifically configured to: perform hyperparameter optimization on the MP problems belonging to the first clustering cluster based on a black-box optimization algorithm to obtain the target hyperparameters for the solver to solve the MP problems belonging to the first clustering cluster, where the first clustering cluster is any one of the M clustering clusters.
[0038] In some optional implementation manners, the processing module is specifically configured to: for the first MP problem, determine the first information and the second information of the first MP problem, where the first information is the solution performance of the first MP problem under different hyperparameters, and the second information is the solution performance of the first MP problem under a benchmark hyperparameter; where the benchmark hyperparameter is a configured value, or the benchmark hyperparameter is any one of the multiple hyperparameters; and determine the solution feature of the first MP problem according to the first information and the second information of the first MP problem.
[0039] In some optional implementation manners, the solution feature of the first MP problem is a set of differences between each first information and the corresponding second information.
[0040] In some optional implementation manners, the solution feature of the first MP problem is a set of first ratios, where the first ratio is the ratio of a first difference to the second information, and the first difference is the difference between the first information and the second information.
[0041] In some optional implementations, the processing module is specifically configured to:
[0042] According to the solution characteristics of each MP problem, the similarity between every two MP problems is determined; the MP problems whose similarity satisfies the constraint conditions are divided into a cluster to obtain M clusters corresponding to multiple MP problems.
[0043] In some optional embodiments, the processing module is specifically used to: construct an affinity matrix, wherein the multiple elements in each row of the affinity matrix are the similarities between the second MP problem and each of the multiple MP problems, the second MP problem is any one of the multiple MP problems, and the second MP problem in each row is not repeated; based on the clustering algorithm and the affinity matrix, cluster the multiple MP problems to obtain M cluster clusters.
[0044] In some optional embodiments, before clustering multiple MP problems based on a clustering algorithm and an affinity matrix to obtain M clustering clusters, the processing module is further used to: sum the elements of each row in the affinity matrix to obtain a summed value for each row, wherein the summed value for each row indicates the degree of compatibility of the second MP problem in the row with MP problems other than the second MP problem in the multiple MP problems; delete the second MP problem corresponding to a target row in the affinity matrix to obtain a processed affinity matrix, wherein the summed value of the target row is less than a first threshold, and the processed affinity matrix is used to cluster multiple MP problems based on a clustering algorithm to obtain M clustering clusters.
[0045] In some optional embodiments, the processing module is further used to: train a first model based on each MP problem and the clustering cluster corresponding to each MP problem, and the first model is used to extract features of the MP problem to obtain corresponding model representations.
[0046] In some optional embodiments, the first model is a graph neural network (GNN) model, and the processing module is specifically used to: generate multiple bipartite graphs based on multiple MP problems, each bipartite graph corresponds to a graph representation of an MP problem; obtain a bipartite graph group from the multiple bipartite graphs, the bipartite graph group includes an anchor bipartite graph, a positive bipartite graph and a negative bipartite graph, the anchor bipartite graph is one of the multiple bipartite graphs, the MP problem corresponding to the positive bipartite graph and the MP problem corresponding to the anchor bipartite graph belong to the same cluster cluster, and the MP problem corresponding to the negative bipartite graph and the MP problem corresponding to the anchor bipartite graph do not belong to the same cluster cluster; input the bipartite graph group into the loss function of the first model to update the first model.
[0047] A fourth aspect of the present application provides a super parameter recommendation device, comprising:
[0048] An acquisition module for acquiring a target mathematical programming (MP) problem; a processing module for extracting features of the target MP problem based on a first model to obtain a model representation of the target MP problem; the processing module is further configured to determine a target hyperparameter matching the target MP problem from a model representation-hyperparameter database according to the model representation of the target MP problem, and the model representation-hyperparameter database stores a plurality of mapping relationships, and each mapping relationship indicates the corresponding relationship between the model representation of an MP problem and the hyperparameter.
[0049] In some alternative embodiments, the processing module is specifically configured to: determine at least one model representation of an MP problem from the model representation-hyperparameter database, where the difference in feature similarity between the model representation of the MP problem and the model representation of the target MP problem is less than a first value; and determine the target hyperparameter corresponding to the target MP problem according to the hyperparameters corresponding to each model representation of the at least one model representation of the MP problem.
[0050] In some alternative embodiments, the first model is a graph neural network (GNN) model, and the processing module is specifically configured to: generate a corresponding bipartite graph according to the target MP problem, where the bipartite graph is a graph representation of the target MP problem; and use the GNN model to process the bipartite graph to obtain a model representation of the target MP problem.
[0051] A fifth aspect of the present application provides a hyperparameter optimization device, which includes: a processor, a memory, and a transceiver. A computer program or computer instruction is stored in the memory, and the processor is configured to call and run the computer program or computer instruction stored in the memory, so that the processor implements the processing operations in the first aspect and any implementation manner in the first aspect, and the transceiver is configured to transmit and receive signals, such as: implementing the receiving and transmitting operations in the first aspect and any implementation manner in the first aspect.
[0052] A sixth aspect of the present application provides a hyperparameter recommendation device, which includes: a processor, a memory, and a transceiver. A computer program or computer instruction is stored in the memory, and the processor is configured to call and run the computer program or computer instruction stored in the memory, so that the processor implements the processing operations in the second aspect and any implementation manner in the second aspect, and the transceiver is configured to transmit and receive signals, such as: implementing the receiving and transmitting operations in the second aspect and any implementation manner in the second aspect.
[0053] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when it runs on a computer, it causes the computer to execute the first aspect and any optional method thereof described above.
[0054] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When it runs on a computer, the computer is caused to execute the method according to the second aspect and any optional method thereof described above.
[0055] In a ninth aspect, an embodiment of the present application provides a computer program. When it runs on a computer, the computer is caused to execute the method according to the first aspect and any optional method thereof described above.
[0056] In a tenth aspect, an embodiment of the present application provides a computer program. When it runs on a computer, the computer is caused to execute the method according to the second aspect and any optional method thereof described above.
[0057] In an eleventh aspect, the present application provides a chip system. The chip system includes a processor for supporting an execution device or a training device to implement the functions involved in the above aspects. For example, sending or processing the data or information involved in the above method. In a possible design, the chip system further includes a memory for storing necessary program instructions and data for the execution device or the training device. The chip system may be composed of chips or may include chips and other discrete devices.
[0058] For the above, the technical effects of the third aspect, the fifth aspect, the seventh aspect, and the ninth aspect of the present application can be understood with reference to the technical effects of the first aspect and any implementation manner of the first aspect. The technical effects of the fourth aspect, the sixth aspect, the eighth aspect, and the tenth aspect of the present application can be understood with reference to the technical effects of the second aspect and any implementation manner of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0060] Figure 1 Schematic diagram of the conventional clustering hyperparameter optimization process for the solver;
[0061] Figure 2 Schematic diagram of a process of the hyperparameter optimization method provided by an embodiment of the present application;
[0062] Figure 3 Experimental effect diagram of the hyperparameter optimization method provided by an embodiment of the present application;
[0063] Figure 4 Clustering effect diagram of the hyperparameter optimization method provided by an embodiment of the present application;
[0064] Figure 5 It is a schematic diagram of the method flow provided by the embodiments of the present application and its application mode in the solver;
[0065] Figure 6 It is a schematic flow diagram of the training and deployment module of the graph neural network in the embodiments of the present application;
[0066] Figure 7 It is a schematic flow diagram of the hyperparameter recommendation method provided by the embodiments of the present application;
[0067] Figure 8 It is a schematic flow diagram of the hyperparameter configuration module provided by the embodiments of the present application;
[0068] Figure 9 It is a schematic diagram of the solver system provided by the embodiments of the present application;
[0069] Figure 10 It is a schematic structural diagram of the hyperparameter optimization device provided by the embodiments of the present application;
[0070] Figure 11 It is a schematic structural diagram of the hyperparameter recommendation device provided by the embodiments of the present application;
[0071] Figure 12 It is a schematic structural diagram of the computing device provided by the embodiments of the present application. Detailed implementation manners
[0072] The embodiments of the present application provide a hyperparameter optimization method, a hyperparameter recommendation method and related devices, which are used to improve the solving performance of the mathematical programming solver when solving mathematical programming problems.
[0073] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0074] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0075] Mathematical programming problems are an important area in optimization problems, and many practical problems faced in operations research can be handled by mathematical programming. An optimization problem generally consists of three elements:
[0076] 1. Decision variables: The variables to be solved in an optimization problem.
[0077] 2. Objective function: An expression that needs to be maximized or minimized.
[0078] 3. Constraint conditions: Equality conditions or inequality conditions that need to be satisfied.
[0079] Exemplarily, for example, the production planning (or: scheduling Production Planning) problem of a supply chain needs to meet objectives such as maximizing factory demand and minimizing total cost, while satisfying the hard constraints and soft constraints specified by the actual business scenario. The mathematical model abstracted by people according to these requirements of the production plan can be simplified to the following mathematical programming problem:
[0080] min c T x
[0081] Ax = b
[0082] x ≥ 0.
[0083] Where x is the decision variable, A is the constraint condition (or constraint matrix), b is a vector corresponding to the scale and the number of constraints, and c is a vector corresponding to the scale and the number of variables.
[0084] In recent years, with the rapid increase in the complexity and scale of the actual problem business scenarios, some mathematical programming problems in the industrial field have become extremely large and complex, posing new requirements for the performance of the algorithms used to solve the problems. The difficulties of mathematical programming problems in the industrial field are as follows:
[0085] 1) Large scale, often with the scale of variables and constraints exceeding the million level, or even reaching the ten million level.
[0086] 2) Short time, often requiring a solution to be output within minutes.
[0087] 3) Fast change, not only the input data changes quickly, but also the iteration frequency of the model is high.
[0088] The industry generally believes that the scale of mathematical programming problems is one of the most fundamental difficulties, because as the scale increases, it will amplify the difficulty of solving other bottlenecks. Currently, there have been a large number of studies on mathematical programming algorithms.
[0089] In addition to improving the solution methods of mathematical programming problems, a simpler and more direct way is to configure the hyperparameters of the mathematical programming solver. Generally speaking, many different modules and types of solution algorithms are integrated in the mathematical programming solver, and a mathematical programming problem is usually only sensitive to some of these algorithms, which means that the calling frequency of other solution algorithms can be appropriately reduced or even directly turned off. Therefore, for a complex mathematical programming problem model, we can often control the solution of the problem model by adjusting the hyperparameter configuration of the mathematical programming solver, thereby optimizing the solution time of the problem model.
[0090] Since this solution involves a large number of applications of mathematical programming and neural networks, for the sake of easy understanding, the relevant terms and concepts involved in the embodiments of this application are introduced below.
[0091] 1. Mathematical programming: The problem of maximizing or minimizing a certain objective under certain resource and condition constraints. The statements such as "mathematical programming" and "mathematical programming (MP) problem" can be understood to have the same meaning.
[0092] 2. Mathematical model: It is a mathematical structure that is generally or approximately expressed in mathematical language for the characteristics or quantitative dependence relationship of a certain thing system. This mathematical structure is a pure relationship structure of a certain system depicted by mathematical symbols.
[0093] It should be noted that the statements such as "mathematical programming problem model", "mathematical programming model", and "problem model" can be understood as the mathematical models for mathematical programming.
[0094] 3. Linear programming (LP): When the constraints and the objective are both linear relationships and there are no integer restrictions, then this mathematical programming model is a linear programming.
[0095] 4. Mixed Integer Programming (MIP): When the constraints and the objective are both linear relationships and some of the variables have integer restrictions, the mathematical programming model is considered a mixed integer programming model.
[0096] 5. Heterogeneous data set: The mathematical programming problem model data set consists of multiple mathematical programming models. The constraint matrix of each mathematical programming model exhibits a certain structure according to the sparsity and arrangement pattern of the variable coefficients. When the mathematical programming models in a data set have similar constraint matrix structures, the data set is a homogeneous data set. Otherwise, the data set is called a heterogeneous data set.
[0097] 6. Hyperparameter configuration: Hyperparameters usually refer to the parameters that must be determined before the start of the calculation of an algorithm or problem model and cannot be updated during the calculation process. Such as optimizers, number of iterations, activation functions, learning rates, etc. in deep learning; coding methods, number of iterations, operator weights, user preferences, etc. in operations research optimization algorithms. Additionally, the algorithm type can also be regarded as a hyperparameter at a higher level. Mathematical programming solvers often contain thousands of configurable hyperparameters. According to the structure and data characteristics of the problem model to be solved, different hyperparameter configurations will directly affect the solving performance of the solver.
[0098] Specifically, a mathematical programming model is usually only sensitive to some of the algorithms, which means that the call frequency of other solving algorithms can be appropriately reduced or even directly turned off. Therefore, for a complex mathematical programming model, we can often control the solution of the model by adjusting the hyperparameter configuration of the mathematical programming solver, thereby achieving the optimization of the solution time.
[0099] 7. Affinity matrix: Also known as a similarity matrix, it is a basic statistical technique used to organize the mutual similarity between a set of data points. Similarity is similar to distance, however, it does not satisfy the properties of a metric. The similarity score for two identical points is 1, while the result of calculating a metric is zero. Typical examples of similarity measures are cosine similarity and Jaccard similarity. These similarity measures can be interpreted as the probability that two points are related. For example, if the coordinates of two data points are close, then their cosine similarity score (or their respective "affinity" scores) will be closer to 1 than for data points with a large space between them.
[0100] 8. Spectral Clustering Algorithm: The spectral clustering algorithm is based on spectral graph theory. Compared with traditional clustering algorithms, it has the advantages of being able to cluster in a sample space of any shape and converging to the global optimal solution. The spectral clustering algorithm first defines an affinity matrix that describes the similarity of paired data points according to the given sample data set, and calculates the eigenvalues and eigenvectors of the matrix. Then, appropriate eigenvectors are selected to cluster different data points.
[0101] 9. Heteroscedastic Evolutionary Bayesian Optimization (HEBO) Algorithm: By analyzing the problems and limitations existing in data input, data output, surrogate models, and acquisition functions in classical Bayesian optimization, corresponding optimizations are made for each problem: transforming and calibrating data input and output; jointly optimizing data transformation calibration and GPs kernel functions; introducing a multi-objective acquisition function to explore candidate points more robustly.
[0102] 10. Bipartite Graph: Also known as a bipartite graph, it is a special model in graph theory. For a mathematical programming problem, the constraints and decision variables in the problem model can be modeled as two different types of vertices. Each constraint corresponds to a constraint vertex, and each decision variable corresponds to a variable vertex.
[0103] For each constraint, connect the corresponding constraint vertex with the variable vertices corresponding to the decision variables involved in the constraint. (Assume a constraint is x1 + x3 < 2, then connect the constraint vertex with the variable vertices corresponding to x1 and x3). Thus, a complete graph can be obtained, where the set of constraint vertices and the set of variable vertices are both non-overlapping vertex subsets.
[0104] 11. Black-Box Optimization Problem: The expression of the objective function in a black-box optimization problem is unknown, and the corresponding objective function values can only be obtained according to the discrete independent variable values to update the values.
[0105] For the embodiments of this application, the influence of the hyperparameter configuration of the solver on the solver performance can be regarded as a function y = f(x), where x represents the hyperparameter configuration of the solver, and y represents the (average) solving time of the solver on a given mathematical programming model. Since the above function cannot be explicitly modeled, this function can be regarded as a black-box function. Therefore, finding the optimal hyperparameter configuration of the solver to minimize the solving time of the solver is a black-box optimization problem.
[0106] Exemplarily, heuristic methods such as simulated annealing algorithms and evolutionary algorithms or black-box optimization methods such as Bayesian optimization can usually be used for solving.
[0107] Current mathematical programming solvers usually treat hyperparameter configuration, or hyperparameter optimization, as a black-box optimization problem. For the dataset uploaded by the user, which includes one (or more) mathematical programming problem models, a black-box function is used to find the optimal hyperparameter configuration for the dataset.
[0108] Exemplarily, taking the Bayesian optimization method as an example, the hyperparameter configuration tool of the solver usually randomly assigns a batch of hyperparameters of the solver, and uses these hyperparameter combinations to solve the solver respectively, and solve the given dataset, and statistically calculate the (average) solving time corresponding to each solver hyperparameter combination. Then, the hyperparameter configuration tool will use the initial random data collected to fit a surrogate model (generally composed of parametric models such as Gaussian processes, random forests, neural networks, etc.), and this surrogate model is an approximation of the above black-box function. After that, the hyperparameter configuration tool will design an acquisition function and recommend a new batch of sampling points (i.e., new solver hyperparameter configurations) with the goal of maximizing the acquisition function. By designing different acquisition functions, users can control the balance between optimizing the performance of the sampling points and exploring the uncertainty of the surrogate model. Finally, call the mathematical programming solver to evaluate the black-box function, that is, evaluate the performance of the recommended hyperparameter configuration in solving the given mathematical programming model.
[0109] Repeating the above process can more efficiently explore the hyperparameter space of the solver and find the optimal hyperparameter configuration of the solver.
[0110] However, in very realistic scenarios, the datasets uploaded by users are often heterogeneous datasets, which means that the optimal solver hyperparameters corresponding to these models may vary greatly, and heterogeneous datasets are likely to fall into local optima when performing black-box optimization.
[0111] Please refer to Figure 1 , in the related technology, the solver first transmits the batch of MP problem models uploaded by the user to the problem model feature extractor to extract the statistical features of each problem model. Then, use the problem model clustering engine to cluster the problem models with similar model statistical features into the same cluster. Finally, perform hyperparameter black-box optimization on the problem models of each cluster to obtain the intra-class hyperparameter optimization results of each cluster.
[0112] Specifically, the model statistical features that divide a batch of MP problems into the same cluster can be the same number of decision variables, or artificially designed conditions such as the same constraint conditions.
[0113] The applicant's research found that the statistical features of the model can only reflect the similarity of MP problems in data distribution to a certain extent, and do not mean that these MP problems have similar performance under the same hyperparameter configuration. In the coefficient matrix of the MP problem, a slight perturbation may also bring about a significant change in the solution performance. For example, in the coefficient matrix of a mixed-integer programming problem, a slight change in the variable coefficient of a constraint may cause the constraint to be relaxed or tightened, or even cause the entire problem to change from solvable to unsolvable. At this time, the feasible solution search process of the entire problem model may also change, and thus the corresponding optimal hyperparameter configuration will also change accordingly. That is to say, it is still very difficult to find a set of optimal solver hyperparameters suitable for all problems in the cluster through hyperparameter black-box optimization for MP problems clustered based on model statistical features.
[0114] Based on this, for the optimal solver hyperparameters of heterogeneous data sets, this application provides two implementation methods in real-time examples. The following is a separate introduction:
[0115] 1. For the batch of MP problem data uploaded by the user, cluster the hyperparameter-solution performance for each MP problem, cluster the MP problems with similar hyperparameter-solution performance into the same cluster, and then perform black-box optimization on the MP problems in each cluster to obtain the corresponding optimal hyperparameters.
[0116] Please refer to Figure 2 , Figure 2 which is a schematic diagram of an embodiment of a hyperparameter optimization method provided by an embodiment of this application:
[0117] 201. Obtain a data set, where the data set includes multiple MP problems;
[0118] First, obtain the data set uploaded by the user. The data set includes multiple MP problem models. As Figure 2 shown, the problem models in the data set are MP problem A, MP problem B, MP problem C, etc., a total of n MP problem models;
[0119] 202. By configuring different hyperparameters for the solver, solve multiple MP problems respectively to obtain the solution characteristics of multiple MP problems;
[0120] As Figure 2 shown, the solver presets m hyperparameters such as hyperparameter 1, hyperparameter 2, and hyperparameter 3, and solves each MP problem in the data set respectively to obtain the solution performance information of each MP problem under different hyperparameters, and determines the solution characteristics of each MP.
[0121] It should be noted that the solution performance information can be understood as the feedback information obtained by the solver during the solution process, including the solution time for the MP problem, the solution success rate under the specified time, the number of iterations and nodes of linear programming (LP), etc. There is no specific limitation here.
[0122] Exemplarily, for the MP problem A, the solution performance of A1 regarding hyperparameter 1, the solution performance of A2 regarding hyperparameter 2, the solution performance of A3 regarding hyperparameter 3, and the solution performance of Am regarding hyperparameter m are obtained.
[0123] For the MP problem B, the solution performance of B1 regarding hyperparameter 1, the solution performance of B2 regarding hyperparameter 2, the solution performance of B3 regarding hyperparameter 3, and the solution performance of Bm regarding hyperparameter m are obtained.
[0124] For the MP problem C, the solution performance of C1 regarding hyperparameter 1, the solution performance of C2 regarding hyperparameter 2, the solution performance of C3 regarding hyperparameter 3, and the solution performance of Cm regarding hyperparameter m are obtained.
[0125] Taking the solution performance information as the solution time of each MP problem as an example, for each MP problem, taking its solution performance under the default hyperparameters of the solver as the benchmark performance, calculate the improvement ratio of the solution performance of this MP problem under each hyperparameter relative to the benchmark performance, and the solution feature is the set of improvement ratios corresponding to all hyperparameters of this MP problem.
[0126] In a possible implementation, the solution performance information of the MP problem A is shown in Table 1, and the solution feature of the MP problem A is determined according to the solution performance information.
[0127] Hyperparameter ID 0 1 2 3 4 5 6 Solving Time (s) 100 110 80 70 120 65 50 Improvement Ratio -10% 20% 30% -20% 35% 50% Solving Feature -0.1 0.2 0.3 -0.2 0.35 0.5
[0128] Table 1
[0129] The solution time corresponding to the default hyperparameter 0 is 100 seconds, and the solution times corresponding to the preset parameter numbers 1 to hyperparameter 6 are 110 seconds, 80 seconds, 70 seconds, 120 seconds, 65 seconds, and 50 seconds; at this time, the solution feature of the MP problem A is [-0.1, 0.2, 0.3, -0.2, 0.35, 0.5], and this solution feature can also be regarded as a set of multiple hyperparameters - solution performance of the MP problem A.
[0130] It should be noted that to solve the characteristics of the MP problem for the performance changes of multiple hyperparameters, it is not limited to calculating the improvement ratios of different hyperparameters. The change amplitude can also be recorded. For example, the change amplitude of the solution time relative to the baseline hyperparameter. Then, the solution characteristics of MP problem A are [-10, 20, 30, -20, 35, 50]. In addition, the solver can preset the default hyperparameter, or any one of the preset multiple hyperparameters can be regarded as the default hyperparameter, which is not specifically limited here.
[0131] 203. Determine the multiple MP problems as M clustering clusters according to the solution characteristics of each MP problem;
[0132] The specific steps include:
[0133] 1) Bad point detection
[0134] First, remove the meaningless feature points, that is, the MP problems that fail to solve for all hyperparameters within the limited time. Such problems may themselves be unsolvable mathematical programming problems.
[0135] Exemplarily, the limited time here may be set to 7200 seconds in some solvers, or it may be other times, which is not specifically limited here.
[0136] 2) Determine the similarity between every two problem models
[0137] In the embodiment of the present application, according to the solution characteristics of each MP problem, each MP problem is mapped to a high-dimensional feature space by using a kernel function, and the similarity between every two MP problems is determined according to the distance between every two MP problems in the high-dimensional feature space.
[0138] Specifically, the embodiment of the present application uses the Laplacian kernel function to calculate the similarity between every two MP problems. The formula of the Laplacian kernel function is as follows:
[0139]
[0140] Where x and y are the solution characteristics of two mathematical programming problem models respectively, and σ is a controllable parameter. The larger the value of the similarity, the more similar the corresponding mathematical programming problem models in the row and column are.
[0141] 3) Construct an affinity matrix
[0142] Use the similarity between the problem models to construct an affinity matrix. For example, there are M problem models {x1, x2, x3... x M}, and the expression of the affinity matrix is as follows:
[0143]
[0144] Among them, the elements of each row (or each column) in the affinity matrix can be regarded as a similarity set of a certain problem model to all problem models.
[0145] 4) Outlier Screening
[0146] The outliers in the data set are filtered out according to the affinity matrix of the problem model. The specific method is to sum the row elements (or column elements) corresponding to a problem model in the affinity matrix. Since each row (or column) element in the affinity matrix represents the similarity set of a problem model relative to all problem models, the sum of the row (or column) elements can be regarded as the degree of sociability of the problem model with other problem models in the large set (all problem models). When the sum value is less than a certain threshold, it means that the degree of sociability of the problem model with other problem models is low, and the similarity with other problem models in the feature space is low, so the problem model is filtered out as an outlier.
[0147] It is understandable that the outliers are screened out only to achieve a better clustering effect when executing step 5). In practical applications, even if step 4) is not executed, step 5) can be executed.
[0148] 5) Use clustering algorithm on the remaining problem model to obtain clusters
[0149] Using a clustering algorithm based on affinity matrix (such as spectral clustering algorithm), unsupervised clustering of MP problem model is completed according to the affinity matrix of the problem model.
[0150] like Figure 2 As shown, in a possible clustering result, multiple clusters are obtained according to the set clustering conditions, wherein cluster 1 includes MP problem A, MP problem C, ..., cluster 2 includes MP problem B, ..., and so on.
[0151] 204. The solver is optimized for the MP problem of each cluster to obtain the target hyperparameters of each cluster.
[0152] Finally, a black box optimization algorithm (such as the HEBO algorithm) is used to perform black box optimization on the problem model in each cluster to obtain the optimal hyperparameters corresponding to the MP problem solved by the solver for each cluster. The process of obtaining the optimal hyperparameters by black box optimization is similar to the process of dealing with the black box optimization problem mentioned above, and will not be repeated here.
[0153] The embodiments of this application use the real leaderboard dataset MIPLIB2017 Benchmark to verify the hyperparameter optimization effect of this hyperparameter optimization method. This leaderboard dataset contains 240 mixed-integer programming models from different scenarios and types, which is a typical heterogeneous dataset and is usually used to evaluate the basic capabilities of different mathematical programming solvers. It is one of the most authoritative performance test leaderboards for commercial solvers.
[0154] For the above-mentioned leaderboard dataset, the embodiments of this application construct a problem model clustering engine based on the spectral clustering algorithm and a black-box optimization engine based on the HEBO algorithm as embodiments of the aforementioned hyperparameter optimization method. The solver performance evaluation metrics include the Shifted Geometric Mean (SGM) of the number of solvable problems and the solving time. The aforementioned number of solvable problems represents the number of problems in the 240 problems of the leaderboard dataset that the solver can complete within 7200 seconds, and the solving time represents the SGM of the solving time of the solver on 240 problems.
[0155] To demonstrate the beneficial implementation effects of the present invention, the embodiments of this application also compare two different solver hyperparameter configuration schemes, namely the solver hyperparameter configuration scheme based on model statistical feature clustering and the HEBO-based solver hyperparameter optimization scheme without model clustering.
[0156] Please refer to Figure 3 , Figure 3 which shows the experimental effects of the clustering optimization of the embodiments of this application on MIPLIB2017 Benchmark.
[0157] When using the HEBO-based scheme to perform hyperparameter optimization on the entire leaderboard dataset, due to the obvious heterogeneity of the dataset, the obtained optimal solver hyperparameters only achieve an increase of 3 in the number of solvable problems (114 -> 117), while the SGM declines by 0.58% (1567.82 -> 1576.92). In the scheme based on model statistical feature clustering, although the leaderboard dataset is clustered, since the similarity of statistical features does not represent the similarity of solving performance, although the SGM is increased by 3.31% by providing multiple groups of hyperparameters, the number of solvable problems only increases to 117. In contrast, the scheme of the present invention significantly improves the performance of clustering hyperparameter optimization, achieving an increase of 11 in the number of solvable problems (114 -> 125) on all 6 types of models (5 types of problem models participating in clustering and the outlier problem model), and the SGM is optimized by 7.28% (1566.36 -> 1452.33).
[0158] At the same time, please refer to Figure 4 , Figure 4This is the clustering result of the embodiments of the present application. Among them, the densities of cluster 1 and cluster 2 are significantly higher than those of the other three clusters, corresponding to Figure 3 and the hyperparameter optimization effect of cluster 1 and cluster 2 in
[0159] is also better. In the problem model clusters with higher clustering density in the embodiments of the present application, the hyperparameter optimization effect is more obvious, indicating that in the feature space constructed based on the solution features of the problem model, the mathematical programming problem models with similar solution features do have similar hyperparameter-solution performance distributions.
[0160] Next, another solution of the embodiments of the present application will be introduced:
[0161] II. Using historical MP problems and corresponding clusters, train a hyperparameter recommendation model, and input the MP problem uploaded by the user into the trained hyperparameter recommendation model to obtain the recommended hyperparameters.
[0162] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the core method process of the embodiments of the present application and its application method in the solver.
[0163] The core method process of the embodiments of the present application can be divided into two systems: a hyperparameter optimization system and a hyperparameter recommendation system. While implementing the aforementioned hyperparameter optimization method, the former can also collect relevant data such as clustering-hyperparameter optimization of heterogeneous data sets into the training database to support the training of the hyperparameter recommendation model of the latter, enabling the latter to provide online hyperparameter configuration for the MP problem uploaded by the user.
[0164] As can be seen from Figure 5 , the training database stores multiple historical MP problems and the corresponding clusters of the MP problems, and the model representation-hyperparameter database stores the model representations of multiple historical MP problems and the optimal hyperparameters corresponding to the MP problems.
[0165] It can be understood that the aforementioned historical MP problems can be the MP problems in the data set uploaded in batches by users during hyperparameter optimization, or the MP problems originally used for training, and there is no specific limitation here.
[0166] In a possible implementation manner, the trained hyperparameter recommendation model is a graph neural network (GNN) model. The GNN model is used to extract the graph embedding representation of the MP problem, and finally, according to the model representation stored in the model representation-hyperparameter database, the corresponding optimal hyperparameters are found to provide online solver hyperparameter configuration for the MP problem.
[0167] Please refer toFigure 6 , Figure 6 This is a schematic diagram of the training and deployment module process of the GNN in the embodiments of this application.
[0168] 1) In the offline state, represent a batch of MP problems in the training database as a bipartite graph, along with the corresponding clustering results of the MP problems, and combine them with the model bipartite graph to form a GNN training dataset;
[0169] 2) Extract triples (a, p, n) from the training set as training samples, where a is a randomly selected anchor sample, p is a positive sample randomly selected from the same-class samples (same cluster) of the anchor sample, and n is a negative sample randomly selected from the non-same-class samples (different clusters) of the anchor sample;
[0170] 3) Input the above triples into the GNN to obtain corresponding graph embedding representations, and construct a contrastive learning loss function, and update the GNN parameters through the backpropagation mechanism. The contrastive learning loss function is as follows:
[0171]
[0172] where represents the graph embedding representation given by the above GNN after receiving the model bipartite graph, represents the model bipartite graph corresponding to the anchor / positive / negative samples input into the GNN in the i-th group of triples, and ε is a parameter that controls the interval between positive and negative samples in the graph representation metric space. The larger ε is, the farther the interval is;
[0173] 4) Use the trained GNN to obtain the final graph embedding representations of all MP problems in the training database, and store them together with the corresponding hyperparameter optimization results in the model representation - hyperparameter database. At the same time, use the trained GNN model as a hyperparameter recommendation model to online extract the graph embedding representation of the user-input model.
[0174] In addition, the hyperparameter recommendation process introduced in the embodiments of this application is to train an online model representator (hyperparameter recommendation model) using a known mathematical programming model and the corresponding clustering results, realize the online representation of the problem model, and complete the online fast hyperparameter configuration of the solver. This hyperparameter recommendation model can also be other neural network models outside the GNN model, such as convolutional neural networks, multi-layer perceptrons, etc., which are not specifically limited here.
[0175] Use the trained hyperparameter recommendation model to recommend the optimal hyperparameters for the MP problems uploaded by the user:
[0176] Please refer to Figure 7 , Figure 7 This is an example schematic of a hyperparameter recommendation method provided by the embodiments of this application:
[0177] 701. Obtain the target MP problem;
[0178] First, take the MP problem uploaded by the user as the target MP problem.
[0179] It can be understood that what the user uploads can be an MP problem or a dataset with multiple MP problems. In this case, any one of them is selected as the target MP problem.
[0180] 702. Obtain the model representation of the target MP problem;
[0181] Input the target MP problem into the trained hyperparameter recommendation model to obtain the model representation of the target MP problem.
[0182] In a possible implementation, the hyperparameter recommendation model is a GNN model; represent the target MP problem as a bipartite graph, and then input the bipartite graph of the target MP problem into the GNN model to obtain the graph embedding representation of the target MP problem.
[0183] 703. Determine the target hyperparameters that match the target MP problem from the model representation - hyperparameter database.
[0184] Please refer to Figure 8 , Figure 8 which is the schematic diagram of the hyperparameter configuration module process.
[0185] 1) After obtaining the user data, the GNN model performs online inference on the bipartite graph of the target MP problem and maps it to the graph embedding representation metric space;
[0186] 2) In the metric space, use the K-Nearest Neighbor (KNN) algorithm to match the input MP problem with the known mathematical programming problem models with similar graph representations;
[0187] As Figure 8 shown, at this time, 3 similar problem models, that is, the known mathematical programming problem models with similar graph representations, are obtained;
[0188] 3) Extract the optimal hyperparameters of the matched model from the aforementioned model representation - hyperparameter database as the recommended hyperparameter configuration of the input model.
[0189] Finally, extract the optimal hyperparameters of this similar problem model from the model representation - hyperparameter database as the hyperparameter configuration of the target MP problem, that is, determine the target hyperparameters.
[0190] It can be understood that in practical applications, there may be more than one optimal hyperparameter for the nearest neighbor, and multiple similar recommended hyperparameters can be provided for the user to select and configure.
[0191] In the embodiments of the present application, the GNN model can infer the graph embedding representation of the bipartite graph of the problem model. Since the training supervision information comes from the solution features of the problem model, theoretically, similar models and the online input model have a similar hyperparameter - solution performance distribution, while general GNN models based on statistical feature clustering and then training do not have this characteristic.
[0192] In addition, in addition to obtaining the graph embedding representation (model representation) of the problem model using the GNN model; in practical applications, the problem model statistical features of the MP problem model can also be extracted by manual extraction, and the structural features of the constraint matrix of the MP problem model can also be extracted using a convolutional neural network.
[0193] Combined with the foregoing embodiments, the embodiments of the present application also provide a solver system based on a pre - solution configuration method (or device).
[0194] Please refer to Figure 9 , this system may include: a mathematical programming solver capable of providing mathematical programming solver solution services, a pre - solution device capable of providing pre - solution services, an operations research optimization platform, and a storage device capable of providing storage services. Among them, after receiving the data of the MP problem to be solved uploaded by the user, the mathematical programming solver solves it and returns the obtained solution result to the user. During the solution process, the mathematical programming solver first requests the pre - solution device to perform parameter configuration. After receiving the request from the mathematical programming solver, the pre - solution device accesses the data of the MP problem and performs pre - solution according to the hyperparameter optimization method provided by the embodiments of the present application, obtains the parameters of each pre - solution operator corresponding to the MP problem, and sends them to the mathematical programming solver.
[0195] Then the mathematical programming solver continues to perform pre - solution parameter configuration according to the parameters of each pre - solution operator, and then simplifies the problem. Finally, it solves the problem according to the simplified problem to obtain the solution result. Among them, the pre - solution device includes configuration management, training and inference engines, and a black - box optimization engine. Configuration management is used to perform parameter configuration of the corresponding pre - solution operator according to the data of the MP problem. The training and inference engines are used to perform network training as needed to optimize the implementation of configuration management. During the training process, training can be performed based on the data batch - uploaded by the user and the historical data saved by the operations research optimization platform together. The black - box optimization engine is used to desensitize the obtained MP problem data to obtain desensitized data.
[0196] The operations research optimization platform can manage the desensitized data, the corresponding solution result, and the corresponding solution parameters formed based on the used MP problem data as sample data for subsequent network training of the pre - solution device. Among them, the operations research optimization platform can store the desensitized data, solution result, and solution parameters into the storage device.
[0197] In a possible implementation, the operations research and optimization platform may be ModelArts or the like, and the storage device may provide storage services based on the object storage service (OBS).
[0198] Figure 10 FIG. is a schematic structural diagram of a hyperparameter search device provided by an embodiment of the present application. As Figure 10 shown, the device includes:
[0199] An acquisition module 1001, configured to acquire a data set, where the data set includes multiple mathematical programming (MP) problems;
[0200] Specific descriptions of the acquisition module 1001 may refer to the description of step 201 in the foregoing embodiment, and will not be elaborated here.
[0201] A processing module 1002, configured to solve multiple MP problems respectively by configuring different hyperparameters for a solver, so as to obtain solution characteristics of the multiple MP problems. Among them, the solution characteristics of the first MP problem indicate the change in the solution performance of the first MP problem when different hyperparameters are configured for the solver, and the first MP problem is any one of the multiple MP problems;
[0202] The processing module 1002 is further configured to divide the multiple MP problems into M clustering clusters according to the solution characteristics of each MP problem; where each clustering cluster includes at least one MP problem, M≥2, and M is an integer;
[0203] The processing module 1002 is further configured to perform optimization processing on the solver for the MP problems in each clustering cluster to obtain target hyperparameters for each clustering cluster; where the target hyperparameters corresponding to different clustering clusters are different, each MP problem in the clustering cluster corresponds to the same target hyperparameter, and the target hyperparameters of each clustering cluster are used to optimize the hyperparameter configuration of the solver.
[0204] Specific descriptions of the processing module 1002 may refer to the descriptions of steps 202 to 204 in the foregoing embodiment, and will not be elaborated here.
[0205] Figure 11 FIG. is a schematic structural diagram of a hyperparameter recommendation device provided by an embodiment of the present application. As Figure 11 shown, the device includes:
[0206] An acquisition module 1101, configured to acquire a target mathematical programming (MP) problem;
[0207] Specific descriptions of the acquisition module 1101 may refer to the description of step 701 in the foregoing embodiment, and will not be elaborated here.
[0208] A processing module 1102, configured to perform feature extraction on the target MP problem based on a first model to obtain a model representation of the target MP problem;
[0209] The processing module 1102 is further configured to determine, according to the model representation of the target MP problem, a target hyperparameter matching the target MP problem from a model representation - hyperparameter database, where the model representation - hyperparameter database stores multiple mapping relationships, and each mapping relationship indicates a correspondence between a model representation of an MP problem and a hyperparameter.
[0210] For the specific description of the processing module 1102, reference may be made to the descriptions of steps 702 to 703 in the foregoing embodiment, which will not be elaborated herein.
[0211] It should be noted that for the information interaction, execution process, etc. among the above - mentioned device modules / units, since they are based on the same concept as the method embodiment of the present application, the technical effects brought by them are the same as those of the method embodiment of the present application. For the specific content, reference may be made to the description in the method embodiment shown in the foregoing of the embodiments of the present application, which will not be elaborated herein.
[0212] Figure 12 FIG. is a possible logical structure diagram of a computing device provided by an embodiment of the present application. Among them, Figure 12 The computing device 1200 may be Figure 10 a hyperparameter optimization device for implementing Figure 2 the steps of the corresponding embodiment; or, Figure 12 The computing device 1200 may be Figure 11 a hyperparameter recommendation device for implementing Figure 7 the steps of the corresponding embodiment. As Figure 12 shown, the computing device 1200 provided by the embodiment of the present application includes: a processor 1201, a communication interface 1202, a memory 1203, and a bus 1204. The processor 1201, the communication interface 1202, and the memory 1203 are interconnected through the bus 1204. In the embodiment of the present application, the processor 1201 is configured to control and manage the operations of the computing device 1200. For example, the processor 201 is configured to solve the MP problem uploaded by the user according to multiple preset hyperparameters to obtain the solution features of the MP problem. The communication interface 1202 is configured to support the computing device 1200 to communicate. For example, the communication interface 1202 may receive the MP problem uploaded by the user and send the hyperparameter optimization or hyperparameter recommendation result. The memory 1203 is used to store the program code and data of the computing device 1200.
[0213] Among them, the processor 1201 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 1204 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 only a thick line is used to represent it in Figure 12 , but it does not mean that there is only one bus or one type of bus.
[0214] The embodiments of this application also provide a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computer device, it causes at least one computer device to execute the method described in the foregoing Figure 2 or Figure 7 illustrated embodiments.
[0215] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the method described in the foregoing Figure 2 or Figure 7 illustrated embodiments.
[0216] The communication device provided by the embodiments of this application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, etc. The processing unit can execute the computer-executable instructions stored in the storage unit to cause the chip to execute the foregoing Figure 2 or Figure 7The method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, cache, etc., and the storage unit can also be a storage unit outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0217] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0218] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0219] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0220] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0221] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the part that essentially contributes to the technical solution of this application, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or an access network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.
Claims
1. A hyperparameter optimization method, characterized in that, Including: Obtain a data set, where the data set includes multiple mathematical programming (MP) problems; By configuring different hyperparameters for the solver, solve the multiple MP problems respectively to obtain the solution characteristics of the multiple MP problems. Among them, the solution characteristic of the first MP problem indicates the change in the solution performance of the first MP problem when the solver configures different hyperparameters to solve the first MP problem respectively, and the first MP problem is any one of the multiple MP problems; According to the solution characteristics of each MP problem, divide the multiple MP problems into M clustering clusters; where each clustering cluster includes at least one of the MP problems, M≥2, and M is an integer; Optimize the solver for the MP problems in each clustering cluster to obtain the target hyperparameters of each clustering cluster; where the target hyperparameters corresponding to different clustering clusters are different, each MP problem in the clustering cluster corresponds to the same target hyperparameter, and the target hyperparameters of each clustering cluster are used to optimize the hyperparameter configuration of the solver.
2. The method according to claim 1, characterized in that, The optimizing the solver for the MP problems in each clustering cluster to obtain the target hyperparameters of each clustering cluster includes: Based on the black-box optimization algorithm, perform hyperparameter optimization on the MP problems belonging to the first clustering cluster to obtain the target hyperparameters for the solver to solve the MP problems belonging to the first clustering cluster, where the first clustering cluster is any one of the M clustering clusters.
3. The method according to claim 1 or 2, characterized in that, The obtaining the solution characteristics of the multiple MP problems by configuring different hyperparameters for the solver respectively includes: For the first MP problem, determine the first information and the second information of the first MP problem. The first information is the solution performance of the first MP problem under different hyperparameters, and the second information is the solution performance of the first MP problem under the benchmark hyperparameters; where the benchmark hyperparameters are configuration values, or the benchmark hyperparameters are any one of the multiple hyperparameters; According to the first information and the second information of the first MP problem, determine the solution characteristic of the first MP problem.
4. The method according to claim 3, characterized in that, The solution characteristic of the first MP problem is a set of differences between each of the first information and the second information.
5. The method according to claim 3, characterized in that, The solution characteristic of the first MP problem is a set of first ratios, where the first ratio is the ratio of the first difference to the second information, and the first difference is the difference between the first information and the second information.
6. The method according to any one of claims 1-5, characterized in that, The dividing the multiple MP problems into M clustering clusters according to the solution characteristics of each MP problem includes: According to the solution characteristics of each MP problem, determine the similarity between every two of the MP problems; Divide the MP problems whose similarity satisfies the constraint conditions into one clustering cluster to obtain the M clustering clusters corresponding to the multiple MP problems.
7. The method according to claim 6, characterized in that, The dividing the MP problems whose similarity satisfies the constraint conditions into one clustering cluster to obtain the M clustering clusters corresponding to the multiple MP problems includes: Constructing an affinity matrix, wherein a plurality of elements in each row of the affinity matrix are respectively the similarities between the second MP problem and each of the plurality of MP problems, the second MP problem is any one of the plurality of MP problems, and the second MP problems in each row are not repeated; Based on the clustering algorithm and the affinity matrix, the multiple MP problems are clustered to obtain M clusters.
8. The method according to claim 7, characterized in that, Before clustering the multiple MP problems based on the clustering algorithm and the affinity matrix to obtain M clusters, the method further includes: Summing the elements of each row in the affinity matrix to obtain a summed value of each row, wherein the summed value of each row indicates the degree of compatibility of the second MP problem in the row with the MP problems other than the second MP problem in the plurality of MP problems; The second MP problem corresponding to the target row in the affinity matrix is deleted to obtain a processed affinity matrix, the sum of the target row is less than a first threshold, and the processed affinity matrix is used to cluster the multiple MP problems based on a clustering algorithm to obtain M cluster clusters.
9. The method according to any one of claims 1 to 8, characterized in that, The method further comprises: A first model is trained based on each MP problem and a cluster corresponding to each MP problem, and the first model is used to extract features of the MP problem to obtain a corresponding model representation.
10. The method according to claim 9, characterized in that, The first model is a graph neural network GNN model, and the first model is trained based on each MP problem and the cluster cluster corresponding to each MP problem, including: According to the multiple MP problems, generate multiple bipartite graphs, each of the bipartite graphs corresponds to a graph representation of an MP problem; Obtaining a bipartite graph group from the multiple bipartite graphs, the bipartite graph group includes an anchor bipartite graph, a positive bipartite graph, and a negative bipartite graph, the anchor bipartite graph is one of the multiple bipartite graphs, the MP problem corresponding to the positive bipartite graph and the MP problem corresponding to the anchor bipartite graph belong to the same cluster, and the MP problem corresponding to the negative bipartite graph and the MP problem corresponding to the anchor bipartite graph do not belong to the same cluster; The bipartite graph group is input into the loss function of the first model to update the first model.
11. A hyperparameter recommendation method, characterized in that, include: Obtain the target mathematical programming MP problem; Extracting features of the target MP problem based on the first model to obtain a model representation of the target MP problem; According to the model representation of the target MP problem, a target hyperparameter matching the target MP problem is determined from a model representation-hyperparameter database, wherein the model representation-hyperparameter database stores a plurality of mapping relationships, each of which indicates a corresponding relationship between a model representation of an MP problem and a hyperparameter.
12. The method according to claim 11, characterized in that, Determining target hyperparameters matching the target MP problem from a model representation-hyperparameter database according to the model representation of the target MP problem includes: Determine from the model representation-hyperparameter database at least one model representation of an MP problem whose difference in feature similarity with the model representation of the target MP problem is less than a first value; Determine the target hyperparameters corresponding to the target MP problem according to the hyperparameters corresponding to the model representation of each MP problem in the model representation of the at least one MP problem.
13. The method according to claim 11 or 12, characterized in that, The first model is a graph neural network GNN model. The feature extraction of the target MP problem based on the first model to obtain the model representation of the target MP problem includes: Generate a corresponding bipartite graph according to the target MP problem, where the bipartite graph is a graph representation of the target MP problem; Use the GNN model to process the bipartite graph to obtain the model representation of the target MP problem.
14. A hyperparameter optimization device, characterized in that, It includes: An acquisition module for acquiring a data set, where the data set includes multiple mathematical programming MP problems; A processing module for solving the multiple MP problems respectively by configuring different hyperparameters for the solver to obtain the solution features of the multiple MP problems. Among them, the solution feature of the first MP problem indicates the change in the solution performance of the first MP problem when the solver configures different hyperparameters to solve the first MP problem, and the first MP problem is any one of the multiple MP problems; The processing module is further configured to divide the multiple MP problems into M clustering clusters according to the solution features of each MP problem; where each clustering cluster includes at least one of the MP problems, M≥2, and M is an integer; The processing module is further configured to optimize the solver for the MP problems in each clustering cluster to obtain the target hyperparameters of each clustering cluster; where the target hyperparameters corresponding to different clustering clusters are different, each MP problem in the clustering cluster corresponds to the same target hyperparameter, and the target hyperparameters of each clustering cluster are used to optimize the hyperparameter configuration of the solver.
15. The device according to claim 14, characterized in that, The processing module is specifically configured to: Perform hyperparameter optimization on the MP problems belonging to the first clustering cluster based on a black-box optimization algorithm to obtain the target hyperparameters for the solver to solve the MP problems belonging to the first clustering cluster, where the first clustering cluster is any one of the M clustering clusters.
16. The device according to claim 14 or 15, characterized in that, The processing module is specifically configured to: For the first MP problem, determine the first information and the second information of the first MP problem. The first information is the solution performance of the first MP problem under different hyperparameters, and the second information is the solution performance of the first MP problem under the benchmark hyperparameters; where the benchmark hyperparameters are configuration values, or the benchmark hyperparameters are any one of the multiple hyperparameters; Determine the solution feature of the first MP problem according to the first information and the second information of the first MP problem.
17. The device according to claim 16, characterized in that, The solution feature of the first MP problem is a set of differences between each of the first information and the second information.
18. The device according to claim 16, characterized in that, The solution feature of the first MP problem is a set of first ratios, where the first ratio is the ratio of the first difference to the second information, and the first difference is the difference between the first information and the second information.
19. The device according to any one of claims 14 - 18, characterized in that, The processing module is specifically configured to: Determine the similarity between every two of the MP problems according to the solution features of each MP problem; Divide the MP problems whose similarity satisfies the constraint conditions into one clustering cluster to obtain M clustering clusters corresponding to the multiple MP problems.
20. The device according to claim 19, characterized in that, The processing module is specifically used for: Constructing an affinity matrix, wherein a plurality of elements in each row of the affinity matrix are respectively the similarities between the second MP problem and each of the plurality of MP problems, the second MP problem is any one of the plurality of MP problems, and the second MP problems in each row are not repeated; Based on the clustering algorithm and the affinity matrix, the multiple MP problems are clustered to obtain M clusters.
21. The device according to claim 20, characterized in that, Before clustering the multiple MP problems based on the clustering algorithm and the affinity matrix to obtain M clusters, the processing module is further used to: Summing the elements of each row in the affinity matrix to obtain a summed value of each row, wherein the summed value of each row indicates the degree of compatibility of the second MP problem in the row with the MP problems other than the second MP problem in the plurality of MP problems; The second MP problem corresponding to the target row in the affinity matrix is deleted to obtain a processed affinity matrix, the sum of the target row is less than a first threshold, and the processed affinity matrix is used to cluster the multiple MP problems based on a clustering algorithm to obtain M cluster clusters.
22. The device according to any one of claims 14 to 21, characterized in that, The processing module is further used for: A first model is trained based on each MP problem and a cluster corresponding to each MP problem, and the first model is used to extract features of the MP problem to obtain a corresponding model representation.
23. The device according to claim 22, characterized in that, The first model is a graph neural network (GNN) model, and the processing module is specifically used to: According to the multiple MP problems, generate multiple bipartite graphs, each of the bipartite graphs corresponds to a graph representation of an MP problem; Obtaining a bipartite graph group from the multiple bipartite graphs, the bipartite graph group includes an anchor bipartite graph, a positive bipartite graph, and a negative bipartite graph, the anchor bipartite graph is one of the multiple bipartite graphs, the MP problem corresponding to the positive bipartite graph and the MP problem corresponding to the anchor bipartite graph belong to the same cluster, and the MP problem corresponding to the negative bipartite graph and the MP problem corresponding to the anchor bipartite graph do not belong to the same cluster; The bipartite graph group is input into the loss function of the first model to update the first model.
24. A hyperparameter recommendation device, characterized in that, include: An acquisition module is used to acquire a target mathematical programming MP problem; A processing module is used to extract features of the target MP problem based on a first model to obtain a model representation of the target MP problem; the processing module is also used to determine a target hyperparameter matching the target MP problem from a model representation-hyperparameter database based on the model representation of the target MP problem, wherein the model representation-hyperparameter database stores a plurality of mapping relationships, each mapping relationship indicating a corresponding relationship between a model representation of an MP problem and a hyperparameter.
25. The device according to claim 24, characterized in that, The processing module is specifically used for: Determine from the model representation-hyperparameter database at least one model representation of an MP problem whose difference in feature similarity with the model representation of the target MP problem is less than a first value; Determine a target hyper-parameter corresponding to the target MP problem according to the hyper-parameter corresponding to the model representation of each MP problem in the model representation of the at least one MP problem.
26. The device according to claim 24 or 25, characterized in that, The first model is a graph neural network (GNN) model, and the processing module is specifically configured to: generate a corresponding bipartite graph according to the target MP problem, where the bipartite graph is a graph representation of the target MP problem; process the bipartite graph using the GNN model to obtain a model representation of the target MP problem.
27. A hyperparameter optimization device, characterized in that, The apparatus includes a memory and a processor; the memory stores code, and the processor is configured to obtain the code and execute the method according to any one of claims 1 to 10.
28. A hyperparameter recommendation device, characterized in that, The apparatus includes a memory and a processor; the memory stores code, and the processor is configured to obtain the code and execute the method according to any one of claims 11 to 13.
29. A computer-readable storage medium, characterized in that, It includes computer-readable instructions that, when running on a computer device, cause the computer device to execute the method according to any one of claims 1 to 10, or cause the computer device to execute the method according to any one of claims 11 to 13.
30. A computer program product, characterized in that, It includes computer-readable instructions that, when running on a computer device, cause the computer device to execute the method according to any one of claims 1 to 10, or cause the computer device to execute the method according to any one of claims 11 to 13.