Machine learning program, determination program, machine learning method, determination method, machine learning device, and determination device

By training a machine learning model to incorporate continuous relaxation and penalty terms, the method addresses the inefficiencies of existing combinatorial optimization techniques, enabling the simultaneous solution of multiple solutions and constraint management.

JP2025104790APending Publication Date: 2025-07-10FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023222858
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing combinatorial optimization methods struggle to find multiple solutions efficiently, particularly due to challenges with continuous relaxation annealing methods and the difficulty in optimizing multiple cost functions and tuning penalty coefficients.

Method used

A machine learning model is trained using a cost function that incorporates continuous relaxation into discrete optimization problems, allowing for the simultaneous optimization of multiple solutions by embedding discrete variables into a continuous matrix, and employing a penalty term to manage constraint violations.

Benefits of technology

This approach enables the simultaneous solution of multiple solutions in combinatorial optimization problems, improving efficiency and flexibility in obtaining diverse and accurate results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025104790000001_ABST
    Figure 2025104790000001_ABST
Patent Text Reader

Abstract

To provide a machine learning program, determination program, machine learning method, determination method, machine learning device and determination device, which enable derivation of multiple solutions to a combinatorial optimization problem.SOLUTION: A machine learning program provided herein causes a computer to perform processing for training a machine learning model through machine learning using a cost function, in which each element of a matrix obtained by relaxing discrete variables to be optimized into a continuous matrix is a discrete optimization problem, as a cost function in a search processing for making a search by incorporating continuous relaxation to the discretion optimization problem.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a machine learning program, a determination program, a machine learning method, a determination method, a machine learning device, and a determination device.

Background Art

[0002] Techniques for optimizing complex combinations have been disclosed (see, for example, Non-Patent Documents 1 and 2).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In combinatorial optimization, it is conceivable to search for an optimal solution by a continuous relaxation solution method using a machine learning model. However, it is difficult to solve multiple solutions.

[0005] On one side, the present invention aims to provide a machine learning program, a determination program, a machine learning method, a determination method, a machine learning device, and a determination device that can solve multiple solutions in a combinatorial optimization problem.

Means for Solving the Problems

[0006] In one aspect, the machine learning program causes a computer to train a machine learning model by machine learning using, as a cost function in a search process of incorporating continuous relaxation into a discrete optimization problem and performing a search, a cost function in which each element of a matrix obtained by relaxing discrete variables to be optimized into a continuous matrix becomes a discrete optimization problem.

[0007] In another aspect, the determination program causes a computer to execute a process of outputting a solution by embedding an optimization problem into a machine learning model trained by the execution of the above machine learning program.

Advantages of the Invention

[0008] Multiple solutions can be solved in a combinatorial optimization problem.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Best Mode for Carrying Out the Invention

[0010] Optimization problems exist at the sites of various industries, including the manufacturing industry and the distribution industry. In particular, the combinatorial optimization problem of optimizing combinations is one of the most important fields in the field of optimization. The combinatorial optimization problem is applied in various fields such as transportation, logistics, communication, and finance.

[0011] General solvers such as the Izing machine search for constraint-satisfied solutions using the penalty method, but it is difficult to solve depending on the penalty coefficient. In local transition algorithms such as the Izing machine, it is difficult to search for anything other than local solutions, so it is difficult to obtain multiple solutions at once.

[0012] With the development of information science, technologies aiming at high-speed solution of combinatorial optimization using machine learning have been developed. One of such technologies is the optimal solution method by the continuous relaxation annealing method. However, in the optimal solution method by the continuous relaxation annealing method, it is difficult to optimize multiple cost functions in one learning. Also, improvement in the efficiency of tuning the penalty coefficient and acquisition of various solutions are required.

[0013] First, the penalty method will be described. The constrained optimization problem is represented by the following formula (1). The following formula (1) is an optimization problem of minimizing C(x). Here, the equality constraint is represented by the following formula (2), and the inequality constraint is represented by the following formula (3).

Number

Number

Number

[0014] The penalty method is represented by the following formula (4). R represents the Euclidean space. In the following formula (4), γ is the penalty coefficient. If the penalty coefficient γ is too large, it is likely to be trapped in a local solution. On the other hand, if the penalty coefficient γ is too small, most of the search time will be spent in the infeasible region. Therefore, it is required to determine the penalty coefficient γ appropriately.

Number

[0015] Next, the continuous relaxation solution method using a machine learning model will be described. The continuous relaxation solution method is a method of relaxing and solving a discrete optimization problem into a continuous optimization problem for the parameter C∈C that characterizes the problem, instead of solving the discrete optimization problem. The parameter that characterizes the problem corresponds to, for example, the price of snacks and the capacity of the knapsack in the case of the knapsack problem.

[0016] As an example, the continuous relaxation solution method in the QUBO format will be described. QUBO stands for Quadratic Unconstrained Binary Optimization, which is a format that allows binary optimization without quadratic constraints.

[0017] Discrete optimization can be expressed as a loss function as in the following formula (5). Generally, in f(x;A), x represents the variable to be optimized, and A represents a constant that is not the optimization target. The variable to be optimized may omit the variable x to be optimized, as in E(·;C,γ) in the following formula (5). Also, E(x;C,γ) represents Cx + γ. As in the following formula (5), the variable x is a vector represented by 0 and 1 and has N elements.

Number

[0018] In the continuous relaxation solution method, the QUBO of the above formula (5) is relaxed into some simple form. As an example, the above formula (5) is relaxed into a hypercubic lattice. For example, the above formula (5) is expressed as the following formula (6). In the following formula, [0,1] N represents an N-dimensional hypercubic lattice that takes values of 0 or 1.

Number

[0019] However, even if the QUBO is continuously relaxed, the Loss Landscape may still become complex. Also, the relaxed optimal solution may be significantly different from the original optimal solution.

[0020] Next, it is conceivable to parameterize the parameter p with a GNN (Graph Neural Network) and optimize it using the loss function of the following formula (7). In the following formula (7), G of the optimization problem is converted into the embedding vector h G,φ G is the feature vector of the graph in the GNN. For example, in Figure 1, the feature vector G of the graph is converted into a graph embedding vector with the number of nodes N = 4 and the number of edges E = 4. For example, the optimization problem of p is reduced to the optimization regarding (θ, φ) of the following formula (7) characterized by the parameters of the GNN. Note that in the following formula (7), R E is the Euclidean space of E×N dimensions. p T is the transposed matrix of p.

Number

Number

[0021] However, in the above continuous relaxation solution method, only one solution can be obtained from the learning result. Also, the learning may fail depending on the random seed.

[0022] Here, the above-mentioned continuous relaxation solution method will be organized. In the above-mentioned continuous relaxation solution method, it is difficult to obtain a plurality of approximate solutions in one learning. Therefore, in Example 1, an example in which a plurality of solutions can be solved will be described. In addition, in the above-mentioned continuous relaxation solution method, there is a possibility that the solution performance may vary greatly depending on the penalty coefficient. Therefore, in Example 1, an example of searching for an appropriate penalty coefficient will also be described.

Example

[0023] First, the principle of this example will be described.

[0024] The discrete variable x ∈ {0, 1} N is relaxed to the continuous matrix P ∈ [0, 1] N×M . The loss function (cost function) of the following formula (9) is optimized so that each column P :,i of the continuous matrix P becomes the solution of the optimization problem. Here, P :,i means the i-th column vector of the continuous matrix P.

Equation

[0025] By relaxing the matrix as in the above formula (9), each column can be regarded as a decision variable, and it becomes possible to solve a plurality of different solutions simultaneously in parallel. In addition, by using the penalty term, it is also possible to obtain various solutions.

[0026] For example, by setting {C s} in the above formula (9) to {C,..., C}, for the same problem as the above formula (9), a plurality of penalty coefficients γ = (γ1,..., γ S ) can be optimized simultaneously as in the following formula (10). When using a GNN, the following formula (10) can be expressed as the following formula (11).

Equation

Equation

[0027] Here, the simultaneous Bayesian optimization of the penalty coefficient γ will be described. Using the penalty coefficient γ obtained by simultaneous solution, the loss function, and the penalty term (learning data) of the following formula (12), a surrogate function is learned. For example, Gaussian Process Regression can be used as the surrogate function.

Number

[0028] Next, an acquisition function is calculated using the surrogate function, and the next search range [γ min , γ max is proposed and divided into grids to sample γ at the S points again. Then, the solution result of γ for γ new is added. Note that the design method of the surrogate function varies depending on the situation. For example, when searching for a constraint satisfaction solution, it is constructed to focus on the γ-dependency of C(P new ). * :,S )

[0029] For example, for the initial penalty coefficient γ, the user may determine it. For example, an approximate hit can be made based on past solution results. Note that by using the penalty term R(p) of the following formula (13), the degree of continuity and discreteness may be controlled. For example, if the penalty coefficient γ < 0, a continuous solution will be preferably searched, and if the penalty coefficient γ > 0, a discrete solution will be preferably searched. As an example, as machine learning progresses, the penalty coefficient γ is gradually changed from a negative value to a positive value. Thereby, the penalty term of the following formula (13) changes from a case where the loss becomes smaller as the discrete vector p is more continuous to a case where the loss becomes larger as the discrete vector p is more continuous as machine learning progresses. The additional computational cost is only that the number of parameters from the layer before the last layer to the last layer increases by O(S) in FIG. 1.

Number

[0030] Subsequently, the above solution principle is verified. The maximum independent set problem on the graph G=(V,ε) is solved simultaneously with multiple penalty coefficients. The function to be optimized is given by the following formula (14).

Number

[0031] The solution of this embodiment is applied to MIS (Maximal Independent Set). MIS is a problem for finding the largest independent set on a certain graph. An independent set is a set of nodes such that if any other vertex is added, both sides of an edge are included in the set. The results are shown in Fig. 2. The horizontal axis represents the value of the penalty coefficient γ, and the vertical axis represents the size of the independent set. The larger the value on the vertical axis, the better the approximate solution is obtained. Since MIS is high, a good approximate solution is obtained in this embodiment. On the other hand, Violation is low. Also, MIS and Violation are obtained simultaneously in parallel. Note that Violation represents the number of variables that violate the constraint conditions.

[0032] It should be noted that by optimizing the following formula (15) for the same problem C and the same penalty coefficient γ as in the above formula (9), it is possible to obtain various solutions.

Number

[0033] Here, the following formula (16) is a penalty term for preventing each column from having the same value. In this case, for example, R(P) can be expressed as in the following formula (17). First, multiple solutions can be solved simultaneously in parallel by matrix relaxation. Furthermore, the penalty term enables the obtaining of various solutions.

Number

Number

[0034] For example, similar to the simultaneous solution of a plurality of penalty coefficients, it may be parameterized as in the following formula (18).

Equation

[0035] Subsequently, the apparatus configuration for realizing the above solution principle will be described. FIG. 3(a) is a functional block diagram showing the overall configuration of the information processing apparatus 100 according to the first embodiment. The information processing apparatus 100 is an optimization processing server or the like. As illustrated in FIG. 3(a), the information processing apparatus 100 functions as an optimization problem storage unit 10, a model parameter storage unit 20, a node embedding unit 30, a relaxation variable unit 40, a loss function calculation unit 50, a gradient storage unit 60, an approximate solution output unit 70, a constraint adjustment unit 80, and the like. The information processing apparatus 100 functions as a machine learning apparatus during machine learning and functions as a determination apparatus during determination. FIG. 3(b) is a block diagram showing the details of the constraint adjustment unit 80. As illustrated in FIG. 3(b), the constraint adjustment unit 80 functions as a coefficient group storage unit 81, a model parameter storage unit 82, an acquisition function calculation unit 83, an update coefficient storage unit 84, and the like.

[0036] FIG. 4 is a hardware configuration diagram of the information processing apparatus 100. As illustrated in FIG. 4, the information processing apparatus 100 includes a CPU 101, a RAM 102, a storage device 103, an input device 104, a display device 105, and the like.

[0037] The CPU (Central Processing Unit) 101 is a central processing unit. The CPU 101 includes one or more cores. The RAM (Random Access Memory) 102 is a volatile memory that temporarily stores programs executed by the CPU 101, data processed by the CPU 101, and the like. The storage device 103 is a non-volatile storage device. As the storage device 103, for example, a ROM (Read Only Memory), a solid state drive (SSD) such as a flash memory, a hard disk driven by a hard disk drive, or the like can be used. The storage device 103 stores a machine learning program and a determination program. The input device 104 is a device for the user to input necessary information, such as a keyboard and a mouse. The display device 105 is a display device that displays the approximate solution output by the approximate solution output unit 70 on the screen. By the CPU 101 executing an arithmetic program or a machine learning program, each part of the information processing device 100 is realized. Note that, as each part of the information processing device 100, hardware such as a dedicated circuit may be used.

[0038] FIG. 5 is a flowchart showing an example of the operation of the information processing device 100 during machine learning (training of a model by machine learning). As illustrated in FIG. 5, the loss function calculation unit 50 initializes the model and the constraint coefficients (step S1). Specifically, the loss function calculation unit 50 sets the model parameters stored in the model parameter storage unit 20 to predetermined initial values, and sets the penalty coefficient γ S stored in the coefficient group storage unit 81 to predetermined initial values.

[0039] Next, the node embedding unit 30 embeds the optimization problem (step S2). For example, in a problem using a graph, the node embedding unit 30 embeds the graph feature vector of the given optimization problem into the embedding vector h φ,GConvert it. Also, the relaxation variable part 40 sets relaxed mechanical variables parameterized by a neural network. Further, the node embedding part 30 uses the penalty coefficients stored in the coefficient group storage part 81. Thereby, the loss function represented by the above formula (9) is obtained.

[0040] Next, the loss function calculation unit 50 updates the model parameters by the gradient method (step S3). The model parameters are updated using the gradients stored in the gradient storage unit 60. When step S3 is executed for the first time, the model parameters are not updated.

[0041] Next, the loss function calculation unit 50 determines whether or not the convergence condition is satisfied (step S4). For example, it is determined whether or not the loss function of the above formula (9) stops decreasing below the specified value even when step S3 is repeatedly executed. If it is determined "No" in step S4, it is executed again from step S3.

[0042] If it is determined "Yes" in step S4, the loss function calculation unit 50 determines whether or not the constraints are satisfied (step S5). For example, the loss function calculation unit 50 checks whether the constraint conditions of the constrained optimization problem are satisfied. For example, when the solver performs the solution by the penalty method, a constraint violation may occur depending on the adjustment of the penalty coefficient. Therefore, if a constraint violation occurs, it is determined that the constraint conditions of the constrained optimization problem are not satisfied.

[0043] If it is determined "No" in step S5 (in other words, if it is determined that there is a constraint violation), the acquisition function calculation unit 83 performs learning of the surrogate function using the above formula (12) (step S6). In this case, the acquisition function calculation unit 83 uses the surrogate function model parameters stored in the model parameter storage unit 82. Further, the acquisition function calculation unit 83 uses the loss function and the penalty coefficient γ obtained by the most recent execution of step S3. Note that the loss function and the penalty coefficient γ obtained by the most recent execution of step S3 are stored in the coefficient group storage unit 81.

[0044] Next, the acquisition function calculation unit 83 calculates an acquisition function using the surrogate function learned in step S6 (step S7).

[0045] Next, the acquisition function calculation unit 83 proposes the following search range [γ min , γ max , divides it into grids, and samples the penalty coefficient γ new of the S point. Then, the solution result of γ S is added to γ new (step S8). The penalty coefficient obtained in step S8 is stored in the update coefficient storage unit 84. Then, the process returns to step S3 and is executed again. When executing step S3 for the second time and later, the penalty coefficient γ new stored in the update coefficient storage unit 84 is used.

[0046] If it is determined as "Yes" in step S5, the execution of the flowchart ends. In this case, the model parameter storage unit 20 stores the model parameters when the loss function is the smallest.

[0047] Through the machine learning shown in FIG. 5, a machine learning model is obtained such that the loss function of the above formula (9) becomes the minimum. The machine learning model (model parameters) is stored in the model parameter storage unit 20.

[0048] FIG. 6 is a flowchart showing an example of the operation of the information processing apparatus 100 when outputting an approximate solution to an optimization problem using the result of the machine learning model obtained by the machine learning shown in FIG. 5. As illustrated in FIG. 6, the node embedding unit 30 embeds the optimization problem (step S11).

[0049] Next, the approximate solution output unit 70 acquires the output of the machine learning model (step S12).

[0050] Next, the approximate solution output unit 70 performs threshold processing on the optimal solution output by the machine learning model (step S13). For example, for each value output by the machine learning model, a threshold is set for binarization. For example, when converting each value into two values of 0 and 1, the threshold is set to 0.5 or the like. Values greater than 0.5 are set to 1, and values less than 0.5 are set to 0.

[0051] In the above embodiment, an optimization problem using a graph as an optimization target has been described. The optimization problem using a graph is not particularly limited, and examples include an energy transport problem. The above embodiment can also be applied to an optimization problem that does not use a graph as an optimization target. The optimization problem that does not use a graph is not particularly limited, and examples include a corporate scheduling problem.

[0052] In the above embodiment, the loss function calculation unit 50 is an example of an execution unit that executes a process of machine learning the model using, as a cost function in a search process of incorporating continuous relaxation into a discrete optimization problem and performing search, a cost function in which each element of a matrix obtained by relaxing discrete variables to be optimized into a continuous matrix becomes a discrete optimization problem. The approximate solution output unit 70 is an example of an output unit that outputs a solution by embedding an optimization problem in the model obtained by executing the machine-learned machine learning program.

[0053] As described above in detail, the embodiments of the present invention have been described in detail. However, the present invention is not limited to such specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims. (Appendix 1) A machine learning program causing a computer to execute a process of training a machine learning model by machine learning using, as a cost function in a search process of incorporating continuous relaxation into a discrete optimization problem and performing search, a cost function in which each element of a matrix obtained by relaxing discrete variables to be optimized into a continuous matrix becomes a discrete optimization problem. (Appendix 2) The cost function includes a penalty term representing the constraints in the search process, A machine learning program according to Appendix 1, wherein when training the machine learning model, a penalty coefficient of the penalty term is learned. (Appendix 3) A machine learning program according to Appendix 1, wherein the discrete optimization problem is represented in a QUBO format. (Appendix 4) A machine learning program according to Appendix 1, wherein when training the machine learning model, a loss term corresponding to the degree of continuity and discreteness of the variables to be optimized is used, and the loss term is changed according to the progress of the search process. (Appendix 5) A machine learning program according to Appendix 4, wherein the computer is caused to execute a process of changing the loss term from one in which the loss decreases as the variable is more continuous to one in which the loss increases as the variable is more continuous as the search process progresses. (Appendix 6) The computer, A determination program, wherein a process of outputting a solution by embedding an optimization problem into a machine learning model trained by executing the machine learning program according to any one of Appendices 1 to 5 is executed. (Appendix 7) The computer, A machine learning method, wherein in a search process of performing a search by incorporating continuous relaxation into a discrete optimization problem, a machine learning model is trained by using, as a cost function in the search process, a cost function in which each element of a matrix obtained by relaxing a discrete variable to be optimized into a continuous matrix becomes a discrete optimization problem. (Appendix 8) The cost function includes a penalty term representing the constraints in the search process, A machine learning method according to Appendix 7, wherein when training the machine learning model, a penalty coefficient of the penalty term is learned. (Appendix 9) The machine learning method according to Appendix 7, wherein the discrete optimization problem is represented in the QUBO format. (Appendix 10) The machine learning method according to Appendix 7, wherein when training the machine learning model, a loss term corresponding to the degree of continuity and discreteness of the variables to be optimized is used, and the loss term is changed according to the progress of the search process. (Appendix 11) The machine learning method according to Appendix 10, wherein as the search process progresses, the computer executes a process of changing the loss term from one in which the loss decreases as the variable is more continuous to one in which the loss increases as the variable is more continuous. (Appendix 12) The computer A determination method characterized by executing a process of outputting a solution by embedding an optimization problem into a machine learning model trained by executing a machine learning program according to any one of Appendices 7 to 11. (Appendix 13) A machine learning apparatus comprising an execution unit that executes a process of training a machine learning model by machine learning using, as a cost function in a search process that performs search by incorporating continuous relaxation into a discrete optimization problem, a cost function in which each element of a matrix obtained by relaxing a discrete variable to be optimized into a continuous matrix becomes a discrete optimization problem. (Appendix 14) The cost function includes a penalty term representing a constraint in the search process, The execution unit learns a penalty coefficient of the penalty term when training the machine learning model. The machine learning apparatus according to Appendix 13. (Appendix 15) The discrete optimization problem is represented in the QUBO format. The machine learning apparatus according to Appendix 13. (Appendix 16) The execution unit uses a loss term corresponding to the degree of continuity and discreteness of the variables to be optimized when training the machine learning model, and changes the loss term according to the progress of the search process. The machine learning apparatus according to Appendix 13. (Supplementary Note 17) As the search process progresses, the execution unit changes the loss term from one in which the loss decreases as the variable becomes more continuous to one in which the loss increases as the variable becomes more continuous. The machine learning device according to Supplementary Note 16, characterized by this. (Supplementary Note 18) A determination device comprising an output unit that outputs a solution by embedding an optimization problem in a machine learning model trained by the machine learning device according to any one of Supplementary Notes 13 to 17, characterized by this.

Explanation of Signs

[0054] 10 Optimization problem storage unit 20 Model parameter storage unit 30 Node embedding unit 40 Relaxed variable unit 50 Loss function calculation unit 60 Gradient storage unit 70 Approximate solution output unit 80 Constraint adjustment unit 81 Coefficient group storage unit 82 Model parameter storage unit 83 Acquisition function calculation unit 84 Update coefficient storage unit 100 Information processing device 101 CPU 102 RAM 103 Storage device 104 Input device 105 Display device

Claims

1. A computer, as a cost function in a search process for incorporating a continuous relaxation into a discrete optimization problem and performing a search, by machine learning using a cost function in which each element of a matrix obtained by relaxing an optimization target discrete variable into a continuous matrix becomes a discrete optimization problem, to execute a process of training a machine learning model. A machine learning program characterized by that.

2. The cost function includes a penalty term representing a constraint in the search process, The machine learning program according to claim 1, characterized in that, when training the machine learning model, a penalty coefficient of the penalty term is learned.

3. The machine learning program according to claim 1, characterized in that the discrete optimization problem is represented in a QUBO format.

4. The machine learning program according to claim 1, characterized in that, when training the machine learning model, a loss term corresponding to the degree of continuity and discreteness of a variable to be optimized is used, and the loss term is changed according to the progress of the search process.

5. The computer is caused to execute a process of changing the loss term from one in which the loss becomes smaller as the variable is more continuous to one in which the loss becomes larger as the variable is more continuous as the search process progresses. The machine learning program according to claim 4, characterized by that.

6. A computer, A determination program characterized by causing the computer to execute a process of outputting a solution by embedding an optimization problem into a machine learning model trained by executing the machine learning program according to any one of claims 1 to 5.

7. A machine learning method, characterized in that a computer executes a process of training a machine learning model by machine learning using a cost function in which each element of a matrix obtained by relaxing an optimization target discrete variable into a continuous matrix becomes a discrete optimization problem, as a cost function in a search process for incorporating a continuous relaxation into a discrete optimization problem and performing a search.

8. A computer, A determination method characterized in that the computer executes a process of outputting a solution by embedding an optimization problem into a machine learning model trained by executing the machine learning program according to claim 7.

9. ​ As a cost function in a search process that incorporates a continuous relaxation into a discrete optimization problem and performs a search, a machine learning apparatus is provided with an execution unit that executes a process of training a machine learning model by machine learning using a cost function in which each element of a matrix obtained by relaxing an optimization target discrete variable into a continuous matrix becomes a discrete optimization problem.

10. A determination apparatus characterized by comprising an output unit that outputs a solution by embedding an optimization problem into a machine learning model trained by the machine learning apparatus according to Claim 9.