Information processing device, method, and program
The information processing apparatus enhances accuracy and efficiency in solving large-scale combinatorial optimization problems by using a dual search approach to refine solutions based on partial derivative comparisons.
Patent Information
- Application Number
- JP2024002774
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-24
AI Technical Summary
Existing methods struggle to achieve high accuracy solutions or require excessive processing time for large-scale combinatorial optimization problems.
An information processing apparatus with a first search unit and a second search unit, utilizing a first and second searcher to find and refine solutions based on the signs of tentative solutions and partial derivative values of the objective function.
Improves calculation accuracy and reduces processing time for combinatorial optimization problems by refining incomplete dimensions as sub-problems.
Smart Images

Figure 2025109070000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to an information processing apparatus, method, and program.
Background Art
[0002] When solving a large-scale combinatorial optimization problem, there is a problem that a solution with high accuracy cannot be obtained, or an enormous processing time is required to obtain a solution with high accuracy.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The problem to be solved by the present invention is to provide an information processing apparatus, method, and program that enable improvement in the calculation accuracy of a combinatorial optimization problem and / or reduction in processing time.
Means for Solving the Problems
[0005] The information processing apparatus according to the embodiment includes a first search unit and a second search unit. The first search unit searches for a first tentative solution of a combinatorial optimization problem whose objective function is differentiable using a first searcher that solves the entire combinatorial optimization problem. The second search unit outputs a final solution of the combinatorial optimization problem based on a comparison of the signs of the first tentative solution and the partial derivative value of the objective function.
Brief Description of the Drawings
[0006]
Fig. 1
Fig. 2
Fig. 3
Fig. 4
Fig. 5
Fig. 6
Fig. 7
Fig. 8
Fig. 9
Fig. 10
Fig. 11
Embodiment for Carrying Out the Invention
[0007] Hereinafter, an information processing apparatus, method, and program according to this embodiment will be described with reference to the drawings.
[0008] FIG. 1 is a diagram showing an example of the hardware configuration of the information processing apparatus 100. As shown in FIG. 1, the information processing apparatus 100 is a computer having a processor 1, a storage device 2, an input device 3, a display device 4, and a communication device 5. The transmission and reception of data and various signals among the processor 1, the storage device 2, the input device 3, the display device 4, and the communication device 5 are performed via a bus.
[0009] The processor 1 is an integrated circuit that controls the overall operation of the information processing apparatus 100. For example, the processor 1 has a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), and / or an FPU (Floating-Point Unit). The processor 1 may include an internal memory and an I / O interface. The processor 1 executes various processes by interpreting and calculating programs pre-stored in the storage device 2 and the like. Note that part or all of the processor 1 may be realized by hardware such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0010] The storage device 2 is a volatile memory and / or a non-volatile memory that stores various data. For example, the storage device 2 stores data and setting values used when the processor 1 executes various processes, data generated by various processes in the processor 1, and the like. The storage device 2 is composed of a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), an integrated circuit storage device, and the like. Note that the storage device 2 may have a non-temporary computer-readable storage medium that stores the program executed by the processor 1.
[0011] The input device 3 receives inputs of various operations from the operator. As the input device 3, a keyboard, a mouse, various switches, a touch pad, a touch panel display, or the like can be used. An electrical signal (hereinafter, an operation signal) corresponding to the received input of the operation is supplied to the processor 1.
[0012] The display device 4 displays various data in accordance with the control by the processor 1. As the display device 4, a CRT (Cathode-Ray Tube) display, a liquid crystal display, an organic EL (Electro Luminescence) display, an LED (Light-Emitting Diode) display, a plasma display, or any other display can be appropriately used. The display device 4 may be a projector.
[0013] The communication device 5 includes a communication interface such as a network interface card (NIC) for performing data communication with various devices connected to the information processing apparatus 100 via a network. Note that an operation signal may be supplied from a computer connected via the communication device 5 or an input device included in the computer, or various data may be displayed on a display device or the like included in the computer connected via the communication device 5. However, for the sake of simplifying the following description, unless otherwise specified, the source of the operation signal is the input device 3, and the destination for displaying various data is the display device 4. The input device 3 can be replaced by a computer connected via the communication device 5 or an input device included in the computer, and the display device 4 can be replaced by a display device or the like included in the computer connected via the communication device 5.
[0014] The information processing apparatus 100 does not necessarily need to include all of the processor 1, the storage device 2, the input device 3, the display device 4, and the communication device 5. A part of the storage device 2, the input device 3, the display device 4, and the communication device 5 may not be provided as necessary. The information processing apparatus 100 may be provided with any additional hardware device useful for executing the processing according to the present embodiment. The information processing apparatus 100 does not necessarily need to be physically configured by one computer, and may be configured by a computer system having a plurality of computers communicably connected via a wired or network line or the like. The allocation of a series of processing according to the present embodiment to the plurality of processors 1 respectively mounted on the plurality of computers can be arbitrarily set. All the processors 1 may execute all the processing in parallel, or a specific processing may be assigned to one or a part of the processors 1, and a series of processing according to the present embodiment may be executed as a whole computer system.
[0015] Processor 1 solves a combinatorial optimization problem. The combinatorial optimization problem according to this embodiment is directed to a problem of obtaining a discretely expressed solution that optimizes a partially differentiable objective function. Such an optimization problem includes an Ising problem. Specifically, the Ising problem according to this embodiment addresses any one of a linear regression problem, a binary classification problem, and a Traveling Salesman Problem (TSP). Here, the outline of the Ising problem will be described.
[0016] The Ising problem is a combinatorial optimization problem in which explanatory variables are composed of lattice points that take two coordination states, and the objective function is set to the Ising energy of the plurality of lattice points, and it is a lattice model that considers only two-dimensional interactions. The Ising model is one method for formulating a combinatorial problem.
[0017] Ising energy E ising is composed of the following formula (1). The vector s representing the spin of each lattice takes a binary state of -1 and +1, J is a coupling coefficient matrix representing the interaction of each lattice, and h is a vector representing an external magnetic field.
[0018]
Equation
[0019] The Ising energy can be rewritten in a form called a Quadratic Unconstrained Binary Optimization (QUBO) problem without quadratic constraints as shown in the following formula (2). The vector w takes a binary state represented by binary variables of 0 and 1.
[0020]
Equation
[0021] There is compatibility between the above equations (1) and (2) of the Ising energy. For example, by converting Q used in the QUBO equation to s in the following equation (3) and then converting it to J shown in the following equation (4) and h shown in the following equation (5).
[0022] [Number]
[0023] An optimization device or optimization method that outputs a vector w or s that minimizes the Ising energy is called an Ising machine.
[0024] In the Ising problem, it is possible to extract a sub-problem (sub-QUBO) by changing only some of the spins and fixing the other spins. For example, as shown in the following equation (6), when the vector w is divided into a part w1 that wants to be changed and a part w2 that wants to be fixed, since Q is a symmetric matrix, as shown in the following equation (7), E ising can be rewritten.
[0025] [Number]
[0026] w2 in equation (7) T Q 22 Since w2 is a fixed constant, E ising and the sub-problem E of the following equation (8) ising ´ can be regarded as equivalent.
[0027] [Number]
[0028] The following describes the Simulated Bifurcation (SB) machine, which is a type of algorithm for solving the Ising problem. The following equations (9) to (13) hold between the momentum y of a particle, the vector x corresponding to the spin, and the Hamiltonian H according to the classical bifurcation theory.
[0029]
number
[0030] The coefficient p(t) increases monotonically according to the number of updates during calculation. For example, p(t)=0 is the initial value, and it increases monotonically so that p(t)=1 at the end. The coefficient a(t) increases with p(t), and can be set to a(t)=√p(t), for example. The coefficient D corresponds to detuning, and the coefficient K corresponds to the positive Kerr coefficient. f i is the external force, z i is the last term of the Hamiltonian H, x i In order to ensure the accuracy of the calculation, it is possible to set the coefficient c to a value close to the reciprocal of the largest eigenvalue of the matrix J. For example, it is possible to set c = 0.5 · D√(N / 2n). N is the number of dimensions of the Ising problem, and n is an arbitrary integer.
[0031] By treating the variable vectors x and y as continuous values, the ordinary differential equations shown in equations (9) and (10) can be solved numerically. For example, by using the symplectic Euler method, the differential equations can be rewritten as discrete recurrence equations as shown in equations (14) and (15) below.
[0032]
number
[0033] t is the time, and Δt is the time step (time increment). If the time step Δt and the number of steps are set as in equations (14) and (15), x and y can be updated iteratively. For example, if Δt = 1, Δt may be omitted from equations (14) and (15).
[0034] In the simulated branching machine, after reaching a predetermined number of updates, the sign of x i can be used to obtain the spin s i of the Ising problem. For example, when x i > 0, s i = +1, and when x i < 0, s i = -1.
[0035] Therefore, in order to reduce the error, the simulated branching machine can be improved. Specifically, when the variable x i > 1 due to an update, the variable x i is set to +1. Also, when x i < -1 due to an update, the variable x i is set to -1. This makes it possible to approximate the spin s i with higher accuracy using the variable x i .
[0036] Next, the process of formulating a linear regression problem as an Ising problem will be described. In a linear regression problem Y = X·β having an objective variable Y, data X, and an explanatory variable β, to obtain an estimated value β^ of β, the loss function represented by the following equation (16) is minimized.
[0037]
Equation
[0038] Although β^ is a continuous value, the continuous value β^ can be represented by the spin s through the product of the spin s and the matrix P. Therefore, the original problem minL(β^) can be converted into an Ising problem of minL(s) as shown in the following equation (17).
[0039]
Number
[0040] Next, the process of formulating a binary classification problem as an Ising problem will be described. In a binary classification problem of predicting a categorical variable target variable Y using multidimensional data X and explanatory variable β, in order to estimate the estimated value β^ of the explanatory variable β, the loss function represented by the following formula (18) is minimized to calculate the mediating variable b. When there is no kernel function, the estimated value β^ can be calculated from the obtained mediating variable b according to the following formula (19).
[0041]
Number
[0042] Similar to the case of formulating a linear regression problem as an Ising problem, the continuous-valued mediating variable b can be represented by the binary variable w by the product of the vector w and the matrix P. Also, the value of the mediating variable b in the binary classification problem is restricted to 0 ≤ b ≤ 1. Therefore, the minL(b) of the original problem can be converted into the Ising problem of minL(w) as shown in the following formula (20).
[0043]
Number
[0044] Also, when using a kernel function, X in the above formula T X can be replaced with a Gram matrix calculated based on the kernel function.
[0045] Next, a process of formulating the traveling salesman problem (TSP) as an Ising problem will be described. The traveling salesman problem is a problem of searching for the shortest route among a plurality of cities on the premise of making a round trip among them. Assuming the number of cities is N, the traveling salesman problem can be converted into a QUBO equation for solving an N×N-dimensional binary variable w represented by the following equation (21). The vector of the k-th row of the solution w means which order the k-th city is visited. For example, when obtaining solutions w1=(0,1,0), w2=(1,0,0), and w3=(0,0,1) as 3×3-dimensional w, w means a route of first visiting the second city, then the first city, and finally the third city.
[0046]
Number
[0047] C dist corresponds to the travel distance, and d l,l´ means the relative distance between the l-th and l'-th cities. The two parts in the constraint term C const guarantee that only one city is visited at a time and that all cities are visited. A is a coefficient for controlling the strength of the constraint term and can be set to, for example, A = 0.5.
[0048] The processor 1 according to the present embodiment solves a combinatorial optimization problem represented by the above Ising problem. Specifically, the processor 1 has functional configurations such as an acquisition unit 11, a first search unit 12, a second search unit 13, and a display control unit 14.
[0049] The acquisition unit 11 acquires various data related to the combinatorial optimization problem. For example, the acquisition unit 11 acquires explanatory variables, target data, and objective variables related to the combinatorial optimization problem. The explanatory variable is a parameter to be updated in the combinatorial optimization problem, and the target data is a parameter treated as a constant that is not an update target. The objective variable is a parameter calculated based on the explanatory variable and the target data. As described above, the combinatorial optimization problem according to the present embodiment is an Ising problem in which the objective function is set to the Ising energy of a plurality of lattice points. As a solution, a vector (hereinafter, solution vector) representing a combination of a plurality of quantum states corresponding to a plurality of lattice points is output. The quantum state takes a binary state such as -1 and +1, 0 and +1, 0 and -1. The solution vector has a number of dimensions corresponding to the number of lattice points. The solution vector is calculated from the explanatory variables. The objective function is designed to be partially differentiable. The partial derivative means the derivative of the objective function with respect to the quantum state.
[0050] The first search unit 12 searches for a first tentative solution of the combinatorial optimization problem in which the objective function is partially differentiable, using a first searcher 15 that solves the entire problem of the combinatorial optimization problem. The entire problem is a problem of solving for all dimensions that make up the solution. The first tentative solution means the solution vector obtained by the first search unit 12. The first search unit 12 inputs the target data, explanatory variables, and objective variables to the first searcher 15. The first searcher 15 means an algorithm used to solve the entire problem. The first searcher 15 updates the estimated value of the explanatory variable so as to optimize the objective function that depends on the input explanatory variables, target data, and objective variables, and outputs a solution vector based on the estimated value when it is evaluated as being optimized. The first search unit 12 acquires the output solution vector as the first tentative solution.
[0051] The second search unit 13 outputs the final solution to the combinatorial optimization problem based on the comparison of the signs of the first tentative solution and the partial differential value of the objective function. The final solution means the solution vector finally obtained by the second search unit 13. More specifically, the second search unit 13 extracts a first dimension, which is the dimension in which the sign of the partial differential value corresponding to the first tentative solution matches, from the first tentative solution, and uses a second searcher 16 that solves a sub-problem of the combinatorial optimization problem to calculate the combinatorial optimization problem with the first dimension as the sub-problem and search for the final solution. The sub-problem is a problem that is solved by limiting it to some of the dimensions among all the dimensions that make up the solution vector. Here, the partial differential value corresponding to the first tentative solution means a value obtained by substituting the quantum state corresponding to the first tentative solution into the partial differential of the objective function regarding the quantum state. The partial differential value is represented by a vector having a dimension corresponding to the dimension of the solution vector. The vector of the partial differential value is also denoted as the partial differential vector or the gradient vector.
[0052] When the signs of the first tentative solution and the partial differential value do not match in all dimensions, the second search unit 13 outputs the first tentative solution as the final solution. On the other hand, when the signs of the first tentative solution and the partial differential value do not match in all dimensions, the second search unit 13 extracts the first dimension from the first tentative solution. More specifically, the second search unit 13 uses the second searcher 16 to calculate the combinatorial optimization problem with the first dimension as the sub-problem and outputs a second tentative solution. The second tentative solution means the solution vector of the sub-problem obtained by the second search unit 13. The solution vector of the sub-problem has the same number of dimensions as the dimension of the sub-problem. Specifically, the second search unit 13 inputs the explanatory variable, target data, and objective variable regarding the first dimension to the second searcher 16. The second searcher 16 updates the estimated value of the explanatory variable so as to optimize the objective function that depends on the input explanatory variable, target data, and objective variable, and outputs a solution vector based on the estimated value when it is evaluated that the optimization has been achieved. The explanatory variable is updated so as to optimize the objective function, and the solution vector when the objective function has been optimized is output. The solution vector means the second tentative solution.
[0053] The second search unit 13 extracts a second dimension, which is a dimension in which the sign matches the partial derivative value corresponding to the second tentative solution, from the second tentative solution, and uses the second searcher 16 to calculate a combinatorial optimization problem with the second dimension as a subproblem to search for a final solution. Here, the partial derivative value corresponding to the second tentative solution means a value obtained by substituting the quantum state corresponding to the second tentative solution into the partial derivative of the objective function regarding the quantum state. The search for the final solution based on the second dimension will be described in more detail. The second search unit 13 generates an integrated solution of the second tentative solution and the remaining dimensions of the first tentative solution, and if the signs of the integrated solution and the partial derivative value corresponding to the second tentative solution do not match for all dimensions, outputs the integrated solution as the final solution, and if the signs of the integrated solution and the partial derivative value corresponding to the second tentative solution do not all mismatch for all dimensions, extracts the second dimension from the second tentative solution. For the extracted second dimension, it is again used for solving the subproblem using the second searcher 16, and subsequent processes such as determination of sign matching or mismatch follow.
[0054] The Ising problem according to this embodiment is any one of a linear regression problem, a classification problem, and a traveling salesman problem. The first searcher 15 and the second searcher 16 are Ising machines that output a vector that minimizes the Ising energy, which is the objective function. Each of the first searcher 15 and the second searcher 16 is a simulated annealing machine or a simulated bifurcation machine. The first searcher 15 and the second searcher 16 may use the same algorithm or different algorithms.
[0055] In this embodiment, it is assumed that the first searcher 15 and the second searcher 16 are algorithms processed by a processor 1 of a digital computer such as the information processing apparatus 100. However, the first searcher 15 and the second searcher 16 may be hardware implementations using physical phenomena. As hardware implementations using physical phenomena, quantum computers such as quantum annealers, coherent Ising machines, and quantum branching machines can be used. A quantum annealer realizes quantum annealing using a superconducting circuit. A coherent Ising machine utilizes the oscillation phenomenon of a network formed by optical parametric oscillators. A quantum branching machine utilizes the quantum mechanical branching phenomenon in a network of parametric oscillators having the Kerr effect.
[0056] The display control unit 14 displays various information on a display device. For example, the display control unit 14 displays the final solution.
[0057] Hereinafter, the operation of the information processing apparatus 100 according to this embodiment will be described in detail.
[0058] First, the relationship between the sign of the solution vector and the partial derivative value of the objective function will be described. In the Ising problem, the Ising energy E ising is used as the objective function, and the solution vector is the spin s of each lattice. Therefore, the partial derivative value of the Ising energy E ising which is the objective function, with respect to the spin s which is the solution vector, that is, the gradient vector g, can be expressed by the following equation (21) based on equation (1).
[0059]
Equation
[0060] FIG. 2 is a diagram showing the relationship between the Ising energy E ising and the spin s. The vertical axis in FIG. 2 is the Ising energy E isingIt is defined as specified, and the horizontal axis is defined as spin s. The solid line in Figure 2 represents the optimization process of the spin with the final solution being -1, and the dotted line represents the optimization process of the spin with the final solution being +1. As shown in Equation (21) and Figure 2, since each dimension of spin s is monotonically increasing or decreasing, in an ideal state of optimization, the sign of the final solution s for the same dimension should match the inverse sign of the gradient vector g. For example, in the case of the dimension where spin s = -1, the corresponding dimension of the gradient vector g should be +, which is the inverse sign of the spin. If it is not the inverse sign, that is, if the signs match, it means that the optimization for that dimension is incomplete. On the contrary, if it is the inverse sign, that is, if the signs do not match, it means that the optimization for that dimension is good. The same applies to the dimension where spin s = +1.
[0061] The reason why the dimension with the matching sign has incomplete optimization will be explained from the general Karush-Kuhn-Tucker condition. Consider the constrained minimization problem represented by the following Equation (22). Here, x is the variable, E(x) is the function to be minimized, and g(x) is the constraint condition.
[0062]
Equation
[0063] In the case of the combinatorial optimization problem, since x is in a binary state of -1 or +1, the constraint condition is represented by the following Equation (23).
[0064]
Equation
[0065] In this case, the steady-state condition of the final solution is represented by the following Equation (24).
[0066]
Equation
[0067] If x in the formula (24) is replaced with the spin s, the same conclusion as the gradient formula (21) of the Ising energy is obtained. Therefore, the fact that "the sign of the final solution s should match the inverse sign of the gradient vector g" can also be understood from the Karush-Kuhn-Tucker conditions.
[0068] On the premise of this understanding, the second search unit 13 compares the sign of the solution of the objective function with the sign of the partial derivative value, extracts the dimension where the signs match as the dimension with incomplete optimization, and re-optimizes the extracted dimension as a sub-problem.
[0069] Next, the processing procedure of the optimization process by the processor 1 will be described. FIG. 3 is a diagram illustrating the processing procedure of the optimization process. FIG. 4 is a diagram schematically showing the transition of the sign of the solution and the sign of the partial derivative in the optimization process shown in FIG. 3. For convenience of explanation, FIG. 4 assumes that the solution vector and the partial derivative value are 5-dimensional vectors. The shaded circles among the vectors represent the dimensions with the sign +, and the unshaded circles represent the dimensions with the sign -.
[0070] As shown in FIG. 3, the acquisition unit 11 acquires the target data, the explanatory variable, and the objective variable (step SA1). The target data and the explanatory variable have the same number of dimensions and sample size. The number of dimensions and the sample size can be set to any number.
[0071] As an example, the processing of step SA1 when solving a linear regression problem as an Ising problem will be described. The acquisition unit 11 randomly generates target data X with a dimension of 1000 × sample size of 2000 from a normal distribution with a mean of 0 and a variance of 1. Further, the acquisition unit 11 randomly generates an explanatory variable β of 1000 dimensions having the same number of dimensions as the target data X from a normal distribution with a mean of 0 and a variance of 0.5. The acquisition unit 11 selects the explanatory variable β whose absolute value does not exceed 1. The acquisition unit 11 calculates the objective variable Y according to Y = X·β. The acquisition unit 11 uses the conversion formula represented by the following formula (25) to convert the explanatory variable β of the linear regression problem into a vector s of spins. p is a sequence of 1, 2, 4, 8, 16,... scaled to an appropriate threshold.
[0072]
Number
[0073] As another example, the processing of step SA1 when solving a binary classification problem as an Ising problem will be described. The acquisition unit 11 randomly generates data X with a dimensionality of 100×sample size 2000 from a normal distribution with a mean of 0 and a variance of 1. The acquisition unit 11 randomly generates an explanatory variable β of 100 dimensions having the same dimensionality as the target data X from a normal distribution with a mean of 0 and a variance of 1. The acquisition unit 11 calculates the target variable Y according to Y = X·β. When Y>0, the acquisition unit 11 sets Y = 1, and when Y<0, the acquisition unit 11 sets Y = -1. The acquisition unit 11 uses the conversion formula represented by the following formula (26) to convert the mediating variable b of the binary classification problem into a binary vector w. p is a sequence of 1, 2, 4, 8, 16, ··· scaled to an appropriate threshold value.
[0074]
Number
[0075] When step SA1 is performed, the first search unit 12 uses the first searcher 15 to calculate a combinatorial optimization problem and search for a first tentative solution regarding the overall problem (step SA2). The Ising problem according to the present embodiment is any one of a linear regression problem, a binary classification problem, and a traveling salesman problem. The first searcher 15 is an Ising machine that outputs a vector that minimizes the Ising energy. The first searcher 15 is a simulated annealing machine or a simulated bifurcation machine. When using a simulated bifurcation machine, it is preferable to set the decoupling coefficient D shown in formulas (9) to (12) as D = 1, the car coefficient K as K = 0, and the coefficient c as c = D / (2σ·√n). Here, N is the dimensionality of the Ising problem, and σ is the standard deviation value of all elements of the matrix J in the Ising problem.
[0076] Specifically, the first searcher 15 calculates an objective function value based on the target data, explanatory variables, and objective variables acquired in step SA1, and updates the explanatory variables so that the objective function value is minimized. The first searcher 15 repeats the calculation of the objective function value and the update of the explanatory variables until a predetermined stop condition is satisfied. The predetermined stop condition can be set, for example, that the number of calculation steps exceeds a threshold value, or that the objective function value or the explanatory variable value exceeds a predetermined value. When the predetermined stop condition is satisfied, the first searcher 15 outputs, as the first tentative solution, the solution vector corresponding to the objective function value at the calculation step when the condition is satisfied. Here, a calculation step means a processing step of performing the calculation of the objective function value and the update of the explanatory variables once each.
[0077] For example, when using a simulated bifurcation machine as the first searcher 15, the first searcher 15 executes the sum-of-products operation of the coupling coefficient matrix and the vector of the Ising problem and the update of the vector based on the result of the sum-of-products operation until the number of predetermined calculation steps is exceeded. When the number of executions of the sum-of-products operation and the vector update exceeds the number of predetermined calculation steps, the first searcher 15 outputs, as the first tentative solution, the vector at the calculation step when the number of predetermined calculation steps is exceeded. In this case, the number of calculation steps for solving the combinatorial optimization problem may be set to a value proportional to the number of dimensions of the Ising problem, or may be set to the number of steps determined by the user or the like.
[0078] When step SA2 is performed, the second search unit 13 determines whether there is a dimension in which the sign of the first tentative solution obtained in step SA2 matches the sign of the partial derivative value (step SA3). Specifically, the second search unit 13 calculates the partial derivative value by substituting the first tentative solution into the partial derivative of the objective function with respect to the spin. The partial derivative value has the same number of dimensions as the first tentative solution (spin). Then, the second search unit 13 compares the sign of the calculated partial derivative value with the sign of the first tentative solution obtained in step SA2 for each same dimension. The dimension in which the signs match is the dimension of incomplete optimization, and the dimension in which the signs do not match is the dimension of complete optimization.
[0079] In the case of the linear regression problem, the objective function is the loss function L(s) shown in equation (17), and for the partial derivative, the partial derivative dL(s) / ds of the loss function L(s) with respect to the spin s is used. In the case of the binary classification problem, the objective function is the loss function L(w) shown in equation (19), and for the partial derivative, the partial derivative dL(w) / dw of the loss function L(w) with respect to the binary variable w is used. In the case of the traveling salesman problem, the objective function is the cost function C_tsp(w) shown in equation (20), and for the partial derivative, the partial derivative dC_tsp(w) / dw of the cost function C_tsp(w) with respect to the binary variable w is used.
[0080] When it is determined that there is no dimension in which the sign of the first tentative solution matches the sign of the partial derivative value (step SA3: NO), since all dimensions of the first tentative solution obtained in step SA2 are non-matching dimensions and are fully optimized dimensions, the second search unit 13 outputs the first tentative solution as the final solution (step SA9).
[0081] On the other hand, when it is determined in step SA3 that there is a dimension in which the sign of the first tentative solution matches the sign of the partial derivative value (step SA3: YES), the second search unit 13 extracts the matching dimension from the first tentative solution as a sub-problem (step SA4). For example, as shown in FIG. 4, the signs of the first tentative solution and the partial derivative value are compared for each dimension. Assuming that the signs of the 1st, 4th, and 5th dimensions of the first tentative solution are +, and the signs of the 2nd and 3rd dimensions are -, and the signs of the 1st and 5th dimensions of the partial derivative are +, and the signs of the 2nd to 4th dimensions are -, among the dimensions constituting the first tentative solution, the 1st, 2nd, 3rd, and 5th dimensions are extracted as matching dimensions (incompletely optimized dimensions). The remaining 4th dimension is fixed as a non-matching dimension (fully optimized dimension).
[0082] When step SA4 is performed, the second search unit 13 uses the second searcher 16 to calculate a combinatorial optimization problem and search for a second tentative solution regarding the sub-problem (step SA5). The second searcher 16 is, like the first searcher 15, an annealing machine, a simulated annealing machine, or a simulated branch machine.
[0083] Specifically, the second search unit 13 calculates an objective function value based on the target data, explanatory variables, and objective variables regarding the matching dimension, and updates the explanatory variables so that the objective function value is minimized. The second search unit 13 repeats the calculation of the objective function value and the update of the explanatory variables until a predetermined stop condition is satisfied. The predetermined stop condition can be set to the same content as the stop condition in the overall problem. When the predetermined stop condition is satisfied, the second search unit 13 outputs, as the first second tentative solution, the solution vector corresponding to the objective function value at the calculation step when the condition is satisfied. The second tentative solution is a solution vector having the number of dimensions of the matching dimension extracted in step SA4.
[0084] When using a simulated bifurcation machine as the second searcher 16, the second searcher 16 performs the sum-of-products operation of the coupling coefficient matrix and the vector of the Ising problem and the update of the vector based on the result of the sum-of-products operation until the number of predetermined calculation steps is exceeded. When the number of executions of the sum-of-products operation and the vector update exceeds the number of predetermined calculation steps, the second searcher 16 outputs, as the first second tentative solution, the vector at the calculation step when the number of predetermined calculation steps is exceeded. In this case, the number of calculation steps for solving the combinatorial optimization problem may be set to a value proportional to the number of dimensions of the Ising problem, or may be set to the number of steps determined by the user or the like.
[0085] When step SA5 is performed, the second search unit 13 integrates the second tentative solution obtained in step SA5 and the non-matching dimensions (fully optimized dimensions) fixed in step SA4 (step SA6). For example, in the example of FIG. 4, when the first second tentative solution is composed of the 1st to 3rd and 5th dimensions and the fully optimized dimension is composed of the 4th dimension, by integrating the second tentative solution and the non-matching dimensions, a first integrated solution of 5 dimensions having the same number of dimensions as the first tentative solution is generated. Since the non-matching dimensions have not been subjected to the sub-problem, there is no sign change from the first tentative solution.
[0086] When step SA6 is performed, the second search unit 13 determines whether there is a dimension in which the sign of the integrated solution obtained in step SA6 matches the sign of the partial derivative value (step SA7). Specifically, the second search unit 13 calculates the partial derivative value by substituting the integrated solution based on the second tentative solution and the inconsistent dimension into the partial derivative of the objective function with respect to the spin. Since the partial derivative value obtained in step SA7 is recalculated based on the integrated solution, the signs of both the consistent dimension and the inconsistent dimension may be different from the partial derivative value obtained in step SA3. The second search unit 13 compares the sign of the calculated partial derivative value with the sign of the integrated solution obtained in step SA6 for each dimension.
[0087] If it is determined in step SA7 that there is a dimension in which the sign of the integrated solution matches the sign of the partial derivative value (step SA7: YES), the second search unit 13 extracts the matching dimension as a sub-problem (step SA8). For example, as shown in FIG. 4, the signs of the first integrated solution and the partial derivative value are compared for each dimension. Assuming that the signs of the first, second, and fourth dimensions of the first integrated solution are +, and the signs of the third and fifth dimensions are -, and the sign of the first dimension of the partial derivative is +, and the signs of the second to fifth dimensions are -, then the second and fourth dimensions among the dimensions constituting the integrated solution are extracted as the matching dimensions (incompletely optimized dimensions). The remaining first, third, and fifth dimensions are fixed as inconsistent dimensions (completely optimized dimensions).
[0088] Thereafter, the second search unit 13 executes steps SA5 to SA7 with the inconsistent dimension extracted in step SA8 as a sub-problem, and repeats steps SA5 to SA8 until it is determined in step SA7 that there is no dimension in which the sign of the integrated solution matches the sign of the partial derivative value. In the example of FIG. 4, a sub-problem is calculated for the consistent dimension extracted from the first integrated solution, and a second provisional solution is obtained. A second integrated solution is generated based on this second provisional solution and the inconsistent dimension extracted from the first integrated solution. The second integrated solution and the partial derivative value based on the second integrated solution are compared for each dimension. In the case of the example of FIG. 4, the signs of the second integrated solution and the partial derivative value do not match in all dimensions. If there is a dimension that matches between the second integrated solution and the partial derivative value, steps SA4 to SA7 are executed with the inconsistent dimension as a sub-problem, and steps SA5 to SA8 are repeated until it is determined in step SA7 that there is no dimension in which the sign of the integrated solution matches the sign of the partial derivative value.
[0089] When it is determined in step SA7 that there is no dimension in which the sign of the integrated solution matches the sign of the partial derivative (step SA7: NO), the second search unit 13 outputs the integrated solution obtained in step SA6 as the final solution (step SA9). The final solution can be displayed in an arbitrary layout by the display control unit 14. The display target is not limited to only the final solution, and the first provisional solution or the second provisional solution may be displayed in an arbitrary layout such as a graph.
[0090] Thus, the optimization process according to the present embodiment ends.
[0091] FIG. 5 is a diagram comparing the prediction accuracy and the number of consistent dimensions (incomplete optimization dimensions) when solving a binary classification problem using a simulated branch algorithm. The left vertical axis in FIG. 5 is defined as AUC (Area Under Curve), which is an example of prediction accuracy, the right vertical axis is defined as the number of incomplete optimization dimensions, and the horizontal axis is defined as the number of calculation steps. The upper limit value of the number of calculation steps was set to 20,000 times. Note that the number of calculation steps of the sub-problem means the number of times the determination process in step SA7 is performed, in other words, the number of times the iterative process from steps SA4 to SA8 is performed. As shown in FIG. 5, as the number of calculation steps of the sub-problem increases, the prediction accuracy improves and the number of incomplete optimization dimensions tends to decrease. Therefore, by setting the overall problem to a small number of calculation steps and re-optimizing the incomplete optimization dimensions as sub-problems, improvement of accuracy deficiency and improvement of calculation efficiency can be expected.
[0092] FIG. 6 is a diagram comparing the prediction accuracies when solving a binary classification problem using a simulated branch machine (SB) and a simulated annealing machine (SA) between this example and a comparative example. The vertical axis in FIG. 6 is defined as AUC, which is an example of prediction accuracy, and the horizontal axis is defined as the calculation time. The comparative example is an example in which only the overall problem is calculated. In the case of the comparative example, for each of SB and SA, the prediction accuracy was calculated for each calculation step of the overall problem. In the case of this example, for each of SB and SA, after calculating the overall problem in the same manner as the comparative example, the sub-problem was additionally calculated, and the prediction accuracy was calculated for each calculation step of the sub-problem. In the case of this example, the total time of the calculation time of the overall problem and the calculation time of the sub-problem was defined as the calculation time. In the case of SB in this example, after solving the overall problem with SB, the sub-problem was also recalculated with SB. The number of calculation steps of the sub-problem was fixed at 50 times. In the case of SA in this example, after solving the overall problem with SA, the sub-problem was also recalculated with SA. The number of calculation steps of the sub-problem was made proportional to the dimension of the sub-problem. As shown in FIG. 6, together with SB and SA, this example was able to obtain better prediction accuracy and calculation efficiency than the comparative example. It can also be seen that this tendency is particularly remarkable in the case of SB.
[0093] FIG. 7 is a diagram comparing the prediction accuracy when solving a linear regression problem using a simulated branching machine (SB) and a simulated annealing machine (SA) between this example and a comparative example. The vertical axis in FIG. 7 is defined as Cor (correlation coefficient), which is an example of prediction accuracy, and the horizontal axis is defined as the calculation time. The processing procedures of this example and the comparative example are the same as those in FIG. 6. As shown in FIG. 7, also for the linear regression problem, together with SB and SA, this example was able to obtain better prediction accuracy and calculation efficiency than the comparative example.
[0094] FIGS. 8 and 9 are diagrams comparing the prediction accuracy when solving a traveling salesman problem using a simulated branching machine (SB) and a simulated annealing machine (SA) between this example and a comparative example. FIG. 8 is an example (gr96) in which 96 cities in Africa Groetschel are applied to the traveling salesman problem, and FIG. 9 is an example (ch130) in which 130 cities in Churritz are applied to the traveling salesman problem. In gr96, a binary variable of 96×96 dimensions and in ch130, a binary variable of 130×130 dimensions are set as the Ising problem to be solved. The coefficient A for controlling the strength of the constraint term when converting the traveling salesman problem into an Ising problem, shown in Equation (20), was set to A = 0.5. For both SB and SA, the number of calculation steps for the entire problem was set to a fixed value. The number of executions was set to 100 times, and the success rate of the solution (probability that the constraint term is satisfied), time (average calculation time), and success distance (average moving distance in the case of success) were calculated. In the case of SB in this example, after solving the entire problem with SB, the partial problem was calculated with SA. The number of calculation steps for the partial problem was made proportional to the dimension of the partial problem. In the case of SA in this example, after approximately solving the entire problem with SA in fewer calculation steps than usual, the partial problem was also calculated with SA. The number of calculation steps for the partial problem was made proportional to the dimension of the partial problem.
[0095] As shown in FIGS. 8 and 9, in the case of SA in this example, compared with SA in the comparative example, the success rate was improved and the calculation time was reduced, but the moving distance was substantially the same. In the case of SB in this example, compared with SB in the comparative example, the success rate was significantly improved. Also, a solution with a shorter moving distance than SA was obtained.
[0096] The processing procedure of the optimization process shown in FIG. 3 is an example, and the present embodiment is not limited thereto, and deletion, addition, and / or change of various elements are possible.
[0097] As an example, in the process shown in FIG. 3, it is assumed that the end condition of the iteration from step SA4 to SA8 is that the signs of the integrated solution and the partial derivative value match in all dimensions. However, the present embodiment is not limited thereto. The end condition of the iteration from step SA4 to SA8 may be set when the scale of the subproblem to be extracted does not change, when the number of calculation steps reaches a threshold value, when the number of matching dimensions is equal to or less than the threshold value, when the prediction accuracy reaches the threshold value, and the like.
[0098] In the above embodiment, the second search unit 13 extracts the incomplete optimization dimensions from the first tentative solution, and calculates the combinatorial optimization problem with the incomplete optimization dimensions as subproblems to search for the final solution. However, the present embodiment is not limited thereto. As another example, the second search unit 13 extracts the first dimension (incomplete optimization dimension), which is the dimension whose sign matches the sign of the partial derivative value corresponding to the first tentative solution, from the first tentative solution, sets the sign of the first dimension to the opposite sign to generate the second dimension, generates the integrated solution of the second dimension and the remaining dimensions of the first tentative solution, and outputs the integrated solution as the final solution. In this case, the second searcher 16 is unnecessary. Hereinafter, the processing procedure of this modification example will be described. In the following description, components having substantially the same functions as those in the above embodiment are denoted by the same reference numerals, and redundant description will be given only when necessary.
[0099] FIG. 10 is a diagram illustrating the processing procedure of the optimization process according to the modification example. FIG. 11 is a diagram schematically showing the transition of the signs of the solution and the partial derivative value in the optimization process shown in FIG. 10. Similar to FIG. 4, for convenience of explanation, FIG. 11 assumes that the solution vector and the partial derivative value are 5-dimensional vectors. The hatched circles in each vector represent the dimensions with the sign +, and the non-hatched circles represent the dimensions with the sign -.
[0100] As shown in FIG. 10, first, the acquisition unit 11 acquires target data and the like (step SB1), the first search unit 12 obtains a first provisional solution regarding the overall problem (step SB2), and a determination is made as to whether the sign of the first provisional solution matches or does not match the sign of the partial differential value (step SB3). Steps SB1, SB2, and SB3 are the same as steps SA1, SA2, and SA3 in FIG. 3, respectively, and the description thereof is omitted.
[0101] If it is determined in step SB3 that there is no dimension in which the sign of the first provisional solution matches the sign of the partial differential value (step SB3: NO), the second search unit 13 outputs the first provisional solution as the final solution because all dimensions of the first provisional solution obtained in step SB2 are non-matching dimensions and are completely optimized dimensions (step SB7).
[0102] If it is determined in step SB3 that there is a dimension in which the sign of the first provisional solution matches the sign of the partial differential value (step SB3: YES), the second search unit 13 extracts the matching dimensions from the first provisional solution (step SB4). For example, as shown in FIG. 11, the signs of the first provisional solution and the partial differential value are compared for each dimension. Assuming that the signs of the 2nd, 4th, and 5th dimensions of the first provisional solution are +, the signs of the 1st and 3rd dimensions are -, the signs of the 2nd and 5th dimensions of the partial differential value are +, and the signs of the 1st, 3rd, and 4th dimensions are -, then among the dimensions constituting the first provisional solution, the 1st, 2nd, 3rd, and 5th dimensions are extracted as matching dimensions (incompletely optimized dimensions). The remaining 4th dimension is fixed as a non-matching dimension (completely optimized dimension).
[0103] When step SB4 is executed, the second search section 13 inverts the codes of the matching dimensions extracted in step SB4 (step SB5), and integrates the matching dimensions after the code inversion obtained in step SB5 and the non-matching dimensions fixed in step SB4 (step SB6). For example, as shown in FIG. 11, in step SB5, the codes of the 1st and 3rd dimensions of the matching dimensions are inverted from "-" to "+", and the codes of the 2nd and 5th dimensions are inverted from "+" to "-", thereby obtaining the matching dimensions after the code inversion. Then, by integrating the non-matching dimension, which is the 4th dimension, into the non-matching dimensions, a 5-dimensional integrated solution is generated. From the above code operations, the signs of the integrated solution and the partial differential values do not match for all dimensions.
[0104] When step SB6 is executed, the second search section 13 outputs the integrated solution generated in step SB6 as the final solution (step SB7). The final solution can be displayed in an arbitrary layout by the display control section 14. The display target is not limited to only the final solution, and the first tentative solution may be displayed in an arbitrary layout.
[0105] Thus, the optimization process according to the modification example ends.
[0106] According to the modification example, compared with the optimization process shown in FIG. 3, since the optimization calculation of the sub-problems is not performed, the calculation time required to obtain the final solution is shortened. Further, since the codes of the matching dimensions (incompletely optimized dimensions) are inverted, an improvement in the prediction accuracy of the final solution is expected compared to the case where only the overall problem is executed.
[0107] Note that the optimization process shown in FIG. 3 and the optimization process shown in FIG. 10 can be combined. For example, in the optimization process shown in FIG. 3, in step SA7, when it is determined that the number of calculation steps exceeds the threshold value and there is a dimension in which the sign of the integrated solution and the sign of the partial derivative value match (step SA7: YES), the process after step SB4 shown in FIG. 10 may be performed. The threshold value can be set to any value of 1 or more. Alternatively, in the optimization process shown in FIG. 3, in step SA7, when it is determined that the number of calculation steps exceeds the threshold value and there is a dimension in which the sign of the integrated solution and the sign of the partial derivative value match (step SA7: YES), whether to continue performing steps SA5 to SA8 shown in FIG. 3 or to perform the process after step SB4 shown in FIG. 10 may be selected by the user.
[0108] As described above, the information processing apparatus 100 according to the present embodiment has the processor 1. The processor 1 searches for a first tentative solution of a combinatorial optimization problem in which the objective function is partially differentiable using a first searcher 15 that solves the entire problem of the combinatorial optimization problem. The processor 1 outputs the final solution of the combinatorial optimization problem based on the comparison of the signs of the first tentative solution and the partial derivative value of the objective function.
[0109] According to the above configuration, even when a good solution cannot be obtained for the entire problem, it is possible to improve the accuracy of the final solution. Also, even if the number of calculations for the entire problem is reduced, it is possible to maintain the accuracy of the final solution at the same level as before.
[0110] Thus, it becomes possible to improve the calculation accuracy and / or reduce the processing time of the combinatorial optimization problem.
[0111] Although some embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are included in the invention described in the claims and its equivalent scope.
Description of Symbols
[0112] 1 Processor 2 Memory Device 3 Input Device 4 Display Device 5 Communication Device 11 Acquisition Unit 12 First Search Unit 13 Second Search Unit 14 Display Control Unit 15 First Explorer 16 Second Explorer 100 Information Processing Device
Claims
1. A first search unit that searches for a first tentative solution of a combinatorial optimization problem for which an objective function is differentiable using a first searcher that solves the overall problem of the combinatorial optimization problem; A second search unit that outputs a final solution of the combinatorial optimization problem based on a comparison of the signs of the first tentative solution and the partial derivative value of the objective function; An information processing apparatus comprising the above.
2. The second search unit: Extracts a first dimension that is a dimension in which the sign matches the partial derivative value corresponding to the first tentative solution from the first tentative solution; Searches for the final solution by calculating the combinatorial optimization problem with the first dimension as a subproblem using a second searcher that solves the subproblem of the combinatorial optimization problem. The information processing apparatus according to Claim 1.
3. The second search unit: When the signs of the first tentative solution and the partial derivative value do not match in all dimensions, outputs the first tentative solution as the final solution; When the signs of the first tentative solution and the partial derivative value do not all mismatch in all dimensions, extracts the first dimension from the first tentative solution. The information processing apparatus according to Claim 2.
4. The second search unit: Outputs a second tentative solution by calculating the combinatorial optimization problem with the first dimension as a subproblem using the second searcher; Extracts a second dimension that is a dimension in which the sign matches the partial derivative value corresponding to the second tentative solution from the second tentative solution; Searches for the final solution by calculating the combinatorial optimization problem with the second dimension as a subproblem using the second searcher. The information processing apparatus according to Claim 3.
5. The second search unit: Generates an integrated solution of the remaining dimensions between the second tentative solution and the first tentative solution; When the signs of the integrated solution and the partial derivative value corresponding to the second tentative solution do not match in all dimensions, outputs the integrated solution as the final solution; When the signs of the integrated solution and the partial derivative value corresponding to the second tentative solution do not all mismatch in all dimensions, extracts the second dimension from the second tentative solution. The information processing apparatus according to Claim 4.
6. The combinatorial optimization problem is an Ising problem in which the objective function is set to the Ising energy of a plurality of lattice points, for outputting vectors representing combinations of quantum states of a plurality of lattice points as the first tentative solution and the final solution, The quantum state is a binary state. The partial derivative of the objective function is the derivative of the objective function with respect to the quantum state. The information processing apparatus according to claim 2.
7. The information processing apparatus according to claim 6, wherein the Ising problem is any one of a linear regression problem, a support vector machine classification problem, and a traveling salesman problem.
8. The information processing apparatus according to claim 6, wherein the first searcher and the second searcher are Ising machines that output the vector that minimizes the Ising energy.
9. The information processing apparatus according to claim 8, wherein each of the first searcher and the second searcher is a simulated annealing machine or a simulated bifurcation machine.
10. When each of the first searcher and the second searcher is a simulated bifurcation machine, the number of calculation steps for solving the combinatorial optimization problem by the first searcher and / or the second searcher is set to a value proportional to the number of dimensions of the Ising problem. The information processing apparatus according to claim 9.
11. The first searcher and / or the second searcher execute a sum-of-products operation of the coupling coefficient matrix of the Ising problem and the vector, and an update of the vector based on the result of the sum-of-products operation until a predetermined number of calculation steps is exceeded. When the number of executions of the sum-of-products operation and the update exceeds the predetermined number of calculation steps, the vector at the calculation step when the predetermined number of calculation steps is exceeded is output as the first tentative solution and / or the final solution. The information processing apparatus according to claim 6.
12. The second search unit extracts a first dimension that is a dimension in which the sign coincides with the partial derivative value corresponding to the first tentative solution from the first tentative solution, generates a second dimension by setting the sign of the first dimension to the opposite sign, generates an integrated solution of the second dimension and the remaining dimensions of the first tentative solution, outputs the integrated solution as the final solution. The information processing apparatus according to claim 1.
13. Search for a first tentative solution of a combinatorial optimization problem for which the objective function is differentiable using a first searcher that solves the entire problem of the combinatorial optimization problem, output the final solution of the combinatorial optimization problem based on a comparison of the sign between the first tentative solution and the partial derivative value of the objective function. An information processing method comprising the steps of:
14. On a computer, A function that causes a first tentative solution of a combinatorial optimization problem with a differentiable objective function to be searched for using a first searcher that solves the overall problem of the combinatorial optimization problem, A function that outputs a final solution of the combinatorial optimization problem based on a comparison of the signs of the first tentative solution and the partial derivative value of the objective function, An information processing program that realizes the above.