Machine learning optimization method and system based on QUBO model and quantum annealing
By discrete the continuous parameters of the machine learning model into binary variables and using the QUBO model and quantum annealing algorithm to optimize, the autoregressive model has limited nonlinear relationship processing capabilities, high support vector computer computing complexity, and large consumption of convolutional neural network training resources is solved, and more efficient calculations and better model performance are achieved.
Patent Information
- Application Number
- CN202510354916.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing autoregressive models have limited processing capabilities for nonlinear relationships, the computational complexity of the support vector machine is high and difficult to scale to large-scale data sets, and the computational resources are consumed during the training process of convolutional neural networks and are easy to overfit.
The continuous parameters of the machine learning model are discretized into binary variables, and the QUBO model and quantum annealing algorithm are optimized to solve the global optimal solution by adjusting the annealing parameters.
It improves the processing ability of nonlinear relationships, reduces the computational complexity, reduces the computing resource consumption in the training process, and avoids overfitting problems.
Smart Images

Figure CN120297437A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning optimization. More specifically, the present invention relates to a machine learning optimization method and system based on the QUBO model and quantum annealing. Background Art
[0002] An autoregressive model (AR) is a statistical model used for time series prediction. It assumes that the value at the current moment is a linear combination of the values at several past moments, plus a random error term. The AR model is usually used to predict future time series data, such as stock prices, weather changes, etc.
[0003] A support vector machine (SVM) is a supervised learning model used for classification and regression. The SVM separates data points of different classes by finding an optimal hyperplane and maximizes the classification margin. The SVM is widely used in fields such as text classification and image recognition.
[0004] A convolutional neural network (CNN) is a deep learning model widely used in fields such as image classification and object detection. The CNN automatically extracts features in an image through structures such as convolutional layers, pooling layers, and fully connected layers, and performs classification or regression.
[0005] The autoregressive model has limited ability to handle non-linear relationships. Since the autoregressive model assumes that the time series is linear, it cannot capture non-linear relationships well;
[0006] The training process of the support vector machine involves solving a quadratic programming problem, with a relatively high computational complexity. At the same time, the support vector machine has a large computational and storage overhead when dealing with large-scale data, and it is difficult to scale to ultra-large-scale data sets.
[0007] For a convolutional neural network, its training process requires a large amount of computational resources and time. Especially in deep networks, when the training data is scarce, the convolutional neural network model is prone to overfitting.
[0008] Therefore, we propose a machine learning optimization method and system based on the QUBO model and quantum annealing to solve the above problems. Summary of the Invention
[0009] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a machine learning optimization method and system based on the QUBO model and quantum annealing to solve the problems raised in the above background art.
[0010] To achieve the above object, the present invention provides the following technical solution: A machine learning optimization method based on the QUBO model and quantum annealing, including the following content:
[0011] First, discretize the continuous parameters of the machine learning model into binary variables to construct a QUBO objective function;
[0012] Then, use the quantum annealing algorithm to optimize the QUBO model;
[0013] Finally, solve the global optimal solution by adjusting the annealing parameters.
[0014] In a preferred embodiment, the annealing parameters include the initial temperature, the cooling coefficient, and the number of iterations.
[0015] In a preferred embodiment, the machine learning model includes at least one of an autoregressive model, a support vector machine, and a convolutional neural network.
[0016] In a preferred embodiment,
[0017] The discretization process adopts a binary coding or a multi - ary coding method. The binary coding is achieved through the following formula:
[0018] The discretization formula for the continuous parameter c is as follows:
[0019]
[0020] where x j is a binary variable;
[0021] The autoregressive coefficient The discretization formula is as follows:
[0022]
[0023] In a preferred embodiment, the parameters of the quantum annealing algorithm are set as:
[0024] The initial temperature T0 ranges from 100 to 3000;
[0025] The cooling coefficient α ranges from 0.90 to 0.99;
[0026] The number of iterations at each temperature is 10 to 100 times.
[0027] In a preferred embodiment,
[0028] The specific construction of the QUBO objective function includes:
[0029] For the autoregressive model, minimize the sum of squared prediction errors:
[0030]
[0031] In a preferred embodiment,
[0032] For the support vector machine (SVM), a penalty term is introduced:
[0033]
[0034] In a preferred embodiment, for the convolutional neural network (CNN), the objective function is constructed by combining the cross-entropy loss and the regularization term.
[0035] In a preferred embodiment, when the method is applied to time series prediction, it specifically includes:
[0036] Use the data of the first 1 - 9 months as the training set and the data of the 10th month as the test set;
[0037] Solve through the QUBO model to obtain the optimal parameters c and The mean squared error of the predicted value drops to 30723.78, and the goodness of fit R 2 Reaches 0.933.
[0038] In a preferred embodiment, a system of a machine learning optimization method based on the QUBO model and quantum annealing includes the following:
[0039] Parameter discretization module: used to convert continuous parameters into binary variables;
[0040] QUBO model construction module: used to define the objective function and constraint conditions;
[0041] Quantum annealing solution module: built-in simulated annealing solver, supporting parameter customization and parallel computing;
[0042] Result output module: generates the optimal parameter table and performance evaluation indicators.
[0043] Technical effects and advantages of the present invention:
[0044] The QUBO model can discretize the continuous parameters in the AR model into binary variables, thus better handling non-linear relationships;
[0045] The QUBO model can discretize the continuous variables in the SVM into binary variables, thus reducing the computational complexity. The QUBO model can be solved in parallel through quantum computing or simulated annealing algorithms, significantly improving the computational efficiency under large-scale data;
[0046] The QUBO model can discretize the weights and biases in the CNN into binary variables, thus reducing the computational complexity during the training process. Through the QUBO model, the regularization term can be better controlled to avoid overfitting problems. Description of the Drawings
[0047] Figure 1It is the flowchart for solving the QUBO model in the present invention;
[0048] Figure 2 It is the flowchart for solving the QUBO problem by simulated annealing in the present invention;
[0049] Figure 3 It is the classification hyperplane in the case of linear separability in the present invention;
[0050] Figure 4 It is the diagram of the convolutional neural network in the present invention;
[0051] Figure 5 It is the image diagram of the first 16 digits in the MNIST dataset in the present invention. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] A machine learning optimization method based on the QUBO model and quantum annealing includes the following contents:
[0054] First, discretize the continuous parameters of the machine learning model into binary variables and construct a QUBO objective function;
[0055] Then use the quantum annealing algorithm to optimize the QUBO model;
[0056] Finally, solve the global optimal solution by adjusting the annealing parameters.
[0057] The annealing parameters include the initial temperature, the cooling coefficient, and the number of iterations.
[0058] The machine learning model includes at least one of an autoregressive model, a support vector machine, and a convolutional neural network.
[0059] The discretization process adopts a binary coding or a multi - ary coding method. The binary coding is achieved through the following formula:
[0060] The discretization formula for the continuous parameter c is as follows:
[0061]
[0062] where x j is a binary variable;
[0063] The discretization formula for the autoregressive coefficient φ i is as follows:
[0064]
[0065] The parameter settings of the quantum annealing algorithm are as follows:
[0066] The range of the initial temperature T0 is from 100 to 3000;
[0067] The range of the cooling coefficient α is from 0.90 to 0.99;
[0068] The number of iterations at each temperature is from 10 to 100 times.
[0069] The specific construction of the QUBO objective function includes:
[0070] For the autoregressive model, minimize the sum of squared prediction errors:
[0071]
[0072] For the support vector machine (SVM), introduce a penalty term:
[0073]
[0074] For the convolutional neural network (CNN), construct the objective function by combining the cross-entropy loss and the regularization term.
[0075] When this method is applied to time series prediction, it specifically includes:
[0076] Use the data of the first 1 - 9 months as the training set and the data of the 10th month as the test set;
[0077] Solve to obtain the optimal parameters c and φ through the QUBO model i , and the mean squared error of the predicted value drops to 30723.78, and the goodness of fit R 2 reaches 0.933.
[0078] A system of a machine learning optimization method based on the QUBO model and quantum annealing includes the following:
[0079] Parameter discretization module: used to convert continuous parameters into binary variables;
[0080] QUBO model construction module: used to define the objective function and constraint conditions;
[0081] Quantum annealing solution module: built-in simulated annealing solver, supporting parameter customization and parallel computing;
[0082] Result output module: generate the optimal parameter table and performance evaluation indicators.
[0083] The autoregressive model (AR) is a model for studying time series, which is characterized by wide applicability and high accuracy. Since the AR model assumes that the time series is linear and cannot capture non-linear relationships well, there are limitations when using the autoregressive model for prediction. The QUBO model can discretize the continuous parameters in the AR model into binary variables, thus better handling non-linear relationships. The specific application here is to more accurately predict the value in October based on the actual values from January to September.
[0084] The AR model mainly makes predictions based on the linear combination of past observed values and current interference values. Generally, the p-order autoregressive process AR(p) is as follows:
[0085]
[0086] Among them, represents the predicted value of the t-th month, Y t-i represents the true value of the (t - 1)-th month; c is a constant, representing the average level of the time series; φ i refers to the autoregressive model parameter, representing the influence of the i-th lag term on the current value; ε t refers to the random noise with a mean of zero; p is the order of the AR model, indicating how many lag terms are used to predict the current value, which is determined by AIC and p = 3.
[0087] On this basis, we transform the above time series prediction problem into a quadratic unconstrained binary optimization (QUBO) model. Its minimum energy state is called the ground state. The QUBO model consists of multiple 0-1 binary variables, and a solution is given after combination. In particular, the combination of binary variables that gives the minimum energy state is called the ground state solution. The QUBO model uses the binary variable x i to represent, and this variable takes 0 or 1 (x i = 0 or x i = 1)
[0088]
[0089] Among them, X = {x1, x2...} is a set of binary variables; a i is the coefficient representing the binary variable x i ; b ij is the interaction coefficient representing the connection weight between the binary variables x i and x j .
[0090] The solution of the QUBO model is a kind of combination of binary variables. The ground state solution of the QUBO model is the combination of binary variables that minimizes f(X). The f(X) value given for a combination of binary variables is called the target value.
[0091] Denote Q as composed of b ijThe composed matrix transforms the objective function into a quadratic form according to the QUBO Hamiltonian it generates. Then, formula (2) can be simplified as follows:
[0092] Q(x) = x T Qx + a T x#(3)
[0093] where x is a binary decision variable, x T represents the transpose of x, a is the coefficient vector of the linear term, Q is the coefficient matrix of the quadratic term of the quadratic unconstrained binary optimization problem, and it is a symmetric matrix. Otherwise, the form of the coefficient matrix Q can be changed through the following formula:
[0094]
[0095] Take 70% of the data as the training set and 30% of the data as the test set. By determining the autoregressive function when the sum of the squared errors of the predicted values is minimized, predict the required values. Obtain the predicted values for October. In this problem, our goal is to minimize the error of the predicted values while satisfying the constraints that vary within a reasonable range. Therefore, based on the above analysis, construct the objective function D for minimizing the predicted values:
[0096]
[0097] Combined with formula (3), the objective function can be expressed as follows:
[0098]
[0099] Based on this, we define a (1 + p)·n - dimensional vector x to discretize the continuous values c and φ i into binary decision variables, where x 1:n is the binary representation of c, x n+1:(p+1)n is the binary representation of φ i , arranged in the order of i, that is:
[0100] x = [c1, c2, … c n , φ 11 , φ 12 , … φ 1n , φ 21 , … φ pn
[0101] The specific mathematical formula for binaryization is as follows:
[0102]
[0103] Substitute formula (7) into formula (6) to simplify and expand the objective function:
[0104]
[0105] Moreover, since the constant term is independent of the values of the decision variables when the function D reaches its minimum value, for the sake of simplicity in calculation, we directly ignore the constant term. Formula (9) can be simplified to the following mathematical expression:
[0106]
[0107] In addition, we impose constraints on the decision variables through the following formula:
[0108]
[0109] Among them, before using the Kaiwu SDK, the QUBO matrix needs to be converted into Ising variables. The Ising Model is a type of stochastic process model that describes the phase transition of substances. Abstracted into a mathematical form as follows:
[0110]
[0111] Among them, σ is the spin variable to be solved, taking values in {-1, 1}, H is the Hamiltonian, J is the quadratic term coefficient, and μ and h are the linear term coefficients, which are known quantities [4].
[0112] We adopt the following formula as the conversion formula between binary variables and Ising variables:
[0113]
[0114] Among them, σ i =-1 corresponds to x i =0, σ i =+1 corresponds to x i =1.
[0115] Here, we summarize the solution idea of this method: Refer to the attached drawings of the specification Figure 1
[0116] 2.1.2 Solving the system of the above method based on the quantum annealing algorithm
[0117] In order to solve for the value corresponding to October when the prediction value error is minimized, we use the simulated annealing solver built into the Kaiwu SDK for the solution. Refer to the attached drawings of the specification Figure 2 .
[0118] For the parameters in the simulated annealing solver in the Kaiwu SDK, the settings are as follows:
[0119] Table 1 Simulated Annealing Parameter Table
[0120]
[0121] Select the initial temperature T0 and the termination temperature T end , and select an appropriate cooling coefficient α during the annealing process so that the iteration depth at each temperature during the annealing process does not exceed 100.
[0122] Through the solution of the simulated annealing solver, the following results are obtained:
[0123] Table 2 Optimal prediction parameter value table
[0124]
[0125] At this time, the mean square error D' reaches the minimum value of 30723.78. Among them, φ i can be regarded as the weight of the first i true values of the predicted value to the predicted value. Then, the phenomenon that φ1 has a large influence on the predicted value conforms to the actual situation, which is reasonable. The prediction result has a high goodness of fit. Based on this, various prediction evaluation indicators are calculated to assist in explaining that the result performance is excellent.
[0126] Table 3 Result evaluation index value statistics
[0127]
[0128] In addition, the change process of the Hamiltonian over time in the simulated annealing algorithm can reflect the energy state of the system. When the system is in the initial state, the energy is relatively high; then, the Hamiltonian drops rapidly, and the system quickly finds a state with a lower energy during the optimization process; there is a small increase in the Hamiltonian when it is close to 0.03 seconds, which is because the algorithm is trying to jump out of the local optimal solution; after the drop, the Hamiltonian tends to be stable, indicating that the algorithm has converged and the system has reached the global optimal state. The solution time of the QUBO model is about 0.07s.
[0129] 2.2.1 Application method of the QUBO model-based optimized support vector machine algorithm (SVM) in three-class classification problems
[0130] The support vector machine algorithm (SVM) is developed from the optimal classification surface in the linearly separable case. The basic idea can be illustrated by Figure 3 the two-dimensional case. As shown in the following figure, solid dots and hollow dots represent two types of samples respectively. If the two are linearly separable, the result of machine learning is a hyperplane (also called a discriminant function), and this hyperplane divides the training samples into positive and negative classes. Refer to the attached drawings of the specification Figure 3 .
[0131] Obviously, there are infinitely many such hyperplanes. According to the requirements of the structural risk minimization principle (SRM), the result of machine learning should be the optimal hyperplane, which not only correctly separates the two types of training samples, but also maximizes the classification margin.
[0132] Since SVM is a binary classification model, in a three-class classification problem, we first split the three-class classification problem into 3 different binary classification problems. By doing so, a three-class classification problem is transformed into three independent binary classification problems.
[0133] Suppose an m-dimensional hyperplane is described by the following equation:
[0134] w·x + b = 0, w ∈ R m , b ∈ R #(14)
[0135] Then, the optimal hyperplane with the largest classification margin can be obtained by finding the minimum value of , and the constraint conditions at this time are:
[0136] y i (w·x i + b) - 1 ≥ 0, i = 1, …, n #(15)
[0137] Among them, y i ∈ {-1, +1}, x i is the sample feature vector. The constraint conditions ensure that the distance from the data point to the hyperplane is at least 1. When the above equation holds with equality, it means that the data point is exactly on the margin boundary, and these data points are called support vectors.
[0138] In the case of linear inseparability, such as in the presence of noisy data, SVM introduces a relaxation term ξ i ≥ 0 in the above equation to achieve a soft margin, that is:
[0139] y i (w T ·x i + b) ≥ 1 - ξ i i = 1, …, n #(16)
[0140] Based on this, the objective function is the following equation:
[0141]
[0142] Among them, w is the normal vector of the hyperplane, b is the bias of the hyperplane, ξ i is the relaxation variable, C is the regularization parameter, which is used to control the trade-off between the classification margin and misclassification, determines their relative importance, and this value is selected by cross-validation, that is, C is adjusted according to the performance on the validation set.
[0143] The training process of SVM involves solving quadratic programming problems, with a relatively high computational complexity. At the same time, SVM has a large computational and storage overhead when dealing with large-scale data, making it difficult to scale to ultra-large-scale datasets. The QUBO model can discretize the continuous variables in SVM into binary variables, thereby reducing the computational complexity. The QUBO model can be solved in parallel through quantum computing or simulated annealing algorithms, significantly improving the computational efficiency under large-scale data. Based on the above analysis, we transform SVM into a QUBO model, which means transforming Equation (17) into Equation (3).
[0144] First, we discretize w, b, and ξ in SVM i by converting them into binary form.
[0145] Assume that w j ∈[w min , w max . It is represented by m binary variables z j,k ∈{0, 1} for w j :
[0146]
[0147] Among them,
[0148]
[0149] Then the mathematical expression for the binary conversion of the w two-norm is as follows:
[0150]
[0151] Similarly, we assume that b ∈ [b min , b max . It is represented by m b binary variables z b,k ∈{0, 1} as follows:
[0152]
[0153] Among them,
[0154]
[0155] And b min , b max represent the minimum and maximum values of b respectively.
[0156] Converting Equation (15) into a QUBO penalty term, we get:
[0157]
[0158] Among them, b is the bias and λ is the penalty term weight.
[0159] Secondly, we substitute the binary expressions of w and b into the above formula respectively, and we get:
[0160]
[0161] Among them, is the constant term, is the part about binary variables. Finally, substitute the above formula into formula (23):
[0162]
[0163] Among them is the constant term, is the linear term about binary variables. Denote the constant term as C i 、L i respectively, then the above formula can be abbreviated as follows:
[0164] 1 - y i (w T x i + b) = C i + L i #(27)
[0165] Then formula (23) can be denoted as:
[0166]
[0167] Therefore, the objective function can be integrated as follows:
[0168]
[0169] 2.2.2 Solving the system of the above method based on quantum annealing algorithm
[0170] First, we perform data preprocessing, shuffle the original ordered data, and use 70% of the data for data training and the remaining for data prediction.
[0171] We still use the simulated annealing solver built in Kaiwu SDK for solving, and the parameters in the solver are set as follows:
[0172] Table 4 Simulated annealing parameter table
[0173]
[0174] Select the initial temperature T0 and the termination temperature T end , and select an appropriate cooling coefficient α during the annealing process so that the iteration depth at each temperature during the annealing process does not exceed 10.
[0175] By solving with a simulated annealing solver, the parameter results of three independent binary classification problems can be obtained. At this time, it can be seen from the comparison of the classification results that the binary classification results of QUBO-SVM are more accurate. Based on this, multiple prediction evaluation metrics are calculated to assist in demonstrating the excellent performance of the results.
[0176] Table 5 Statistical Values of Result Evaluation Metrics (First and Second Planes)
[0177]
[0178] Table 6 Statistical Values of Result Evaluation Metrics (First and Third Planes)
[0179]
[0180] Table 7 Statistical Values of Result Evaluation Metrics (Second and Third Planes)
[0181]
[0182] In addition, the change process of the Hamiltonian with time in the simulated annealing algorithm during the solution of the three binary classification problems can reflect the energy state of the system. It can be seen that the Hamiltonian drops rapidly at the beginning, indicating that the energy of the system decreases rapidly. There are some fluctuations later, which means that during the optimization process, the system explores different states to find a state with lower energy. As time goes by, the dropping speed gradually slows down and the Hamiltonian tends to be stable, which may mean that the system has approached or reached a stable state. And the solution time of the QUBO model is between 0.07 - 0.09 s.
[0183] 2.3.1 Application Method of Optimizing Convolutional Neural Network (CNN) Based on QUBO Model in Image Classification Problem of MNIST Handwritten Digit Dataset
[0184] So far, the pattern recognition system based on convolutional neural network is one of the best-performing systems currently, especially in the field of handwritten character recognition, and has been used as the evaluation standard for the performance of machine recognition systems. Convolutional neural networks reduce the number of trainable parameters in the network by exploring the spatial correlations in the data, thereby improving the efficiency of the backpropagation algorithm for the forward propagation network.
[0185] As shown in the figure below, after the input layer reads in the pictures of unified size, the input image is convolved with three filters and an additive bias. After convolution, three feature maps are generated in the C1 layer. Then, four adjacent pixels in the feature maps are grouped and averaged, and then weighted and biased. Three feature maps of the S2 layer are obtained through the ReLU activation function.
[0186] Among them,
[0187] ReLU(x) = max(0, x) #(30)
[0188] These mapping graphs are then filtered accordingly to obtain the C3 layer. This layer then generates S4 in the same way as S2. Finally, these pixel values are rasterized and connected into a one-dimensional vector, which is input into a traditional neural network and then enters the fully connected layer to obtain the output. Refer to the attached drawings of the specification Figure 4 .
[0189] Among them, the C layer is the convolutional layer, that is, the feature extraction layer. A convolutional layer usually contains multiple feature maps with different weight vectors, so that multiple different features can be obtained at the same position; the S layer is the pooling layer, which reduces the resolution of the feature map by performing local averaging and downsampling operations, and at the same time reduces the sensitivity of the network output to displacement and deformation. The subsequent convolutional layers and pooling layers are alternately connected to form a "double pyramid" structure: the number of feature maps gradually increases, and the resolution of the feature maps gradually decreases.
[0190] The convolutional layer is the core of the CNN. Each convolutional layer uses several convolutional layers to scan the input image. For each convolutional kernel, the convolution operation calculates the weighted sum of the local area between the input image and the convolutional kernel. Assume that the size of the input image is H×W×D (height, width, depth), the convolution operation uses a convolutional kernel of size k×k, the depth is D, and the number of convolutional kernels (output channels) is F. For each local area X in the input image X i and the convolutional kernel W, the convolution operation is as follows:
[0191]
[0192] where W m,n,d represents the (m, n, d) - th element of the convolutional kernel, X i+m-1,j+n-1,d is the (i + m - 1, j + n - 1, d) - th pixel of the input image, and b is the bias term.
[0193] The pooling layer is used to downsample the feature map output by the convolutional layer. The pooling operation performed here is max - pooling, that is, for each k×k region in the feature map, the maximum value in this region is selected as the output.
[0194] Each image in the MNIST dataset corresponds to a digit label from 0 to 9 and is composed of NIST Special Database 1 and NIST Special Database 3. In MNIST, 30,000 examples from SD-3 and SD-1 are selected for the training set, and the 60,000 selected examples are from the handwritten data of approximately 250 different individuals. Similarly, 5,000 examples from SD-3 and SD-1 are selected as the test set. The images in the MNIST dataset are all of size 28×28. Images of some digits in the dataset are as follows:
[0195] In a multi-class classification problem, the objective function is to minimize the difference between the predicted result and the true label. Suppose the result predicted by the model is The actual label is y i , and we use the cross-entropy loss function to measure the error:
[0196]
[0197] where C is the total number of classes. For this dataset, C is the constant 10; is the k-th element of the true label of the sample x i . If the sample belongs to class k, then otherwise it is 0; is the probability that the model predicts the sample x i belongs to class k. The output layer of the CNN uses the softmax activation function to calculate the probability of each class. The softmax activation function is expressed as the following mathematical formula:
[0198]
[0199] where, is the raw score of class k output by the fully connected layer, is the predicted probability of class k.
[0200] For the convolutional layer l and the fully connected layer, their outputs can be expressed as follows:
[0201] Z (l) = ReLU(W (l) *X + b (l) )#(34)
[0202] where, W (l) refers to the weight matrix of the convolutional layer and the fully connected layer, X is the output of the previous layer (or the input image), and b (l) is the bias term. In addition, denoting y i as the target value, the error term of the convolutional layer can be expressed as follows:
[0203]
[0204] However, for CNN, its training process requires a large amount of computing resources and time, especially in deep networks. In particular, when the training data is scarce, the CNN model is prone to overfitting. The QUBO model can discretize the weights and biases in CNN into binary variables, thereby reducing the computational complexity during the training process. Through the QUBO model, the regularization term can be better controlled to avoid overfitting problems.
[0205] To combine the variables involved in the above theoretical analysis with QUBO, we define x i as the binary activation variable of the i-th neuron, indicating whether the neuron is activated, with values of 0 or 1; x k as the binary variable for the predicted class k of the model, indicating whether the sample is classified as class k, with values of 0 or 1; In addition, we also need to replace W ij , b i with binary expressions.
[0206] To prevent overfitting, we add a regularization term, especially the L2 regularization term.
[0207]
[0208] Based on the above analysis, we determine the objective function as:
[0209]
[0210] where λ conv , λ fc , λ crcoss-entrcopy are penalty coefficients that control the importance of the error term.
[0211] On this basis, we determine the constraint conditions. The classification result needs to satisfy "only one class is predicted", the activation value of the neuron can only be 0 or 1, the convolution operation is discretized into binary decision variables, the neuron satisfies the calculation rule of the weighted sum, the cross-entropy loss function is minimized, and the regularization constraint. The specific mathematical expressions are as follows:
[0212]
[0213] where x k is the predicted binary variable for class k, and w1, w2,..., w n are the discretized weight values.
[0214] 2.3.2 System for Solving the Above Method Based on Quantum Annealing Algorithm
[0215] We still use the simulated annealing solver built into the Kaiwu SDK for the solution. The parameters in the solver are set as follows:
[0216] Table 8 Simulated Annealing Parameter Table
[0217]
[0218]
[0219] Select the initial temperature T0 and the termination temperature T end , and select an appropriate cooling coefficient α during the annealing process so that the iteration depth at each temperature during the annealing process does not exceed 100.
[0220] Through the solution, the images in the MNIST dataset and their recognition results are obtained. In addition, the loss during the model training process decreases as the training cycle increases, and both the training loss and the validation loss are decreasing, which indicates that the model is learning and gradually reducing the prediction error. The validation loss tends to be stable after the initial decrease and is always lower than the training loss, indicating that the model performs well on the training set and there is no obvious overfitting phenomenon. The accuracy rate increases as the training cycle increases, and both the training accuracy rate and the validation accuracy rate are rising, and both are close to 1.0, which indicates that the prediction of the model is becoming more and more accurate. Also, the training accuracy rate and the validation accuracy rate are very close, which further indicates that the model is not overfitting and has good generalization ability on both the training set and the validation set.
[0221] In addition, we calculated a variety of prediction evaluation metrics to assist in explaining that the image recognition classification results are excellent:
[0222] Table 9 Statistical Values of Result Evaluation Metrics (Iris-versicolor and Iris-virginica)
[0223]
[0224] As can be seen from the above table, the precision, recall rate, and F1 score of the model recognition results are all close to 0.99 or 1.00, indicating that the performance of the model in the handwritten digit recognition task is very excellent. Due to the huge amount of data and the individual differences in hardware performance, the solution time of the QUBO model is 211s.
[0225] Alternative solutions for the construction and discretization of the QUBO model:
[0226] In addition to binary discretization, other discretization methods (such as multi-valued discretization) can be used to convert continuous parameters into discrete variables.
[0227] Alternative solutions for the solution of the QUBO model based on the simulated annealing algorithm:
[0228] Other classical optimization algorithms such as genetic algorithms, particle swarm optimization (PSO), ant colony algorithms, etc. can replace the simulated annealing algorithm to solve the QUBO model.
[0229] Alternative solutions for the application of the QUBO model in time series prediction:
[0230] Other time series models: Deep learning models such as long short-term memory networks (LSTM), gated recurrent units (GRU), etc. can replace the AR model for time series prediction.
[0231] Alternative solutions for the application of the QUBO model in classification problems:
[0232] Other classification models such as random forests, XGBoost, neural networks, etc. can replace SVM for classification tasks.
[0233] Alternative solutions for the application of the QUBO model in image classification:
[0234] Other image classification models such as deep learning models like ResNet, EfficientNet, etc. can replace CNN for image classification.
[0235] Alternative solutions for the parallel solution ability of quantum computing:
[0236] Classical parallel computing: Use classical parallel computing hardware such as GPUs or TPUs to accelerate the solution process of the QUBO model.
[0237] Hybrid quantum-classical computing: Combine quantum computing with classical computing, using quantum computing to process key parts and classical computing to process other parts.
[0238] Construction and discretization processing of the QUBO model:
[0239] Discretize the continuous parameters (such as autoregressive coefficients, hyperplane normal vectors, convolutional kernel weights, etc.) in traditional machine learning models (such as AR, SVM, CNN) into binary variables to construct the QUBO model.
[0240] Solving the QUBO model based on the simulated annealing algorithm:
[0241] Use the simulated annealing algorithm to solve the QUBO model. By setting parameters such as the initial temperature, cooling coefficient, and termination temperature, gradually optimize the objective function to find the global optimal solution.
[0242] Application of the QUBO model in time series prediction:
[0243] Convert the autoregressive model (AR) into a QUBO model. By discretizing the autoregressive coefficients and constant terms, construct an objective function, and use the simulated annealing algorithm to solve it to predict future time series data.
[0244] Application of QUBO model in classification problems:
[0245] Convert the support vector machine (SVM) into a QUBO model. By discretizing the normal vector and bias of the hyperplane, construct an objective function, and use the simulated annealing algorithm to solve it to achieve efficient classification.
[0246] Application of QUBO model in image classification:
[0247] Convert the training optimization problem of the convolutional neural network (CNN) into a QUBO model. By discretizing the convolutional kernel weights and biases, construct an objective function, and use the simulated annealing algorithm to solve it to achieve efficient image classification.
[0248] Parallel solution ability of quantum computing:
[0249] Utilize the parallel computing ability of quantum computing to efficiently solve the QUBO model, significantly improving the computational efficiency under large-scale datasets.
[0250] Binary discretization of continuous values:
[0251] In the QUBO model for time series prediction, we define a vector to discretize continuous values into binary decision variables, which is considered complete and helps to establish a better QUBO prediction model.
[0252] QUBO classification model based on SVM:
[0253] Based on the principle of SVM, convert the three-class classification problem of the dataset into three independent two-class classification problems for complete classification. Add a penalty term to this model and obtain the optimal classification result of the model by finding the optimal penalty coefficient.
[0254] Convert the classification of MNIST handwritten digit dataset images by CNN into a QUBO model:
[0255] Utilize the optimization ability of quantum computing to convert the CNN model in the classical image classification scenario of the MNIST handwritten digit dataset into a QUBO model. The prediction accuracy of this model on the test dataset reaches 99%, indicating that the performance of this model in the handwritten digit recognition task is very remarkable.
[0256] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A machine learning optimization method based on the QUBO model and quantum annealing, characterized in that; It includes the following: First, discretize the continuous parameters of the machine learning model into binary variables and construct a QUBO objective function; Then, use the quantum annealing algorithm to optimize the QUBO model; Finally, solve the global optimal solution by adjusting the annealing parameters.
2. The machine learning optimization method based on the QUBO model and quantum annealing according to claim 1, characterized in that: The annealing parameters include the initial temperature, the cooling coefficient, and the number of iterations.
3. The machine learning optimization method based on the QUBO model and quantum annealing according to claim 1, characterized in that: The machine learning model includes at least one of an autoregressive model, a support vector machine, and a convolutional neural network.
4. A machine learning optimization method based on the QUBO model and quantum annealing according to claim 1, characterized in that: The discretization process adopts a binary coding or a multi - ary coding method. The binary coding is implemented by the following formula: The discretization formula for the continuous parameter c is as follows: where x j is a binary variable; Autoregressive coefficient The discretization formula is as follows:
5. A machine learning optimization method based on the QUBO model and quantum annealing according to claim 1, characterized in that: The parameter settings of the quantum annealing algorithm are: The initial temperature T0 ranges from 100 to 3000; The cooling coefficient α ranges from 0.90 to 0.99; The number of iterations at each temperature is 10 to 100 times.
6. The machine learning optimization method based on the QUBO model and quantum annealing according to claim 1, characterized in that: The specific construction of the QUBO objective function includes: For the autoregressive model, minimize the sum of squared prediction errors:
7. A machine learning optimization method based on the QUBO model and quantum annealing according to claim 1, characterized in that: For the support vector machine (SVM), introduce a penalty term:
8. A machine learning optimization method based on the QUBO model and quantum annealing according to claim 1, characterized in that: For the convolutional neural network (CNN), construct the objective function by combining the cross - entropy loss and the regularization term.
9. A machine learning optimization method based on the QUBO model and quantum annealing according to claim 1, characterized in that: When this method is applied to time - series prediction, it specifically includes: Use the data of the first 1 - 9 months as the training set and the data of the 10th month as the test set; The optimal parameters c are obtained by solving the QUBO model and the mean squared error of the predicted value is reduced to 30723.78, and the goodness of fit R 2 reaches 0.
933.
10. A system for a machine learning optimization method based on the QUBO model and quantum annealing according to any one of claims 1-9, characterized in that: It includes the following: Parameter discretization module: used to convert continuous parameters into binary variables; QUBO model construction module: used to define the objective function and constraints; Quantum annealing solution module: built - in simulated annealing solver, supporting parameter customization and parallel computing; Result output module: generate the optimal parameter table and performance evaluation indicators.
Citation Information
Cited By
Fire detection method and system based on binocular vision and ultraviolet
CN120852809A
Medical image classification method and device based on coherent Isin machine, equipment and medium
CN121147591A
Rail transit connection service quality evaluation and optimization method based on intelligent traffic technology
CN121724493A