Interactive mixed quantum-classical machine learning model training method, device and medium

By using the output of classical neural networks as parameters of quantum machine learning, the problems of insufficient gradient calculation accuracy and large measurement overhead in quantum computing devices containing noise-containing medium-scale quantum computing devices are solved, and efficient quantum computing forward computing and gradient updates are achieved, improving pattern recognition capabilities.

CN120124768APending Publication Date: 2025-06-10上海量感智能科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510285148.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the quantum machine learning, in the medium-scale quantum computing devices with noise, there are problems such as insufficient gradient calculation accuracy, large measurement overhead, insufficient noise resistance and insufficient pattern recognition capabilities.

Method used

By taking the output of classical neural networks as parameters of quantum machine learning, efficient quantum computing forward computing is achieved, and the parameters are updated in combination with the gradient backpropagation algorithm of classical neural networks, the problems of insufficient gradient calculation accuracy and measurement overhead of quantum machine learning are overcome.

Benefits of technology

It effectively avoids the problem of large overhead of quantum machine learning training, overcomes the problem of disappearing quantum gradients, makes full use of the functions of quantum and classical neural networks, and improves the time series prediction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124768A_ABST
    Figure CN120124768A_ABST
Patent Text Reader

Abstract

The invention relates to an interactive hybrid quantum-classical machine learning model training method and device and a medium, and the method comprises the steps: loading training data, and initializing a classical neural network and a quantum machine learning model; inputting the classic data into a classic neural network to obtain a classification result as prediction output; inputting the quantum data into a quantum machine learning model, taking the output of the classical neural network as a parameter of quantum machine learning, and obtaining an expected value of a quantum bit under considerable measurement as a prediction output; calculating a weighted joint loss function based on a classic neural network, prediction output of quantum machine learning and a true value of a data set; updating classic neural network parameters based on a weighted joint loss function; when the weighted joint loss function converges or reaches the maximum number of iterations, training is ended. Compared with the prior art, the method has the advantages that the expense of quantum machine learning training is reduced, and the potential gradient disappearance problem of a quantum machine learning algorithm is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of quantum computing and artificial intelligence, and more particularly to an interactive hybrid quantum-classical machine learning model training method, device, and medium. Background Art

[0002] Quantum machine learning utilizes its unique quantum entanglement and superposition properties. Compared with classical machine learning algorithms, it has the potential for exponential acceleration in computing. In the era of big data and large models, it can effectively reduce energy consumption and increase computing power requirements. However, in noisy medium-scale quantum computing devices, due to the limitations of quantum computer hardware, there are many obstacles in the expressiveness, trainability, etc. of quantum machine learning algorithms. To achieve efficient prediction and training, quantum machine learning needs to overcome the interference of noise to accurately estimate observables. However, research has shown that accurately estimating the expectation of observables requires consuming a large amount of quantum resources. Especially in quantum machine learning algorithms, learning from big data and performing parameter iteration require a still extremely large overhead of quantum computing resources. Therefore, it further limits the parameter scale of quantum machine learning. At this time, the acceleration advantage of quantum computing will be reduced, and due to the limitation of the parameter scale, the learning ability and expressiveness of quantum machine learning algorithms will also be severely limited. However, the success of classical machine learning in many fields has brought new opportunities to quantum machine learning.

[0003] CN116432710A discloses a machine learning model construction method, a machine learning framework, and related devices, which are applied to an electronic device including a first machine learning framework and not including a second machine learning framework. The first machine learning framework includes a compatible quantum computing program encapsulation unit. By determining the compatible quantum computing layer interface corresponding to the second machine learning framework, the compatible quantum computing layer interface is used to provide a quantum program created based on a quantum computing programming library included in the second machine learning framework; by using the compatible quantum computing program encapsulation unit to call the compatible quantum computing layer interface to construct a compatible quantum computing layer; constructing a machine learning model including the compatible quantum computing layer, realizing the construction of a machine learning model across quantum machine learning frameworks. However, the purpose of this method is to achieve the compatible combination of classical machine learning models and quantum computing, and solve the compatibility problem, and cannot solve the overhead problem of quantum computing. Summary of the Invention

[0004] The object of the present invention is to overcome the defects existing in the above-mentioned prior art and provide an interactive hybrid quantum-classical machine learning model training method, device and medium, aiming to directly learn classical data through a classical neural network algorithm to obtain an output, and use it as the parameter of quantum machine learning to achieve efficient quantum computing forward calculation, thereby overcoming the problems of insufficient accuracy of quantum machine learning gradient calculation and large measurement overhead, giving full play to the functions of quantum and classical neural networks, and providing a new solution for solving the limitations of large measurement overhead, insufficient anti-noise ability and insufficient pattern recognition ability in quantum machine learning.

[0005] The object of the present invention can be achieved by the following technical solutions:

[0006] According to the first aspect of the present invention, an interactive hybrid quantum-classical machine learning model training method is provided, and the method includes the following steps:

[0007] Data loading and model initialization: Load training data, and initialize a classical neural network and a quantum machine learning model;

[0008] Interactive quantum-classical machine learning calculation: Input the classical data in the training data into the classical neural network, execute the classical calculation process, and obtain the classification result as the prediction output; Input the quantum data in the training data into the quantum machine learning model, and use the output of the classical neural network as the parameter of the quantum machine learning model, and sequentially execute the quantum calculation process and the ordinary quantum measurement process to obtain the expected value of the quantum bit under the observable as the prediction output;

[0009] Joint loss function calculation: Calculate the classical loss based on the prediction output of the classical neural network and the true value of the data set, calculate the quantum loss based on the prediction output of the quantum machine learning model and the true value of the data set, and calculate the weighted joint loss function based on the classical loss and the quantum loss;

[0010] Update of classical neural network parameter gradient: Update the classical neural network parameters by using the gradient backpropagation algorithm based on the weighted joint loss function;

[0011] Iteration termination determination: When the weighted joint loss function converges or reaches the maximum number of iterations, end the training; otherwise, according to the updated classical neural network parameters, return to the interactive quantum-classical machine learning calculation step for the next round of iteration.

[0012] As a preferred technical solution, the data loading and model initialization is specifically:

[0013] Judge the type of training data, and the type of training data includes classical data and quantum data;

[0014] If the training data only includes classical data, it is serialized, converted into sequence data, loaded through a classical neural network, and the qubits of the quantum machine learning model are reset, that is, the input of the quantum machine learning model does not contain any information;

[0015] If the training data only includes quantum data, it is loaded using a quantum machine learning model, and fixed classical data is randomly generated, that is, the input of the classical neural network does not contain any information;

[0016] If the training data includes both classical data and quantum data, the classical data is serialized, converted into sequence data, loaded through a classical neural network, and the quantum data is loaded through a quantum machine learning model.

[0017] As a preferred technical solution, in the interactive quantum-classical machine learning calculation step, a classical neural network is used to calculate the serialized classical data, encode the classical data in the hidden state, obtain the encoded hidden state at each moment, and obtain the output of the classical neural network according to the hidden state at the current moment; wherein, the number of network layers of the classical neural network is the same as the time series length of the serialized classical data.

[0018] As a preferred technical solution, in the interactive quantum-classical machine learning calculation step, the quantum machine learning model is mathematically characterized as a parameterized unitary matrix, the parameters of the quantum machine learning model are set as the output of the classical neural network at each moment, and the maximum number of parameters in each layer of the quantum machine learning model does not exceed the dimension of the encoded hidden state. The quantum calculation process is performed according to the set parameters, and the Pauli Z operator is measured to estimate the observable expectation value as the prediction output. When the quantum machine learning model has a structure with alternating variational layers and entanglement layers, each layer of the quantum machine learning model contains a fixed number of parameters, a fixed number of data is selected from the hidden state as the parameters of each layer of the quantum machine learning, and different moments correspond to different layer numbers of the quantum machine learning.

[0019] As a preferred technical solution, when the training data is only classical data, the qubits of the quantum machine learning model are reset to |0 n ><0 n |, the target quantum state after quantum machine learning calculation is:

[0020]

[0021] where θ t is the quantum gate parameter of the t-th layer of the quantum machine learning, and ρ t (θ t) represents the target quantum state at the $t$-th moment after quantum machine learning calculation. The quantum calculation process is characterized by the operations of a number of single-qubit gates and multi-qubit gates, denoted as $C$ represents the complex number field, $n$ is the number of qubits, represents the quantum machine learning calculation process of the $t$-th layer, represents the unitary inverse transformation of the operation operator of the quantum machine learning of the $t$-th layer; assume the quantum gate parameter $\theta$ t The number of is $N$. Through mathematical transformation, the output of the classical neural network is converted into a feature of length $N$ as $\theta$ t .

[0022] As a preferred technical solution, when the training data is only quantum data, the target quantum state obtained after calculation at the initial moment is:

[0023]

[0024] where $\theta$ 1 is the quantum gate parameter of the first layer of quantum machine learning, $\rho$ e is the input quantum data, $\rho$ 1 $(\theta$ 1 ) is the quantum state obtained after the initial state passes through the first layer of quantum machine learning model. The quantum calculation process is characterized by the operations of a number of single-qubit gates and multi-qubit gates, denoted as $C$ represents the complex number field, $n$ is the number of qubits;

[0025] For subsequent moments, the calculated target quantum state is:

[0026]

[0027] where $\theta$ t is the quantum gate parameter of the $t$-th layer of quantum machine learning, $\rho$ t $(\theta$ t ) represents the target quantum state at the $t$-th moment after quantum machine learning calculation, represents the quantum machine learning calculation process of the $t$-th layer, represents the unitary inverse transformation of the operation operator of the quantum machine learning of the $t$-th layer.

[0028] As a preferred technical solution, the classical loss is calculated using the mean square error loss function:

[0029]

[0030] where $M$ is the number of samples, is the predicted output of the classical neural network, $y$ T is the true value;

[0031] The quantum loss is calculated using the mean squared error loss function:

[0032]

[0033] where is the estimated value of the observable of quantum machine learning for the final state ρ T (θ), and ρ T (θ) is the quantum state obtained after T layers of calculation;

[0034] The weighted joint loss function is expressed as:

[0035] L = αL c +(1 - α)L q

[0036] where α ∈ [0, 1] is the weighting coefficient of classical and quantum losses.

[0037] As a preferred technical solution, the update of the classical neural network parameter gradient is specifically as follows: The gradient of the classical neural network parameter is calculated using the gradient backpropagation algorithm. Assuming the classical neural network parameter is represented as ν, the stochastic gradient descent method is used to update the parameter:

[0038]

[0039] where β is the learning rate hyperparameter and L is the weighted joint loss function.

[0040] According to the second aspect of the present invention, an electronic device is provided, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.

[0041] According to the third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0042] Compared with the prior art, the present invention uses the output of the classical neural network as the parameter of quantum machine learning, avoiding the problem of large training overhead of quantum machine learning; using the output correlation characteristics of classical machine learning algorithms and using them as quantum machine learning parameters to effectively overcome the potential gradient disappearance problem of quantum machine learning algorithms; making full use of the characteristics of quantum and classical machine learning, having the potential to improve the time series prediction ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the flowchart of the method of the present invention;

[0044] Figure 2 is the overall structure diagram of the interactive hybrid quantum-classical machine learning model;

[0045] Figure 3 Schematic diagram of encoding classical data by a classical neural network;

[0046] Figure 4 Quantum circuit diagram of quantum machine learning;

[0047] Figure 5 Schematic diagram of parameters output from a classical neural network to a quantum machine learning model. Specific implementation manners

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0049] In this application, referring to "embodiment" means that specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in this application may be combined with other embodiments without conflict.

[0050] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those of ordinary skill in the technical field to which this application belongs. The "one", "a", "kind", "the" and other similar words involved in this application do not indicate a quantity limit and may represent a singular or plural number. The terms "including", "comprising", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0051] Embodiment 1

[0052] This embodiment provides an interactive hybrid quantum-classical machine learning model training method. Classical data is used as the input of the classical neural network, and quantum data such as quantum states is used as the input of the quantum machine learning model. Only the classical neural network part of the hybrid quantum-classical machine learning model contains training parameters, and the quantum machine learning algorithm part does not contain training parameters. The quantum gate parameters of the quantum machine learning algorithm are the output of the classical neural network. For classical data, first, the classical neural network processes it, and the obtained output is used as the parameters of the quantum machine learning and the quantum computing process is completed. For quantum data, the input of the classical neural network is fixed, and its output is still used as the parameters of the quantum machine learning. The quantum machine learning directly loads the quantum state input and completes the quantum computing process. The loss function is calculated by the weighted sum of the loss of the classical neural network and the loss of the quantum machine learning. The gradient of the classical neural network parameters is calculated through the neural network backpropagation algorithm, and the classical neural network parameters are updated.

[0053] Specifically, as Figure 1 shown, the method includes the following steps:

[0054] S1, Data loading and model initialization: Load the training data and initialize the classical neural network and the quantum machine learning model.

[0055] First, determine the type of training data. The type of training data includes classical data (classical images or other modal data, such as speech sequence data, etc.) and quantum data.

[0056] If the training data only includes classical data, serialize it into sequence data {x 1 ,…x t …,x n}, and load it through the classical neural network. This process requires the help of a classical memory. In addition, the qubits of the quantum machine learning model need to be reset to |0 n ><0 n |, that is, the input of the quantum machine learning model does not contain any information.

[0057] If the training data only includes quantum data ρ e , load it using the quantum machine learning model. This process requires the help of a quantum memory. In addition, randomly generate fixed classical data {r 1 ,r 2 ,…,r t ,…,r n}, that is, the input of the classical neural network does not contain any information.

[0058] If the training data includes both classical data and quantum data, the classical data is serialized and converted into sequence data, which is loaded through a classical neural network, and the quantum data is loaded through a quantum machine learning model. This process requires the use of a quantum memory and a classical memory respectively.

[0059] S2. Interactive quantum-classical machine learning calculation: Input the classical data in the training data into the classical neural network, perform the classical calculation process, and obtain the classification result as the prediction output; Input the quantum data in the training data into the quantum machine learning model, and use the output of the classical neural network as the parameter of the quantum machine learning model. Then, perform the quantum calculation process and the ordinary quantum measurement process in sequence to obtain the expected value of the qubit under the observable as the prediction output.

[0060] In this embodiment, the quantum neural network can be mathematically characterized as a parameterized unitary matrix, and its structure is not limited in any way. The classical neural network structure mainly adopts a sequence structure model, such as a recurrent neural network, a Transformer network, etc. For noisy intermediate-scale quantum computing, the quantum machine learning model usually adopts a hardware-efficient quantum circuit as the network structure. The interaction algorithm mainly uses the output of the classical neural network as the parameter of the quantum machine learning model, makes full use of the memory correlation characteristics between the outputs of the classical neural network, realizes the constraint on the parameter space of the quantum machine learning model, and overcomes the quantum gradient disappearance problem. The overall structure of the specific interactive hybrid quantum-classical machine learning model is as Figure 2 shown.

[0061] For classical data, a sequence structure classical neural network is used to learn the data and obtain the output. The qubits of the quantum computer are reset, and the output of the classical neural network is used as the parameter of the quantum neural network. Quantum calculation and classical calculation are performed simultaneously, where the quantum calculation measures the Pauli Z operator, estimates the expected value of the observable, and makes a prediction output. The classical neural network directly obtains the prediction output.

[0062] The specific structure of the classical input data encoded by the classical neural network is Figure 3 shown. The models of the classical neural network are ordinary recurrent neural network models, LSTM network models, and GRU network models. The present invention is not limited to the 3 recurrent neural network models as Figure 3 shown, and other sequence neural network models such as the Transformer model can also be used.

[0063] For quantum data, after the quantum computer is loaded, fixed classical data is randomly generated and loaded into the classical neural network to obtain the output of the classical neural network. The output of the classical neural network is used as the parameter of the quantum machine learning model. At the same time, quantum computing and classical computing are performed. Among them, quantum computing measures the Pauli Z operator, and its expected value is used as the prediction output. The classical neural network directly obtains the prediction output.

[0064] Specifically, a classical neural network is used to calculate the serialized classical data, encode the classical data in the hidden state, obtain the hidden state encoded at each moment, and obtain the output of the classical neural network according to the hidden state at the current moment; among them, the number of network layers of the classical neural network is the same as the time series length of the serialized classical data. The calculation result of the classical neural network does not need to be measured and is directly used as the prediction output. In this embodiment, the long short-term memory recurrent neural network (LSTM) is taken as an example for detailed description.

[0065] The mathematical calculation formula of LSTM can be briefly described as:

[0066] c t ,h t =LSTMCerll(x t ,c t-1 ,h t-1 )

[0067] In the above formula, x t is the classical data at the t-th moment, c t-1 ,h t-1 respectively represent the hidden states at the previous moment, which are used to construct memory, c t ,h t are the hidden states at the current moment t. LSTMCell is a single classical neural network calculation unit, which is used to encode the input data x at the current moment and calculate the hidden state at the next moment according to the hidden state at the previous moment.

[0068] The calculation units inside LSTMCell follow the following calculation rules:

[0069] i t =σ(W ii x t +b ii +W hi h t-1 +b hi )

[0070] f t =σ(W if x t +b if +W hf h t-1 +b hf)

[0071] g t = tanh(W ig x t + b ig + W hg h t-1 + b hg )

[0072] o t = σ(W io x t + b io + W ho h t-1+ b ho )

[0073] c t = f t ☉ c t-1 + i t ☉ g t

[0074] h t = o t ☉ tanh(c t )

[0075] In the above formula, W and b are parameters of the classical neural network, and σ(·) is the Sigmoid function. The LSTM neural network consists of T LSTM Cells, and the input of each cell corresponds to the input data at each moment. At the initial moment, c 0 , h 0 are computational tensors initialized following the Gaussian distribution. When batch training is adopted, the input data contains the batch dimension. At this time, the hidden layer output of each cell contains the batch dimension. At this time, the dimension array of the h t tensor is [B, H out , where B is the batch size and H out is the hidden feature dimension. Therefore, quantum machine learning needs to run B copies of quantum circuits simultaneously. Each layer of each quantum circuit only accepts 1 copy of parameters, and the number of its parameters is less than H out . For the second moment, similarly, each layer of each quantum circuit only accepts 1 copy of parameters. And so on, for the final moment, the last layer of quantum machine learning only needs to accept 1 copy of parameters.

[0076] If the length of the time series is T, then T LSTM Cell functions are needed to learn classical data. When T LSTM Cell functions are connected in series, they form an LSTM neural network.

[0077] In practice, h t can also be used as the output at the current moment, that is, y t = ht , where y t is used for prediction. At this time, the classical input data is encoded in the hidden state.

[0078] When the input data is quantum data, expressed as ρ e , the input quantum state is directly loaded from the quantum system into the quantum computer, facilitating direct processing by the quantum machine learning algorithm. At this time, a sequence {r 1 , r 2 , …, r t , …, r n} is randomly generated for classical data, which contains no information. The same function in LSTMCell is also used to learn the randomly generated sequence {r 1 , r 2 , …, r t , …, r n} to obtain the encoded state h t at each moment. At this time, the randomly generated classical data, its length depends on the depth of the quantum machine learning. If the depth is T, the length of the randomly initialized classical sequence data is T. At this time, the batch size of the classical input data is set to 1, which contains no information and only serves as the data source for driving the classical neural network. An LSTM neural network is obtained by connecting T LSTMCell in series, and the hidden layer state h t at each moment is calculated. At this time, the dimension array of h t is [1, H out . The quantum computer only needs to run 1 copy of the quantum circuit, and the number of parameters in each layer of the quantum circuit is less than H out .

[0079] In this embodiment, the specific quantum machine learning structure is as Figure 4 shown. Since the quantum machine learning parameters are the outputs of the classical neural network at each moment, its structure needs to be restricted. The maximum number of parameters in each layer of the quantum machine learning algorithm does not exceed the dimension of h t . A most intuitive quantum machine learning structure is: a structure where the variational layer and the entanglement layer appear alternately. At this time, each layer of the quantum machine learning contains a fixed number of parameters. Therefore, only a fixed number of data need to be selected from h t as the parameters of each layer of the quantum machine learning. At this time, each moment t corresponds to the number of the corresponding layer of the quantum machine learning.

[0080] In quantum computing, quantum machine learning is characterized by the operations of a number of single-qubit gates and multi-qubit gates, which is expressed as C represents the complex number field, and n is the number of qubits.

[0081] When the training data is only classical data, the qubits of the quantum machine learning model are reset to |0 n ><0 n |, the target quantum state after quantum machine learning calculation is:

[0082]

[0083] where θ t is the quantum gate parameter of the t-th layer of quantum machine learning, and ρ t (θ t ) represents the target quantum state at the t-th moment after quantum machine learning calculation, represents the calculation process of the t-th layer of quantum machine learning, represents the unitary inverse transformation of the operation operator of the t-th layer of quantum machine learning. Assuming the number of quantum gate parameters θ t is N, the output h t of the classical neural network can be transformed into a feature of length N as θ t .

[0084] The structure of

[0085]

[0086] can be an alternating structure of variable layers and entanglement layers, which is easy to implement in quantum hardware in the noisy intermediate-scale framework. Mathematically, it can be expressed as: zz In the above formula, R z is a two-bit entanglement gate, which can be implemented by two CNOT gates and one R Figure 4 gate. Each layer of the calculation module of the quantum machine learning model contains N = 3n - 1 parameters. The present invention is not limited to the quantum machine learning model characterized by the above formula, and any other structure can be a quantum machine learning model. As Figure 4 shown in a, a quantum machine learning model with an alternating structure of z-y-z variable layers and entanglement layers is easy to deploy and implement in a quantum computer. Figure 4 a shows a quantum machine learning model using CNOT gates as the entanglement layer. zz b shows a quantum machine learning model using R

[0087] When the quantum gates of the T layers of quantum machine learning are calculated, the quantum state after the T-th moment is ρ T . After the last calculation unit of the classical neural network is calculated, the classical state at the last moment is h T .

[0088] When the training data is only quantum data, quantum machine learning directly learns and processes the loaded quantum data. The target quantum state calculated at the initial moment is:

[0089]

[0090] Among them, θ 1 is the quantum gate parameter of the first-level quantum machine learning, ρ e is the input quantum data, ρ 1 (θ 1 ) is the quantum state obtained after the initial state passes through the first layer of quantum machine learning model.

[0091] For subsequent moments, the calculated target quantum state is:

[0092]

[0093] Among them, θ t is the quantum gate parameter of the t-th layer quantum machine learning, ρ t (θ t ) represents the target quantum state at time t after quantum machine learning calculation.

[0094] When the calculation is done at time T, the quantum state is ρ T . The mathematical model is the same as that in the classical data. The input data of the classical neural network is randomly generated Gaussian data, which does not contain any information, but the classical final state h at the final moment can also be obtained. T .

[0095] S3, joint loss function calculation: calculate the classical loss based on the predicted output of the classical neural network and the true value of the data set, calculate the quantum loss based on the predicted output of quantum machine learning and the true value of the data set, and calculate the weighted joint loss function based on the classical loss and quantum loss.

[0096] Final state ρ for quantum machine learning T Perform multiple measurements to obtain the expected value of the observable Pauli operator O <o>The output state of the classical neural network is h T It can be directly read without measurement. According to the true value label y of the data set T , the loss value is calculated using the mean square error loss or cross-entropy loss function. In this embodiment, the classical loss is calculated using the mean square error loss function:

[0097]

[0098] where M is the number of samples, is the predicted output of the classical neural network, which is the hidden state of the last computational unit, that is, there is y T is the true value label of the data set, such as the category of the image, etc.

[0099] The quantum loss is also calculated using the mean square error loss function:

[0100]

[0101] where, <o> πT Estimate of the observable for the final state ρ T (θ), where ρ T (θ) is the quantum state obtained after T layers of calculations.

[0102] Then, the weighted joint loss function is expressed as:

[0103] L = αL c +(1 - α)L q

[0104] where α ∈ [0, 1] is the weighting coefficient for the classical and quantum losses.

[0105] S4, Update of classical neural network parameters: Based on the weighted joint loss function, use the gradient backpropagation algorithm to update the classical neural network parameters.

[0106] Because in the interactive hybrid quantum-classical machine learning model, only the classical neural network contains trainable parameters, the quantum machine learning algorithm does not contain any parameters, and the parameters of the quantum machine learning come from the output of the hidden layer of the classical neural network. The parameters of the classical neural network can be efficiently completed on a classical computer. Using the gradient backpropagation algorithm to calculate the classical neural network parameters, the gradient value of the loss value with respect to the classical neural network parameters can be obtained in polynomial time. Assuming the classical neural network parameters are represented by ν, use the stochastic gradient descent method to update the parameters:

[0107]

[0108] where β is the learning rate hyperparameter and L is the weighted joint loss function.

[0109] S5, Iteration termination determination: When the weighted joint loss function converges or reaches the maximum number of iterations, end the training; otherwise, according to the updated classical neural network parameters, return to the interactive quantum-classical machine learning calculation steps for the next iteration.

[0110] In this embodiment, the judgment criterion for the convergence of the weighted joint loss function is: Use the classical register collection algorithm to accumulate the loss value during the iteration process, calculate the absolute value difference of the loss values in two consecutive rounds of iterations. If the difference is less than 0.0001, it is considered that the loss value converges.

[0111] Embodiment 2

[0112] The electronic device of the present invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus. The electronic device of the present invention further includes a quantum processing unit (QPU), which can also be referred to as a quantum processor or a quantum chip, and can relate to a physical chip including a plurality of qubits interconnected in a specific manner.

[0113] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0114] The processing unit executes the various methods and processes described above, such as methods S1 to S5. For example, in some embodiments, methods S1 to S5 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S5 described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute methods S1 to S5 by any other appropriate means (e.g., by means of firmware).

[0115] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0116] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or a controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or the controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0117] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0118] It should be understood that various forms of the flow shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this embodiment can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this embodiment can be achieved, and no limitation is made herein.

[0119] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.< / o> < / o>

Claims

1. An interactive hybrid quantum-classical machine learning model training method, characterized in that: The method comprises the following steps: Data loading and model initialization: load training data and initialize classical neural network and quantum machine learning models; Interactive quantum-classical machine learning calculation: input the classical data in the training data into the classical neural network, execute the classical calculation process, and obtain the classification result as the prediction output; input the quantum data in the training data into the quantum machine learning model, and use the output of the classical neural network as the parameter of quantum machine learning, execute the quantum calculation process and the ordinary quantum measurement process in sequence, and obtain the expected value of the quantum bit under the observable measurement as the prediction output; Joint loss function calculation: Calculate the classical loss based on the predicted output of the classical neural network and the true value of the data set, calculate the quantum loss based on the predicted output of quantum machine learning and the true value of the data set, and calculate the weighted joint loss function based on the classical loss and quantum loss; Classical neural network parameter gradient update: Based on the weighted joint loss function, the gradient back propagation algorithm is used to update the parameters of the classical neural network; Iteration termination judgment: When the weighted joint loss function converges or reaches the maximum number of iterations, the training ends; otherwise, according to the updated classical neural network parameters, the interactive quantum-classical machine learning calculation step is returned to perform the next round of iteration.

2. An interactive hybrid quantum-classical machine learning model training method according to claim 1, characterized in that: The data loading and model initialization are specifically as follows: Determining a training data type, where the training data type includes classical data and quantum data; If the training data only includes classical data, it is serialized and converted into sequence data, loaded through the classical neural network, and the quantum bits of the quantum machine learning model are reset, that is, the input of the quantum machine learning model does not contain any information; If the training data only includes quantum data, the quantum machine learning model is used to load it and randomly generate fixed classical data, that is, the input of the classical neural network does not contain any information; If the training data includes both classical data and quantum data, the classical data is serialized, converted into sequence data, loaded through the classical neural network, and the quantum data is loaded through the quantum machine learning model.

3. The interactive hybrid quantum-classical machine learning model training method according to claim 1, characterized in that: In the interactive quantum-classical machine learning calculation step, a classical neural network is used to calculate the serialized classical data, the classical data is encoded in a hidden state, the hidden state encoded at each moment is obtained, and the output of the classical neural network is obtained according to the hidden state at the current moment; wherein the number of network layers of the classical neural network is the same as the time series length of the serialized classical data.

4. An interactive hybrid quantum-classical machine learning model training method according to claim 3, characterized in that: In the interactive quantum-classical machine learning calculation step, the quantum machine learning model is mathematically characterized as a parameterized unitary matrix, the parameters of the quantum machine learning model are set to the output of the classical neural network at each moment, and the maximum number of parameters of each layer of the quantum machine learning model does not exceed the dimension of the encoded hidden state. The quantum computing process is performed according to the set parameters, and the Pauli Z operator is measured to estimate the expected value of the observable as the predicted output. When the quantum machine learning model is a structure in which variable layers and entangled layers appear alternately, each layer of the quantum machine learning model contains a fixed number of parameters, and a fixed number of data is selected from the hidden state as the parameters of each layer of the quantum machine learning, and different moments correspond to the numbers of different layers of the quantum machine learning.

5. The interactive hybrid quantum-classical machine learning model training method according to claim 1, characterized in that: When the training data is only classical data, reset the quantum bits of the quantum machine learning model to |0 n ><0 n |, the target quantum state after quantum machine learning calculation is: Among them, θ t is the quantum gate parameter of the t-th layer quantum machine learning, ρ t (θ t ) represents the target quantum state at time t after quantum machine learning calculation. The quantum computing process is characterized by the operation of several single-qubit gates and multi-qubit gates, expressed as C represents the complex field, n is the number of quantum bits, represents the t-th layer quantum machine learning calculation process, represents the unitary inverse transformation of the t-th layer quantum machine learning operator; assuming that the quantum gate parameter θ t The number of is N, and the output of the classical neural network is converted into a feature of length N as θ through mathematical transformation. t .

6. The interactive hybrid quantum-classical machine learning model training method according to claim 1, characterized in that: When the training data is only quantum data, the target quantum state calculated at the initial moment is: Among them, θ1 is the quantum gate parameter of the first layer of quantum machine learning, ρ e is the input quantum data, ρ1(θ1) is the quantum state obtained after the initial state passes through the first layer of quantum machine learning model, and the quantum computing process is characterized by the operation of several single-qubit gates and multi-qubit gates, expressed as C represents the complex field, n is the number of quantum bits; For subsequent moments, the calculated target quantum state is: Among them, θ t is the quantum gate parameter of the t-th layer quantum machine learning, ρ t (θ t ) represents the target quantum state at time t after quantum machine learning calculation, represents the t-th layer quantum machine learning calculation process, Represents the inverse unitary transformation of the t-th layer quantum machine learning operation operator.

7. The interactive hybrid quantum-classical machine learning model training method according to claim 1, characterized in that: The classical loss is calculated using the mean square error loss function: Where M is the number of samples, is the predicted output of the classic neural network, y T is the true value; The quantum loss is calculated using the mean square error loss function: in, The observable quantity for quantum machine learning is the final state ρ T The estimated value of (θ), ρ T (θ) is the quantum state obtained by T-layer calculation; The weighted joint loss function is expressed as: L=αL c +(1-α)L q Among them, α∈[0,1] is the weighting coefficient of classical and quantum losses.

8. The interactive hybrid quantum-classical machine learning model training method according to claim 1, characterized in that: The classical neural network parameter gradient update is specifically as follows: the gradient back propagation algorithm is used to calculate the gradient of the classical neural network parameters. Assuming that the classical neural network parameters are represented by v, the stochastic gradient descent method is used to update the parameters: Among them, β is the learning rate hyperparameter and L is the weighted joint loss function.

9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Neural network construction method and training method based on heterogeneous quantum computing resources

    CN121902850A