Quantum machine learning training methods and electronic devices

JP2026137625AActive Publication Date: 2026-08-27HON HAI PRECISION INDUSTRY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025027780
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-17
Filing Date
2025-02-25
Publication Date
2026-08-27
Estimated Expiration
2045-02-25

AI Technical Summary

Benefits of technology

【0014】 本発明は、量子回路を組み合わせてニューラルネットワークのモデルパラメータをトレーニングする。この方法により、トレーニングが必要なパラメータの数を減らすことができ、トレーニングされたニューラルネットワークは、従来のコンピュータにも適用することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026137625000001_ABST
    Figure 2026137625000001_ABST
Patent Text Reader

Abstract

Traditional machine learning models are very large, require training on many parameters, and the methods for encoding data into quantum states are relatively complex. [Solution] The present invention provides a training method and electronic apparatus for quantum machine learning. The training method includes the steps of setting up a quantum circuit and outputting multiple probability values ​​of multiple qubits, the quantum circuit comprising multiple gates, these gates having multiple circuit parameters; mapping the qubits to multiple model parameters of a neural network, the qubits being used to compute multiple basis sets, the number of which is greater than or equal to the number of model parameters; inputting data into the neural network and calculating a loss based on the output of the neural network; and updating the circuit parameters in the quantum circuit based on the loss.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a training method that combines quantum computing and machine learning. [Background technology]

[0002] Machine learning techniques face significant challenges in terms of the increasing complexity of models and the efficient use of computational resources. As machine learning models become more complex, the number of required parameters increases dramatically, leading to problems with the inefficiency of resource use during training and deployment. Furthermore, with the rise of Quantum Machine Learning (QML), this field faces additional technical challenges, including methods for encoding conventional data into quantum states. For example, when processing a single image, each pixel in the image can be encoded into a corresponding quantum state, and inference can be performed using quantum computation. However, as image resolution increases, complex transformations are required to convert pixels into quantum states, and such methods may have limitations on the input size that can be processed and may result in information loss. [Overview of the project] [Problems that the invention aims to solve]

[0003] Traditional machine learning models are very large, require training on many parameters, and the methods for encoding data into quantum states are relatively complex. [Means for solving the problem]

[0004] The present invention provides a quantum machine learning training method applicable to electronic devices. The training method includes the steps of: setting up a quantum circuit and outputting multiple probability values ​​for multiple qubits, wherein the quantum circuit includes multiple gates, and the multiple gates have multiple circuit parameters; mapping the multiple qubits to multiple model parameters of a neural network, wherein the qubits are used to compute multiple basis sets, the number of which is greater than or equal to the number of which is which is which; inputting data into the neural network and calculating a loss based on the output of the neural network; and updating the multiple circuit parameters in the quantum circuit based on the loss.

[0005] In one embodiment of the present invention, the plurality of gates includes a plurality of controlled NOT gates and a plurality of rotary gates, and the plurality of circuit parameters correspond to the plurality of rotary gates.

[0006] In one embodiment of the present invention, the step of mapping the plurality of qubits to the plurality of model parameters of the neural network includes determining a first basis for the plurality of qubits, calculating a first probability value for the first basis, generating a first vector based on the first basis and the first probability value, and inputting the first vector into the mapping neural network to obtain one of the plurality of model parameters.

[0007] In one embodiment of the present invention, the step of mapping the plurality of qubits to the plurality of model parameters of the neural network further includes determining a second basis for the plurality of qubits, calculating a second probability value for the second basis, determining that the second basis is different from the first basis, generating a second vector based on the second basis and the second probability value, and inputting the second vector into the mapping neural network to obtain another one of the plurality of model parameters.

[0008] In one embodiment of the present invention, the training method further includes the step of updating a plurality of model parameters in the mapping neural network based on the loss.

[0009] In one embodiment of the present invention, generating a first vector based on the first basis and the first probability value includes amplifying the first probability value and generating the first vector based on the amplified first probability value.

[0010] In one embodiment of the present invention, the amplification of the first probability value is performed based on the following formula: P j →α×tanh(α×2 N-1 P j ) In the formula, P j is the first probability value, α is a constant, j is a positive integer, and N is the number of the plurality of qubits.

[0011] In one embodiment of the present invention, the step of mapping the plurality of qubits to the plurality of parameters of the neural network includes determining a first basis for the plurality of qubits, calculating a first probability value for the first basis, and inputting the first probability value into a mapping function to obtain one of the plurality of model parameters.

[0012] In one embodiment of the present invention, the training method further includes the step of selecting a portion of the plurality of model parameters and changing the sign of the portion.

[0013] To put it another way, embodiments of the present invention provide an electronic device including a memory and a processor. The memory stores a plurality of instructions, and the processor is communicatively connected to the memory and executes the instructions in the memory to complete the training method described above. [Effects of the Invention]

[0014] The present invention trains the model parameters of a neural network by combining quantum circuits. By this method, the number of parameters that need to be trained can be reduced, and the trained neural network can also be applied to conventional computers.

Brief Description of the Drawings

[0015] [Figure 1] It is an explanatory diagram showing an electronic device according to an embodiment. [Figure 2] It is an explanatory diagram showing the operations of a quantum circuit and a neural network according to an embodiment. [Figure 3] It is a flowchart showing a training method for quantum machine learning according to an example. [Figure 4] It is an explanatory diagram showing the training of quantum machine learning according to another embodiment.

Embodiments for Carrying Out the Invention

[0016] To make the above features and advantages of the present invention clearer and easier to understand, embodiments will be specifically cited below and described in detail together with the accompanying drawings.

[0017] Some embodiments of the present invention will be described in detail below with reference to the accompanying drawings. When the reference numerals of members cited in the following description appear in different drawings, they are regarded as the same or similar members. These embodiments are only a part of the present invention and do not disclose all possible implementation manners of the present invention. Rather, these embodiments are merely examples of the systems and methods within the scope of the claims of the present invention.

[0018] <tmp The terms "first", "second", etc. used herein do not particularly refer to any order or rank, and are only used to distinguish members or operations described by the same technical terms.

[0019] FIG. 1 is an explanatory diagram showing an electronic device according to an embodiment. Referring to FIG. 1, the electronic device 100 may be a tablet computer, a personal computer, a notebook computer, a server, a distributed computer, a cloud server, an industrial computer, or various electronic devices having computing capabilities, etc., and the present invention is not limited thereto. The electronic device 100 includes a processor 110 and a memory 120. The processor 110 is communicatively connected to the memory 120, and the communication connection can be realized via any wired or wireless communication means, or can also be realized via the Internet. The processor 110 may be a central processor, a microprocessor, a microcontroller, a deep-learning processing unit (DPU), a neural network processing unit (NPU), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), a programmable logic device (PLD), etc. The memory 120 may be a random access memory, a read-only memory, a flash memory, a floppy disk, a hard disk, an optical disk, a USB flash drive, a tape, or a database accessible via the Internet, and stores a plurality of instructions. The processor 110 executes them to complete a quantum machine learning training method, and the method will be described below.

[0020] FIG. 2 is an explanatory diagram showing the operations of a quantum circuit and a neural network according to an embodiment. FIG. 3 is a flowchart showing a quantum machine learning training method according to an example. Referring to FIGS. and FIG. 3, this embodiment mainly relates to a quantum circuit 210, a mapping neural network 220, and a neural network 230. A total of N quantum bits are used in the quantum circuit in which these quantum bits are 2 NIt is used to compute basis vectors. On the other hand, the neural network 230 contains M model parameters. The neural network 230 may include a convolutional neural network (CNN), a recurrent neural network (RNN), a transformer, an attention layer, etc., and the present invention does not limit the configuration of the neural network 230. In particular, each model parameter corresponds to one basis vector, and the number of basis vectors is greater than or equal to the number of model parameters, that is, the number of basis vectors calculated by an exponential function corresponding to the number of qubits is greater than or equal to the number of model parameters, for example, 2 N ≥ M. In some embodiments, N = [log2M], but the present invention is not limited thereto. In other words, the above basis can cover M model parameters. The mapping neural network 220 is used to map N qubits to M model parameters. For example, the probability value of each basis corresponds to one model parameter, the range of probability values ​​is from 0 to 1, the model parameter can be a negative value, and the mapping neural network 220 can map the range from 0 to 1 to the range from -∞ to ∞.

[0021] In step 301, the quantum circuit 210 is set up to output the probability values ​​of multiple qubits. Specifically, the input to the quantum circuit 210 is the initial state 211 of the qubits, and these initial states 211 are, for example, |0〉, but the present invention is not limited to this. The quantum circuit 210 includes multiple gates, for example, multiple rotation gates 212 and multiple controlled-NOT gates (CNOT gates) 213. The rotation gate 212 is represented by the following equation 1. [Formula 1]

number

[0022] Here, μ is referred to as a circuit parameter. In this embodiment, the rotation gate 212 rotates around the Y axis, but in other embodiments, it may rotate around the X axis or the Z axis. Alternatively, in other embodiments, the quantum circuit 210 may include a swap gate, a controlled Z gate (CZ gate), a Hadamard gate, etc., and the present invention is not limited thereto. The embodiment in Figure 2 shows gates within a single block, that is, a rotation gate and a controlled NOT gate are placed in the layer corresponding to each qubit. However, such a block can be repeated multiple times, and the present invention does not limit the number of repetitions.

[0023] The output of the quantum circuit 210 is the probability value 214 of the qubit, and the probability value of each basis can be calculated through measurement. Here, the embodiment simulates the quantum circuit 210 using software, but in other embodiments the quantum circuit 210 may be implemented using a physical circuit, and the present invention is not limited thereto.

[0024] In step 302, the qubits are mapped to the model parameters of the neural network. In some embodiments, a vector 221 can be generated based on one basis and its corresponding probability value, which can then be used as input to the mapping neural network 220. For example, if there are seven qubits and one of the basis is "0100100", the probability value of this basis is |〈0100100|Ψ〉| 2This is expressed as =0.023, where Ψ is the quantum state. Therefore, the formed vector 221 is [0,1,0,0,1,0,0,0.023] and its length is 8. The mapping neural network 220 is, for example, a multilayer perceptron (MLP), but the present invention is not limited to this. The output 222 of the mapping neural network 220 is one model parameter (also called a weight) in the neural network 230. Since the mapping neural network 220 outputs one model parameter each time, it is necessary to repeatedly input different basis vectors into the mapping neural network 220 to calculate different model parameters. For example, another basis is "0100101", and the probability value of this basis is 0.017. Based on this basis and probability value, another vector 221 can be generated, which is represented as [0,1,0,0,1,0,1,0.017]. This vector 221 is input into the mapping neural network 220 to obtain another model parameter.

[0025] In summary, the calculations of the mapping neural network 220 can be expressed as shown in equation 2 below. [Formula 2]

number

[0026] JPEG2026137625000004.jpg41164

[0027] Since the sum of the probability values ​​of all basis vectors is 1, these probability values ​​are usually small. In some embodiments, the probability values ​​of each basis vector can be amplified, and a vector 221 can be generated based on the amplified probability values. For example, the probability values ​​can be amplified based on the following equation 3. [Formula 3]

number

[0028] Here, P j is the aforementioned first probability value. α is a constant that can be determined experimentally. j is a positive integer. In other embodiments, the probability value may be input to an arbitrary function for amplification, which may include, but is not limited to, a polynomial function, an exponential function, and the like.

[0029] In step 303, data 231 is input to the neural network 230, and a loss 240 is calculated based on the output 232 of the neural network 230. The data 231 may be image data, text data, audio data, binary data, or any other type of data, and the present invention is not limited thereto. The output 232 may contain one or more numerical values. If the neural network 230 is used to process a classification problem, the output 232 represents one category. If the neural network 230 is used to process a regression problem, the output 232 is numerical. If the neural network 230 is used as a generative model, the output 232 may be image data, text data, audio data, binary data, etc. The present invention does not limit the content of the output 232.

[0030] Taking the classification problem as an example, the loss function can be expressed using the following equation 4. [Equation 4]

number

[0031] JPEG2026137625000007.jpg35165

[0032] JPEG2026137625000008.jpg81168

[0033] On the other hand, the trained neural network 230 can be applied to a regular computer (no quantum computer is required), and the input to the neural network 230 is conventionally used data, so there is no need to encode this data into qubits, thus reducing hardware requirements.

[0034] Figure 4 is an explanatory diagram showing the training of quantum machine learning according to another embodiment. Referring to Figure 4, the difference between Figure 4 and Figure 2 is that the mapping neural network 220 is replaced by a mapping function 410. When the base probability value 411 is input into the mapping function 410, one model parameter 412 of the neural network 230 is obtained. For example, a part of the model parameters can be made to correspond to one base, and another part of the model parameters can be made to correspond to multiple bases. When corresponding to one base, this model parameter is the same as the probability value corresponding to the base. When corresponding to multiple bases, this model parameter is the same as the average of the probability values corresponding to the bases. In some embodiments, first, 2 N -M model parameters are selected. These model parameters are probability values corresponding to two different bases. At this time, the mapping function is represented as shown in Equation 5 below. [Equation 5] [Number]

[0035] Here, i and k are positive integers. θ j is the j-th model parameter. For other model parameters, they are set to be the same as the probability values corresponding to the bases. At this time, the mapping function is represented as shown in Equation 6 below. [Equation 6] [Number]

[0036] Since the range of the probability value is from 0 to 1, next, negative model parameters are generated. In some embodiments, a part (for example, half) of the model parameters is selected, and the positive and negative signs of this part are changed. For example, when the positive integer j is even, the model parameter θ j is not changed, but when the positive integer j is odd, the model parameter θ jThe sign of the positive or negative value is changed. In this embodiment as well, the model parameter θ j The method for amplifying this is the same as in equation 3 above.

[0037] The mapping function described above is merely an example, and those skilled in the art can design other mapping functions based on the disclosed means, and the present invention is not limited thereto.

[0038] Although the present invention has been disclosed as described above by its embodiments, this does not limit the invention, and those skilled in the art can make some modifications and alterations without departing from the spirit and scope of the invention. Accordingly, the scope of protection of the present invention shall be based on the definitions in the claims described below. [Industrial applicability]

[0039] The method and electronic apparatus provided by the present invention can be used for training any neural network. [Explanation of symbols]

[0040] 100 Electronic equipment 110 processors 120 memory 210 Quantum Circuit 211 Initial state 212 Rotating Gate 213 Controlled NOT Gate 214,411 probability values 220 Mapping Neural Networks 221 Vectors 222,232 output 230 Neural Networks 231 Data 240 loss Steps 301-304 410 Mapping Functions 412 Model Parameters

Claims

1. A training method for quantum machine learning applied to electronic devices, wherein the training method is: The steps include setting up a quantum circuit, outputting multiple probability values ​​for multiple qubits, the quantum circuit including multiple gates, and the multiple gates having multiple circuit parameters, The steps include mapping the plurality of qubits to a plurality of model parameters of a neural network, using the qubits to compute a plurality of basis sets, wherein the number of basis sets is greater than or equal to the number of the plurality of model parameters, The steps include inputting data into the neural network and calculating the loss based on the output of the neural network, The steps include updating the plurality of circuit parameters in the quantum circuit based on the loss, Training methods for quantum machine learning, including [specific data / features].

2. The training method according to claim 1, wherein the plurality of gates include a plurality of controlled NOT gates and a plurality of rotary gates, and the plurality of circuit parameters correspond to the plurality of rotary gates.

3. The step of mapping the plurality of qubits to the plurality of model parameters of the neural network is: The first basis of the plurality of qubits is determined, and the first probability value of the first basis is calculated, To generate a first vector based on the first basis and the first probability value, The first vector is input to the mapping neural network to obtain one of the multiple model parameters, The training method according to claim 1, including the following:

4. The step of mapping the plurality of qubits to the plurality of model parameters of the neural network is: Determine the second basis of the plurality of qubits, calculate the second probability value of the second basis, and confirm that the second basis is different from the first basis. To generate a second vector based on the second basis and the second probability value, The second vector is input to the mapping neural network to obtain another one of the multiple model parameters, The training method according to claim 3, further comprising:

5. The training method according to claim 3, further comprising the step of updating a plurality of model parameters in the mapping neural network based on the loss.

6. Generating a first vector based on the first basis and the first probability value is: Amplifying the aforementioned first probability value, The first vector is generated based on the amplified first probability value, The training method according to claim 3, including the following:

7. The amplification of the first probability value is performed based on the following formula: P j →α×tanh(α×2 N-1 P j ) In the formula, P j The training method according to claim 6, wherein is the first probability value, α is a constant, j is a positive integer, and N is the number of the plurality of qubits.

8. The step of mapping the plurality of qubits to the plurality of model parameters of the neural network is: The first basis of the plurality of qubits is determined, and the first probability value of the first basis is calculated, The first probability value is input into the mapping function to obtain one of the multiple model parameters, The training method according to claim 1, including the following:

9. The training method according to claim 8, further comprising the step of selecting a portion of the plurality of model parameters and changing the sign of the portion.

10. Memory for storing multiple instructions, A processor that is communicatively connected to the aforementioned memory, Includes, The processor executes the plurality of instructions, The steps include setting up a quantum circuit, outputting multiple probability values ​​for multiple qubits, the quantum circuit including multiple gates, and the multiple gates having multiple circuit parameters, The steps include mapping the plurality of qubits to a plurality of model parameters of a neural network, using the qubits to compute a plurality of basis sets, wherein the number of basis sets is greater than or equal to the number of the plurality of model parameters, The steps include inputting data into the neural network and calculating the loss based on the output of the neural network, The steps include updating the plurality of circuit parameters in the quantum circuit based on the loss, An electronic device that completes the process.

11. The electronic device according to claim 10, wherein the plurality of gates include a plurality of controlled NOT gates and a plurality of rotary gates, and the plurality of circuit parameters correspond to the plurality of rotary gates.

12. The step of mapping the plurality of qubits to the plurality of model parameters of the neural network is: The first basis of the plurality of qubits is determined, and the first probability value of the first basis is calculated, To generate a first vector based on the first basis and the first probability value, The first vector is input to the mapping neural network to obtain one of the multiple model parameters, The electronic device according to claim 10, including the following:

13. The step of mapping the plurality of qubits to the plurality of model parameters of the neural network is: Determine the second basis of the plurality of qubits, calculate the second probability value of the second basis, and confirm that the second basis is different from the first basis. To generate a second vector based on the second basis and the second probability value, The second vector is input to the mapping neural network to obtain another one of the multiple model parameters, The electronic device according to claim 12, further comprising:

14. The electronic device according to claim 12, further comprising the step of updating a plurality of model parameters in the mapping neural network based on the loss.

15. Generating a first vector based on the first basis and the first probability value is: Amplifying the aforementioned first probability value, The first vector is generated based on the amplified first probability value, The electronic device according to claim 12, including the following:

16. The amplification of the first probability value is performed based on the following formula: P j →α×tanh(α×2 N-1 P j ) In the formula, P j The electronic device according to claim 15, wherein is the first probability value, α is a constant, j is a positive integer, and N is the number of the plurality of qubits.

17. The step of mapping the plurality of qubits to the plurality of model parameters of the neural network is: The first basis of the plurality of qubits is determined, and the first probability value of the first basis is calculated, The first probability value is input into the mapping function to obtain one of the multiple model parameters, The electronic device according to claim 10, including the following:

18. The electronic device according to claim 17, further comprising the step of selecting a portion of the plurality of model parameters and changing the sign of the portion.