Quantum Transformer-based Classification Method, Device, Medium and Equipment
By introducing quantum computing characteristics and quantum wavelet KAN network into the Transformer model, the problems of high computational complexity and insufficient expression capabilities of the traditional Transformer model are solved, and more efficient computing and stronger interpretability are achieved.
Patent Information
- Application Number
- CN202510297898.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The traditional Transformer model faces the problems of high computational complexity and insufficient expression ability when dealing with highly abstract input information tasks.
By introducing the parallel computing characteristics of quantum computing and the superposition and entanglement properties of quantum states, combining quantum variational lines and quantum wavelet KAN networks, a classification method based on quantum Transformer is constructed.
It improves the computing efficiency of quantum Transformer and the interpretability of the model, improves the resolution of complex modes, and reduces the amount of parameters.
Smart Images

Figure CN119807863B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of quantum computing, and in particular, to a classification method, device, medium and equipment based on quantum Transformer. Background Art
[0002] Since the Transformer model was introduced in 2017, it has become the mainstream architecture in the fields of natural language processing, speech recognition, image classification, etc. Its core multi-head self-attention mechanism (MHSA) and feed-forward neural network (FFN) perform excellently in capturing the global dependencies and feature representations of data. By means of the attention mechanism, the Q, K, V matrices of information such as text and images are finally used to obtain attention scores, thereby improving the accuracy of image recognition.
[0003] With the rapid growth of data scale and model complexity, traditional Transformer faces problems such as high computational complexity and large number of running parameters during training and inference. As a future mainstream computing method, quantum computing combines quantum computing and Transformer by introducing the parallel computing characteristics of quantum computing and the superposition and entanglement properties of quantum states, which can overcome the limitation of the insufficient expression ability of traditional Transformer models when dealing with tasks of highly abstract input information. Summary of the Invention
[0004] The purpose of the present invention is to provide a classification method, device, medium and equipment based on quantum Transformer, which effectively improves the computing efficiency of quantum Transformer and also enhances the interpretability of the model.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] On the one hand, the present invention provides a classification method based on quantum Transformer, including:
[0007] Obtain the feature vector of the object to be classified, and perform quantum encoding on the feature vector to generate a quantum initial state;
[0008] Use quantum variational circuits to respectively construct a quantum query layer, a quantum key layer and a quantum value layer to form a quantum self-attention network;
[0009] Construct a quantum multi-head attention network based on the quantum self-attention network, and input the quantum initial state into the quantum multi-head attention network;
[0010] Use a quantum wavelet KAN network as a feed-forward layer to perform feed-forward processing on the output result of the quantum multi-head attention network;
[0011] A linear layer is used to perform a linear transformation on the result of the feed-forward processing to obtain the recognition result of the object to be classified.
[0012] According to an embodiment of the present invention, the quantum variational circuit is constituted by applying a combined U gate between every two adjacent qubits; the combined U gate includes a CNOT gate configured between two adjacent qubits, an RX rotation gate and an RZ rotation gate configured on the first qubit among two adjacent qubits, and a RY rotation gate configured on the second qubit among two adjacent qubits.
[0013] According to an embodiment of the present invention, the combined U gate is a forward combined U gate, and applying the combined U gate between every two adjacent qubits includes: from to sequentially applying a forward combined U gate to adjacent qubits and , the forward combined U gate includes a CNOT gate configured between qubits and , an RX and an RZ rotation gate configured on qubit , and a RY rotation gate configured on qubit ; or,
[0014] the combined U gate is a reverse combined U gate, and applying the combined U gate between every two adjacent qubits includes: from to sequentially applying a reverse combined U gate to adjacent qubits and , the reverse combined U gate includes a CNOT gate configured between qubits and , an RX and an RZ rotation gate configured on qubit , and a RY rotation gate configured on qubit ;
[0015] wherein, the forward combined U gate is expressed as:
[0016] ;
[0017] the reverse combined U gate is expressed as:
[0018] ;
[0019] represents the identity matrix, and n is the number of qubits.
[0020] According to an embodiment of the present invention, the method further includes:
[0021] The quantum multi - head attention network is trained using the data to be classified in the training set. The training process uses the quantum gradient descent method and the hybrid quantum - classical optimization algorithm to minimize the error function and gradually optimize the parameters of the quantum multi - head attention network; the initial parameters of the quantum multi - head attention network parameters are randomized parameters.
[0022] According to an embodiment of the present invention, the quantum wavelet KAN network includes a plurality of stacked quantum wavelet basis function quantum circuits; the wavelet basis function is a continuous wavelet basis function, and the continuous wavelet basis function can be generally composed of a combination or deformation of a complex sine wave and a Gaussian envelope. The complex sine wave of the continuous wavelet basis function is simulated by an RX gate, and the Gaussian envelope of the wavelet basis function is simulated by an amplitude embedding gate, realizing the construction of continuous wavelets based on quantum circuits. The quantization framework of the continuous wavelet basis can be Morlet, Gabor, Mexican Hat, or Gaussian continuous wavelets, all of which are continuous wavelet basis functions centered on complex sine waves and Gaussian envelopes.
[0023] According to an embodiment of the present invention, the quantum self - attention network is expressed as:
[0024] ;
[0025] wherein, the quantum query layer evolves to form a query matrix , the quantum key layer evolves to form a key matrix , and the quantum value layer evolves to form a value matrix .
[0026] According to an embodiment of the present invention, the quantum multi - head attention network is expressed as:
[0027] ;
[0028] The quantum multi - head attention network performs h - times linear projections on Q, K, and V through different linear projections, executes the quantum self - attention network in parallel, connects the h output results, and uses another learned linear projection to project the connected output results.
[0029] On the other hand, the present invention also provides a classification device based on a quantum Transformer, including:
[0030] A quantum encoding unit, configured to obtain the feature vector of the object to be classified and generate a quantum initial state by quantum - encoding the feature vector;
[0031] The self-attention unit is configured to respectively construct a quantum query layer, a quantum key layer, and a quantum value layer by using a quantum variational circuit to form a quantum self-attention network;
[0032] The multi-head self-attention unit is configured to construct a quantum multi-head attention network based on multiple quantum self-attention networks and input the quantum initial state into the quantum multi-head attention network;
[0033] The feed-forward unit is configured to perform feed-forward processing on the output result of the quantum multi-head attention network by using a quantum wavelet KAN network as a feed-forward layer;
[0034] The output unit is configured to perform a linear transformation on the feed-forward processing result by using a linear layer to obtain the recognition result of the object to be classified.
[0035] On the other hand, the present invention also provides a computer storage medium, in which instructions are stored, and when the instructions are run, the classification method based on the quantum Transformer is implemented.
[0036] On the other hand, the present invention also provides a computing device, characterized in that it includes a processor and a communication interface coupled to the processor; the processor is used to run a computer program or instructions to implement the classification method based on the quantum Transformer.
[0037] The classification method, device, medium, and device based on the quantum Transformer provided by the present invention combine a quantum variational circuit and a quantum wavelet KAN network to implement a quantum Transformer. By incorporating a variational quantum circuit into the Transformer, the construction of the Q, K, and V matrices is completed, and the classical architecture is transformed into a quantum attention architecture; and by using a quantum wavelet KAN network as a feed-forward layer, a learnable quantum gate basis function is used to replace the linear weight. Quantum KAN can approximate the function better with fewer parameters, which not only improves the calculation efficiency but also enhances the interpretability of the model.
[0038] Beneficial effects
[0039] The classification method, device, medium, and device based on the quantum Transformer proposed by the present invention have the following beneficial effects compared with the prior art:
[0040] 1. Improved the accuracy of the quantum Transformer: The quantum Transformer based on the quantum variational circuit and the quantum wavelet KAN network constructs the query layer, key layer, and value layer through the quantum variational circuit, captures long-range dependencies through the quantum superposition state, strengthens the interaction between neighboring features using quantum entanglement, integrates multi-dimensional information, and enhances the model's ability to distinguish complex patterns; the feed-forward layer uses the powerful fitting ability of the quantum wavelet KAN network for non-linear relationships to effectively capture the local and global characteristics of the data and achieve more accurate feature extraction. When the model processes targets in complex datasets such as CIFAR-10, its accuracy is higher than that of general quantum networks. Experiments show that on the CIFAR-10 image dataset, the accuracy of the quantum wavelet KAN reaches 98% with the same number of parameters, a 4% improvement compared to traditional neural networks;
[0041] 2. Reduced the number of parameters: The feed-forward layer of the quantum Transformer uses the quantum wavelet KAN network, saving the number of parameters. Based on the quantum wavelet KAN network, it avoids the fully connected weight matrix of the traditional MLP, fundamentally reducing the parameter complexity. Compared with traditional deep learning models, the quantum wavelet KAN network uses quantum wavelet basis function nodes instead of constant nodes, and the nodes in different layers are connected through activation functions to implement the quantum wavelet KAN network with quantum wavelet basis functions as nodes. While the traditional MLP requires multiple layers of neurons stacked, and each neuron contains multiple weight parameters. The formula for the number of parameters of KAN is: , where is the input dimension, is the parameter of the quantum wavelet basis function, and the number of parameters grows linearly with the input dimension ( ). The formula for the number of parameters of the traditional MLP is: , where is the input dimension, is the width of the hidden layer, and the number of parameters grows quadratically with and ( ). If the input dimension is 100, the number of parameters of the traditional MLP reaches , and the number of parameters of KAN is only The number of parameters of KAN is only 20% of that of the MLP. Each node function can approximate complex non-linear relationships with a small number of parameters, saving the number of parameters;
[0042] 3. Improved interpretability: By using the quantum wavelet KAN network in the feed-forward layer of the quantum Transformer, the behavior of the model becomes more transparent, which helps to understand data features and the decision-making process in practical applications. The quantum wavelet KAN network uses quantum wavelet basis function nodes instead of constant nodes. By decomposing the input variables layer by layer, each node is a learnable quantum wavelet basis function, and the nodes between different layers are connected by activation functions. This design makes the mapping path from input to output clearer, and the influence of each feature can be directly traced through the quantum wavelet basis functions on the path, thus improving interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0044] Figure 1 is a flowchart of a classification method based on a quantum Transformer according to an exemplary embodiment of the present invention;
[0045] Figure 2 is a schematic diagram of a classification device based on a quantum Transformer according to an exemplary embodiment of the present invention;
[0046] Figure 3 is a schematic diagram of the architecture of a quantum Transformer according to an exemplary embodiment of the present invention;
[0047] Figure 4 is a schematic diagram of a quantum circuit of a quantum wavelet basis function according to an exemplary embodiment of the present invention;
[0048] Figure 5 is a schematic diagram of a quantum wavelet KAN network according to an exemplary embodiment of the present invention;
[0049] Figure 6 is a schematic diagram of a V-type quantum variational circuit according to an exemplary embodiment of the present invention;
[0050] Figure 7 is a schematic diagram of a forward combined U gate according to an exemplary embodiment of the present invention;
[0051] Figure 8 is a schematic diagram of a reverse combined U gate according to an exemplary embodiment of the present invention;
[0052] Figure 9 is a schematic diagram of a loss function for classification based on a quantum Transformer according to an exemplary embodiment of the present invention;
[0053] Figure 10Schematic diagram of the accuracy of quantum Transformer-based classification according to an exemplary embodiment of the present invention. Detailed implementation manners
[0054] In order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish identical items or similar items with basically the same functions and effects. For example, the first threshold and the second threshold are only used to distinguish different thresholds, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily mean different.
[0055] It should be noted that in the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0056] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. The following at least one (item) or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b or c can represent: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b and c can be single or multiple.
[0057] The present invention constructs a quantum Transformer based on a quantum variational circuit and a quantum wavelet KAN network, combines the quantum variational circuit into the Transformer, constructs the Q, K, and V matrices, and transforms the classical attention architecture into a quantum multi-head attention architecture. At the same time, a quantum wavelet KAN network is used as the feed-forward layer, and a learnable quantum gate basis function is used to replace the linear weight. The quantum wavelet KAN can approximate the function better with fewer parameters, which not only improves the computational efficiency but also enhances the interpretability of the model.
[0058] The quantum Transformer is a quantum multi-head attention network with a quantum wavelet KAN network as the feed-forward layer; the method includes: obtaining a feature vector of an object to be classified, performing quantum encoding on the vector to generate a quantum initial state; respectively constructing a quantum query layer, a quantum key layer, and a quantum value layer by using a quantum variational circuit to form a quantum self-attention network, and then stacking multiple self-attention networks to construct a quantum multi-head attention network, and inputting the quantum initial state into the quantum multi-head attention network; performing feed-forward processing on the output result of the quantum multi-head attention network through a quantum wavelet KAN network; and performing a linear transformation on the feed-forward processing result to obtain the recognition result of the object to be classified. Applying the quantum Transformer proposed by the present invention can improve the calculation efficiency and enhance the interpretability of the network at the same time.
[0059] Next, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0060] As Figure 1 shown, a flowchart of a classification method based on a quantum Transformer is given, including the following steps:
[0061] Step S1: Obtain a feature vector of an object to be classified, and perform quantum encoding on the feature vector to generate a quantum initial state;
[0062] Step S2: Respectively construct a quantum query layer, a quantum key layer, and a quantum value layer by using a quantum variational circuit to form a quantum self-attention network;
[0063] Step S3: Construct a quantum multi-head attention network based on multiple quantum self-attention networks, and input the quantum initial state into the quantum multi-head attention network;
[0064] Step S4: Use a quantum wavelet KAN network as a feed-forward layer to perform feed-forward processing on the output result of the quantum multi-head attention network;
[0065] Step S5: Perform a linear transformation on the feed-forward processing result by using a linear layer to obtain the recognition result of the object to be classified.
[0066] As Figure 2 shown, a schematic diagram of a classification device based on a quantum Transformer is given, including:
[0067] A quantum encoding unit, configured to obtain a feature vector of an object to be classified, and perform quantum encoding on the feature vector to generate a quantum initial state;
[0068] A self-attention unit, configured to respectively construct a quantum query layer, a quantum key layer, and a quantum value layer by using a quantum variational circuit to form a quantum self-attention network;
[0069] The multi-head self-attention unit is configured to construct a quantum multi-head attention network based on multiple quantum self-attention networks and input the quantum initial state into the quantum multi-head attention network;
[0070] The feed-forward unit is configured to perform feed-forward processing on the output result of the quantum multi-head attention network by using a quantum wavelet KAN network as a feed-forward layer;
[0071] The output unit is configured to perform a linear transformation on the feed-forward processing result by using a linear layer to obtain the recognition result of the object to be classified.
[0072] As Figure 3 shown, a schematic diagram of the quantum Transformer architecture is given. The overall architecture of the quantum Transformer unfolds sequentially according to the data flow processing process. First, the feature vector of the object to be classified is converted into a quantum initial state by a quantum encoding unit. Subsequently, the quantum initial state is input into the self-attention unit, which constructs a quantum query layer, a quantum key layer, and a quantum value layer respectively through a quantum variational circuit to form a quantum self-attention network for capturing the correlation between different features in the quantum state. On this basis, the multi-head self-attention unit constructs a quantum multi-head attention network by parallelizing and stacking multiple quantum self-attention networks, forming a multi-head attention mechanism. The quantum initial state is decomposed and parallelly input into multiple quantum heads to comprehensively extract global and local feature information. Then, the output result of the quantum multi-head attention network is processed by a feed-forward unit composed of a quantum wavelet KAN network, and the characteristics of the wavelet basis function are used to deeply extract and analyze the multi-scale features in the quantum state, thereby enhancing the classification ability of the model. Finally, the processed quantum state enters the output unit, and through a linear transformation, it is mapped into classical data to generate the final recognition result of the object to be classified, thus completing the data processing flow of the entire quantum Transformer.
[0073] The quantum Transformer is constructed as a quantum multi-head attention network, and the quantum multi-head attention network is constructed based on a quantum self-attention network. The quantum self-attention network includes a quantum query layer, a quantum key layer, and a quantum value layer.
[0074] The quantum query layer is used to calculate the similarity between each piece of information of the object to be classified in the vector encoding and the information of other objects to be classified, and determine the degree of attention to each piece of information of the object to be classified. A query matrix is generated through a quantum variational circuit ;
[0075] The quantum key layer is used to describe the feature representation of each piece of information of the object to be classified in the vector encoding, and compare it with the quantum query layer to determine the degree of correlation with the quantum query layer. A key matrix is generated through a quantum variational circuit ;
[0076] The quantum value layer assigns weights to the quantum values of each object information to be classified according to the degree of correlation between the feature representation of each object information to be classified in the vector encoding and the quantum query layer, and extracts the important structural feature information corresponding to the object information to be classified in the vector encoding. A value matrix is generated through a quantum variational circuit 。
[0077] The quantum variational circuit is used to generate the quantum query layer matrix 、the quantum key layer matrix and the quantum value layer matrix 。The above-mentioned quantum query layer, quantum key layer and quantum value layer are evolved and generated through a quantum variational circuit. The quantum variational circuit is the core component of the quantum attention mechanism and the multi-layer perceptron. Given the input matrix X and the parameter matrix W, the obtained behavior is similar to matrix multiplication. The final parameter matrix W can be obtained by evolving the initial parameters
[0078] In the quantum variational circuit, each feature of the vector is embedded into the quantum bits by encoding them as rotation angles. Next, the parameter matrix W acts, and a single-parameter single-qubit rotation of one layer acts on each quantum circuit. These parameters are trained together with other parameters of the model. Then, CNOT gates are connected to entangle the quantum bit states. Therefore, the obtained behavior is similar to matrix multiplication. Finally, each quantum bit is measured, and the output constitutes the Q, K, V matrices of the output
[0079] The expression of the quantum self-attention network generated through the quantum variational circuit is as follows
[0080] ;
[0081] where the given query matrix 、key matrix and value matrix 。
[0082] Furthermore, a multi-head quantum self-attention network can be constructed based on multiple quantum self-attention networks. Multi-head attention is an extension of the attention mechanism, which allows the model to simultaneously focus on information from different representation subspaces. At different positions, instead of performing a single attention function, the query, key, and value are linearly projected h times through different linear projections, the attention functions are executed in parallel, the results are concatenated, and another learned linear projection matrix is used to project the concatenated output. The mathematical definition of multi-head attention is
[0083] 。
[0084] For the multi-head quantum self-attention network, a quantum wavelet KAN network is adopted N is used as a feed - forward network. The output result of the structure MHA of the multi - head attention mechanism is input into the quantum wavelet KAN network as a feed - forward layer to obtain the result Z; at the same time, the result Z is used as the input for re - normalization and residual connection:
[0085] ;
[0086] ;
[0087] Among them
[0088] ;
[0089] Among them
[0090] ;
[0091] Here represents the connection layer and the activation function of layer . Each element represents the activation function of the th neuron of the connection layer and the th neuron of layer .
[0092] Taking the quantum Morlet wavelet transform basis as an example, the quantum wavelet transform basis is expressed as:
[0093] ;
[0094] The elements of each row of the quantum wavelet transform basis matrix are summed, and the result vector is output. It is defined as follows:
[0095] ;
[0096] The elements of the vector are defined as:
[0097] ;
[0098] The quantum wavelet KAN network quantizes the KAN neural network. In the quantum wavelet KAN network, a basis function is designed based on quantum gates, so that the output of the quantum circuit state is , and the basis function can be constructed through the parameters of the quantum wavelet function.
[0099] The feedforward output result Z of the quantum wavelet KAN network enters the linear layer. The linear layer transforms the feedforward output result Z to obtain a deeper structural representation of the information of each object to be classified, thereby obtaining the output probability corresponding to the important structural feature information of the information of each object to be classified and obtaining the recognition result of the information of each object to be classified :
[0100] ;
[0101] Taking the cifar10 images as the objects to be classified, the simulation results show that the quantum Transformer based on the quantum variational circuit and the quantum wavelet KAN network has better classification performance for the cifar10 dataset and higher resolution for different images under the same number of parameters
[0102] As Figure 4 shown, a schematic diagram of the quantum circuit of the quantum wavelet basis function is given. The quantum circuit of the wavelet basis function includes the amplitude embedding gate passed first, the RX gate passed then, and the measurement performed finally. The wavelet basis function is a continuous wavelet basis function, and the continuous wavelet basis function can be generally composed of the combination or deformation of a complex sine wave and a Gaussian envelope. The complex sine wave of the continuous wavelet basis function is simulated by the RX gate, and the Gaussian envelope of the wavelet basis function is simulated by the amplitude embedding gate, realizing the construction of the continuous wavelet based on the quantum circuit. The quantization framework of the continuous wavelet basis can be mainstream continuous wavelets such as Morlet, Gabor, Mexican Hat, Gaussian, etc., all of which are continuous wavelet basis functions centered on the complex sine wave and the Gaussian envelope
[0103] The constructed quantum continuous wavelet basis is expressed as:
[0104] ;
[0105] Where: represents the Gaussian envelope part, is the scale parameter controlling the time domain width of the wavelet; represents the complex sine wave part, is the parameter controlling the center frequency of the wavelet; is the wavelet basis function constructed by the product of the above two parts
[0106] According to an embodiment of the present invention, the Gaussian envelope part of the continuous wavelet basis is input into the amplitude embedding gate to obtain the quantum Gaussian envelope function for amplitude encoding, and the Gaussian envelope quantum state is expressed as:
[0107] ;
[0108] Where represents the input variable; represents the standard deviation, which determines the width or spread of the Gaussian function.
[0109] The frequency variable of the continuous wavelet basis the RX quantum gate to obtain a quantum complex sine wave for encoding the rotation angle with a complex sine wave , and the quantum state of the complex sine wave is expressed as:
[0110] ;
[0111] where is the RX quantum gate, represents the angular frequency, which determines the oscillation frequency of the sine wave.
[0112] The output of the quantum circuit of the previous - stage wavelet basis function in the quantum wavelet KAN network is used as the input of the quantum circuit of the next - stage wavelet basis function. The quantum wavelet KAN network is expressed as:
[0113] ;
[0114] ;
[0115] where is regarded as a matrix including only the input vector, representing the initial quantum state; L represents the number of layers of the quantum circuit of the wavelet basis function, represents the previous - stage connection layer and the next - stage connection layer 's activation function, represents the quantum continuous wavelet basis function, expressed as: ; represents the Gaussian envelope part, represents the complex sine wave part, represents the angular frequency, which determines the oscillation frequency of the sine wave; represents the input variable; is the scale parameter that controls the time - domain width of the wavelet and is used to determine the width or spread of the Gaussian function.
[0116] Next, taking the Morlet wavelet in the continuous wavelet basis function as an example, the construction of the quantum wavelet circuit is described. The Morlet wavelet is expressed as , where the Gaussian envelope part is expressed as: , and the complex sine wave part is expressed as: , where represents the angular frequency, which determines the oscillation frequency of the sine wave; represents the input variable; Denotes the standard deviation, which determines the width or spread of the Gaussian function.
[0117] For a complex sine wave, the rotation angle can be encoded using the RX quantum gate. The matrix transformation of the RX quantum gate is represented as:
[0118] ;
[0119] The role of the Rx gate is to rotate the quantum state around the X-axis, and its operation is defined as: ;
[0120] Therefore, for the quantum state , after applying , the quantum state becomes:
[0121] .
[0122] For the modulated Gaussian function, the amplitude can be encoded using the amplitude embedding gate. The Gaussian function to be modulated is .
[0123] where t is the input variable, is the amplitude modulation parameter for modulating the shape of the Gaussian function. The input variable can be encoded into the amplitude of a single qubit. Since the qubit state must be normalized, the constructed quantum state is represented as:
[0124] , where the amplitude satisfies . , then . The initial qubit state after embedding is:
[0125] .
[0126] Furthermore, the quantum wavelet basis function quantum circuit is constructed by fitting the wavelet basis function. To derive the function obtained from the quantum circuit of the wavelet basis function.
[0127] As Figure 4 shown, first perform amplitude embedding on , then perform a rotation, and finally perform a measurement. The initial qubit state after embedding is:
[0128] ;
[0129] The gate represents a rotation of angle around the X-axis. Its matrix representation is:
[0130] ;
[0131] Apply Applied to a quantum state , an output quantum state is obtained :
[0132] ;
[0133] Calculate the components of the output quantum state :
[0134] Ground state component: ;
[0135] Excited state component: ;
[0136] In quantum mechanics, a single measurement cannot directly obtain the expectation value of an operator. The result of a single measurement is an eigenvalue of the operator. For example, for the quantum state , the measurement result can only be or . To obtain the expectation value of , a large number of samples need to be measured, and then the average value is calculated, or it is obtained by calculating the probability distribution.
[0137] Specifically: A single measurement can only obtain one of the measurement results of or .
[0138] Expectation value: Multiple measurements are required, and the probabilities and of the measurement results being and are statistically obtained, and then the expectation value is calculated:
[0139] ;
[0140] Suppose the expectation value of the measurement , which distinguishes the probabilities of obtaining and :
[0141] ;
[0142] Calculate :
[0143] ;
[0144] Similarly, calculate :
[0145] ;
[0146] Simplify the expectation value and calculate the difference:
[0147] ;
[0148] Further simplify:
[0149] ;
[0150] Using trigonometric identities :
[0151] ;
[0152] Because and , so:
[0153] ;
[0154] Obtain the expected value of :
[0155] ;
[0156] Therefore, the function measured for is: . In this way, the construction of the quantum Morlet wavelet basis function is completed.
[0157] Furthermore, a quantum wavelet KAN network is constructed using the quantum wavelet basis function quantum circuit. The multiple quantum wavelet basis function quantum circuits are stacked to form a quantum wavelet KAN network. As Figure 5 shown, a schematic diagram of the quantum wavelet KAN network is given. The quantum wavelet KAN network includes:
[0158] A quantum input layer, where classical data X is encoded into a quantum state.
[0159] A quantum KAN network layer, which constructs a transform basis function using the quantum wavelet transform basis.
[0160] The input data of the quantum wavelet KAN network first passes through the quantum input layer to map the classical data to a quantum state; then the quantum wavelet KAN network consists of multiple KAN network layers, and the output of each layer serves as the input of each node in the next layer. The data starts from the quantum input layer and is processed through each layer of the quantum KAN network layer, gradually transforming and extracting features.
[0161] The quantum Transformer feed-forward layer uses a quantum wavelet KAN network, saving the number of parameters. Based on the quantum wavelet KAN network, it circumvents the fully connected weight matrix of the traditional MLP, fundamentally reducing the parameter complexity. Compared with traditional deep learning models, the quantum wavelet KAN network uses quantum wavelet basis function nodes instead of constant nodes, and the nodes between different layers are connected through activation functions to implement a quantum wavelet KAN network with quantum wavelet basis functions as nodes. While the traditional MLP requires multiple layers of neurons to be stacked, and each neuron contains multiple weight parameters. The formula for the number of parameters of KAN is: , where is the input dimension, is the parameter of the quantum wavelet basis function, and the number of parameters grows linearly with the input dimension ( ). The formula for the number of parameters of the traditional MLP is: , where is the input dimension, is the width of the hidden layer, and the number of parameters grows quadratically with and ( ). If the input dimension is 100, the number of parameters of the traditional MLP reaches , and the number of parameters of KAN is only The number of parameters of KAN is only 20% of that of the MLP. Each node function can approximate complex non-linear relationships with a small number of parameters, saving the number of parameters.
[0162] As Figure 6 shows, a schematic diagram of a V-shaped quantum variational circuit is given. In the quantum variational circuit, the combination of quantum logic gates that repeatedly appears in the quantum variational circuit is defined as a combined U gate, which serves as the basic building block of the quantum variational circuit.
[0163] In Figure 6 the shown quantum variational circuit, n = 8 qubits are configured. For each qubit, first an H gate is applied to put the qubit in a superposition state, and then a rotation gate RX is applied for the input quantum initial state. Subsequently, a quantum logic gate module, called a combined U gate, is applied between every two adjacent qubits. The combined U gate includes multiple quantum gates, specifically including a CNOT gate configured between two adjacent qubits, an RX rotation gate and an RZ rotation gate configured on the first qubit among two adjacent qubits, and an RY rotation gate configured on the second qubit among two adjacent qubits.
[0164] The combined U gate is divided into a forward combined U gate and a reverse combined U gate. During the construction of the quantum variational circuit, the forward combined U gate is applied first. As Figure 6 shows, from to the adjacent qubits and Apply a forward combined U gate, the forward combined U gate including a CNOT gate configured between qubits and , an RX and an RZ rotation gate configured on qubit , and an RY rotation gate configured on qubit .
[0165] As Figure 7 shown, a schematic diagram of the forward combined U gate applied between adjacent qubits and is given.
[0166] Controlled-NOT gate (CNOT): , used to generate quantum entanglement between adjacent qubits and .
[0167] The RY rotation gate configured on qubit is used to perform a rotation in the Y-axis direction on , thereby adjusting its quantum state and realizing the manipulation of the amplitude of this qubit.
[0168] The RX rotation gate configured on qubit is used to perform a rotation in the X-axis direction on , adjusting the amplitude and phase of its quantum state.
[0169] The RZ rotation gate configured on qubit is used to perform a rotation in the Z-axis direction on , adjusting the phase difference of its quantum state.
[0170] The forward combined U gate provides a basis for the training and parameter optimization of quantum variational circuits by introducing entanglement (through the CNOT gate) between adjacent qubits and applying parameterizable single-qubit rotation gates (RX, RY, RZ) to the qubits.
[0171] After applying the forward combined U gate, apply the reverse combined U gate. As Figure 6 shown, from to , apply the reverse combined U gate to adjacent qubits and in sequence. The reverse combined U gate includes a CNOT gate configured between qubits and , an RX and an RZ rotation gate configured on qubit , and an RY rotation gate configured on qubit RY rotation gates on [object], where .
[0172] As Figure 8 shown, a schematic diagram of the reverse combined U gate applied between adjacent qubits and is given.
[0173] Controlled NOT gate (CNOT): , used to generate quantum entanglement between adjacent qubits and .
[0174] The RY rotation gate configured on qubit is used to perform a rotation in the Y-axis direction on so as to adjust its quantum state and achieve the manipulation of the amplitude of this qubit.
[0175] The RX rotation gate configured on qubit is used to perform a rotation in the X-axis direction on to adjust the amplitude and phase of its quantum state.
[0176] The RZ rotation gate configured on qubit is used to perform a rotation in the Z-axis direction on to adjust the phase difference of its quantum state.
[0177] The reverse combined U gate constructs a parameterized quantum circuit by introducing entanglement (through the CNOT gate) between adjacent qubits and applying parameterizable single-qubit rotation gates (RX, RY, RZ) to the qubits. By reversely connecting the quantum circuits, the expressive power of the quantum variational circuit is further improved, providing a basis for training and parameter optimization.
[0178] By comparing Figure 7 the forward combined U gate shown with Figure 8 the reverse combined U gate shown, it can be seen that the difference between the two lies in the different directions of the CNOT gate.
[0179] Referring to Figure 6 the V-shaped quantum variational circuit shown, it is also possible to first apply the reverse combined U gate and then apply the forward combined U gate to construct an inverted V-shaped quantum variational circuit, including the following steps:
[0180] During the construction of the quantum variational circuit, first apply the reverse combined U gate. From to successively apply the reverse combined U gate to adjacent qubits and The reverse combined U gate includes being configured on the qubit and the CNOT gate between, configured on the qubit the RX and RZ rotation gates, configured on the qubit the RY rotation gate on, where .
[0181] After applying the reverse combined U gate, apply the forward combined U gate. From to successively apply the forward combined U gate to adjacent qubits and , the forward combined U gate includes a CNOT gate configured between the qubits and , the RX and RZ rotation gates configured on the qubit the RY rotation gate configured on the qubit . Thus, an inverted V-shaped quantum variational circuit is constructed and formed.
[0182] After constructing and forming the quantum variational circuit, the quantum variational circuit can be used as a network for training to obtain the quantum variational circuit parameters.
[0183] Example 1: Construct the forward combined U gate.
[0184] For adjacent qubits and , define the forward combined U gate . As Figure 7 shown, the forward combined U gate includes the following parts:
[0185] Controlled-NOT gate (CNOT gate): ;
[0186] The RX and RZ rotation gates configured on the qubit : ;
[0187] The RY rotation gate configured on the qubit : .
[0188] Therefore, the overall operation of the forward combined U gate is expressed as:
[0189] ;
[0190] where represents the identity matrix, indicating no operation on the corresponding qubit.
[0191] The matrix representation of the RZ rotation gate is: 。
[0192] Therefore, for RZ 0 (θ 3 ) ⊗ I 1 its matrix representation A is:
[0193] ;
[0194] The matrix representation of the RX rotation gate is: ;
[0195] Therefore, for RX 0 (θ 1 ) ⊗ I 1 its matrix representation B is:
[0196] ;
[0197] The matrix representation of the RY rotation gate is: ;
[0198] Therefore, for I 0 ⊗ RY 1 (θ 2 ) its matrix representation C is:
[0199] ;
[0200] The controlled-NOT gate CNOT 0 , 1 its matrix representation D is:
[0201] Finally, the transformation matrix of the combined U gate can be calculated step by step:
[0202] First, calculate E = B * A: Since A is a diagonal matrix, directly multiply the diagonal elements of A to the corresponding columns of B.
[0203] Then, calculate F = C * E: For the specific structure of C, the matrix is partitioned to simplify the calculation.
[0204] Finally, calculate U = D * F: The role of the matrix representation D of the controlled-NOT gate is to swap certain rows of F. Therefore, the transformation matrix U of the combined U gate can be obtained by rearranging the rows of F. Combining the above calculations, the transformation matrix U is:
[0205] ;
[0206] Where: , represents the embedded RX rotation parameter, Represents the embedded RY rotation parameter, Represents the embedded RZ rotation parameter.
[0207] Embodiment 2: Construct a reverse U gate.
[0208] For adjacent qubits and , define the reverse combined U gate . As Figure 8 shown, the reverse combined U gate includes the following parts:
[0209] Controlled NOT gate (CNOT gate): ;
[0210] RX and RZ rotation gates configured on qubit : ;
[0211] RY rotation gate configured on qubit : .
[0212] Therefore, the overall operation of the reverse combined U gate is expressed as:
[0213] ,
[0214] wherein, represents the identity matrix, indicating no operation on the corresponding qubit.
[0215] The reverse combined U gate has a similar structure to the forward combined U gate , the difference being that the direction of the controlled NOT gate is opposite. Thus, the reverse combined U gate can be constructed with reference to Embodiment 1.
[0216] Embodiment 3: Construction of a V-shaped variational quantum circuit.
[0217] Apply the combined U gate between every two adjacent qubits, including:
[0218] From to successively apply the forward combined U gate to adjacent qubits and , the forward combined U gate includes a CNOT gate configured between qubits and , RX and RZ rotation gates configured on qubit , and a RY rotation gate configured on qubit ;
[0219] From to successively apply reverse combined U gates to adjacent qubits and The reverse combined U gate includes a CNOT gate configured between qubits and RX and RZ rotation gates configured on qubit and a RY rotation gate configured on qubit wherein, . Finally, a V-shaped variational quantum circuit is formed.
[0220] Example 4: Construction of an inverted V-shaped variational quantum circuit.
[0221] Applying the combined U gate between every two adjacent qubits includes:
[0222] From to successively apply reverse combined U gates to adjacent qubits and The reverse combined U gate includes a CNOT gate configured between qubits and RX and RZ rotation gates configured on qubit and a RY rotation gate configured on qubit ;
[0223] From to successively apply forward combined U gates to adjacent qubits and The forward combined U gate includes a CNOT gate configured between qubits and RX and RZ rotation gates configured on qubit and a RY rotation gate configured on qubit wherein, . Finally, an inverted V-shaped variational quantum circuit is formed.
[0224] Example 5: Data processing of a V-shaped variational quantum circuit.
[0225] After constructing the quantum variational circuit, the quantum variational circuit can be used as a network for training to obtain the quantum variational circuit parameters. The specific training process includes the following steps:
[0226] Step S51: Convert classical training data into quantum-encoded training data;
[0227] Convert classical data into a quantum initial state through the RX gate. The classical data can then be processed in the quantum circuit.
[0228] Step S52: Initialize the parameters of the quantum circuit;
[0229] Initialize the adjustable parameters (such as rotation angles) in the quantum circuit. Initialize the parameters by adding small random perturbations to zero.
[0230] Step S53: Input the quantum-encoded training data into the quantum variational circuit for training to obtain the parameters of the quantum variational circuit;
[0231] The training process includes the following sub-steps:
[0232] Step S5301 Forward propagation: Input the quantum-encoded training data into the quantum variational circuit and obtain the output quantum state through a series of parameterized quantum gate operations.
[0233] Step S5302 Measure and calculate the expectation value: Measure the output quantum state and calculate the relevant physical quantity or expectation value to evaluate the performance of the circuit.
[0234] Step S5303 Calculate the loss function: Calculate the loss function based on the expectation value and the target value :
[0235] ;
[0236] where, is a quantum circuit with parameters, is the Hamiltonian for measurement, is the initial quantum state.
[0237] Step S5304 Gradient calculation: Use the Parameter-Shift method to calculate the gradient of the loss function with respect to the circuit parameters.
[0238] For each parameter , the gradient can be expressed as:
[0239] ;
[0240] Step S5305 Parameter update: Update the circuit parameters according to the gradient and the selected optimization algorithm .
[0241] Update the parameters using the calculated gradient according to the selected optimization algorithm:
[0242] ;
[0243] Among them, is the learning rate.
[0244] Step S5306 iterative training: Repeat the above steps S5301 to S5305 until the loss function converges or reaches a predetermined number of training rounds.
[0245] The circuit of the quantum variational circuit adopts a symmetric structure (forward and backward CNOT chains), forming a V-shaped or inverted V-shaped topological structure, which improves the coupling of the variational quantum network. This symmetry helps to maintain a balanced gradient in different parts of the circuit; between the entangled CNOT gates, a variety of single-qubit rotation gates (R X, R Y, R Z) are applied. These rotation gates increase the expressive power of the quantum variational circuit, increase the number of parameters, and at the same time control the growth of entanglement.
[0246] Example 6: Handwritten digit classification and recognition based on quantum Transformer.
[0247] The quantum Transformer based on the quantum variational circuit and the quantum wavelet KAN network performs handwritten digit recognition, including the following steps:
[0248] Step 1: Extract CIFAR-10 picture features.
[0249] Randomly select 60,000 pictures from the CIFAR-10 picture database. The size of each picture is 32×32×3 pixels. Each picture is directly converted into matrix form for input. Finally, the matrix data of 60,000 pictures is divided into a training set and a test set, where the size of the training set is 50,000 and the size of the test set is 10,000.
[0250] Step 2: Linear projection and tiled image patches:
[0251] The image is first segmented into several image patches, and these image patches are tiled through linear projection.
[0252] Step 3: Convert each image patch into an embedding vector of a fixed dimension. Add positional encoding to each image patch to preserve its position information.
[0253] Step 4: Convert the embedding vector into a quantum state through a quantum state encoder. Using techniques such as amplitude encoding or density matrix encoding, the input vector is transformed into a superposition state of several qubits.
[0254] Step 5: The quantum variational circuit generates a quantum query layer matrix , a quantum key layer matrix and a quantum value layer matrix The above quantum query layer, quantum key layer, and quantum value layer are generated through a quantum variational circuit.
[0255] The quantum variational circuit is the core component of the quantum attention mechanism and the multi-layer perceptron. Given an input matrix X and a parameter matrix W, it behaves similarly to matrix multiplication. The expression for generating the quantum self-attention mechanism module through the quantum variational circuit is as follows:
[0256] ;
[0257] where the given query matrix is , key matrix , and value matrix .
[0258] Step 6: Construct the multi-head attention mechanism.
[0259] Multi-head attention is an extension of the attention mechanism that allows the model to simultaneously focus on information from different representation subspaces. Instead of performing a single attention function at different positions, the query, key, and value are linearly projected h times through different linear projections, the attention functions are executed in parallel, the results are concatenated, and the concatenated output is projected using another learned linear projection.
[0260] The mathematical definition of multi-head attention is:
[0261] ;
[0262] The quantum wavelet KAN network serves as a feed-forward network. The structure of the multi-head attention mechanism is input into the quantum wavelet KAN network as a feed-forward layer to obtain the result Z; at the same time, the result Z is used as the input for re-normalization and residual connection:
[0263] ;
[0264] ;
[0265] where
[0266] ;
[0267] where
[0268] ;
[0269] ;
[0270] ;
[0271] The quantum wavelet KAN network quantizes the latest network architecture, the KAN neural network. In the KAN architecture, basis functions are designed based on quantum gates, such that the output of the quantum circuit state is , and basis functions can be constructed through the parameters of the quantum wavelet function.
[0272] Step 7: Output results of the linear layer. Through transformation, the linear layer obtains a deeper structural representation of each picture's information, thereby obtaining the output probabilities corresponding to the important structural feature information of each picture's information, and obtaining the recognition results of each picture's information:
[0273]
[0274] Refer to Figure 3 , the overall architecture of the quantum Transformer unfolds sequentially according to the data flow processing process. First, the feature vector of the object to be classified is converted into a quantum initial state. Subsequently, the quantum initial state is passed through a quantum variational circuit to construct a quantum query layer, a quantum key layer, and a quantum value layer respectively, forming a quantum self-attention network for capturing the correlation between different features in the quantum state. On this basis, by parallelizing multiple quantum self-attention networks, a multi-head attention mechanism is formed. The quantum initial state is decomposed and parallelly input into multiple quantum heads to comprehensively extract global and local feature information. Then, the output results of the quantum multi-head attention network are processed by a feed-forward unit composed of a quantum wavelet KAN network, and the multi-scale features in the quantum state are deeply extracted and analyzed using the characteristics of the wavelet basis function, thereby enhancing the classification ability of the model. Finally, the processed quantum state is mapped to classical data through a linear transformation to generate the final recognition result of the object to be classified, thus completing the entire data processing process of the quantum Transformer.
[0275] Step 8: Input the features of the handwritten pictures in the training set into the neural network, and perform 50 iterations of training on the network to finally obtain a trained neural network;
[0276] Using the operation steps of the above embodiment, in the training, Pytorch and Pennylane function libraries are jointly used for training to classify the cifar10 picture set. This picture set has 10 classification labels and a total of 60,000 classified pictures, all of which are used in the simulation experiment of the present invention.
[0277] It can be seen from Figure 9 that the effective Loss value of the training decreases, and compared with the Loss value of a classical neural network with the same number of parameters, it is lower and the accuracy is higher. From Figure 10As can be seen from the curves in [the figure], the classification accuracy of the optimal convolutional neural network continues to improve. After 50 optimizations, the classification accuracy can reach 0.98. This is because the Transformer based on the quantum variational circuit and the quantum wavelet KAN network proposed in the present invention has exerted great fitting potential and shown strong optimization ability during the construction process, thus continuously improving the classification accuracy of the overall neural network.
[0278] In addition, according to an exemplary embodiment of the present invention, a computer-readable storage medium storing a computer program may also be provided. The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to execute the classification method based on the quantum Transformer according to the exemplary embodiment of the present invention. The computer-readable recording medium is any data storage device that can store data read by a computer system. Examples of the computer-readable recording medium include: read-only memory, random access memory, compact disc read-only memory, magnetic tape, floppy disk, optical data storage device, and carrier waves (such as data transmission via the Internet through wired or wireless transmission paths).
[0279] In addition, according to an exemplary embodiment of the present invention, a computing device may also be provided. The computing device includes a processor and a memory. The memory is used to store a computer program. The computer program, when executed by the processor, causes the processor to execute the computer program of the classification method based on the quantum Transformer according to the exemplary embodiment of the present invention.
[0280] Although the present invention has been described in connection with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the drawings, the disclosure content, and the like. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of cases. A single processor or other unit can implement several functions listed in the specification. Certain measures are recited in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0281] Although the present invention has been described in combination with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present invention. Accordingly, the present specification and the drawings are only exemplary descriptions of the present invention and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A classification method based on quantum Transformer, characterized in that: include: Acquire a feature vector of an object to be classified, and perform quantum coding on the feature vector to generate a quantum initial state; the object to be classified is text, speech or image; The quantum query layer, quantum key layer and quantum value layer are constructed using quantum variational circuits to form a quantum self-attention network. Building a quantum multi-head attention network based on multiple quantum self-attention networks, and inputting the quantum initial state into the quantum multi-head attention network; The quantum wavelet KAN network is used as the feedforward layer to feedforward the output of the quantum multi-head attention network; The linear layer is used to perform linear transformation on the feedforward processing result to obtain the recognition result of the object to be classified; The quantum variation circuit is formed by applying a combined U gate between every two adjacent quantum bits; The combined U gate includes a CNOT gate configured between two adjacent qubits, an RX revolving gate and an RZ revolving gate configured on the first qubit of the two adjacent qubits, and an RY revolving gate configured on the second qubit of the two adjacent qubits; The combined U gate is a forward combined U gate, and the combined U gate is applied between every two adjacent quantum bits, including: arrive For adjacent quantum bits and Apply a forward combination U gate, the forward combination U gate includes configuring a quantum bit and CNOT gates between them, configured on the quantum bit The RX and RZ revolving gates on the qubit RY revolving door on; or, The combined U gate is a reverse combined U gate, and applying the combined U gate between every two adjacent quantum bits includes: arrive For adjacent quantum bits and Applying a reverse combination U gate, the reverse combination U gate includes configuring a quantum bit and CNOT gates between them, configured on the quantum bit The RX and RZ revolving gates on the qubit RY revolving door on; Wherein, the forward combination U gate is expressed as: ; The reverse combination U gate is expressed as: ; represents the unit matrix, and n is the number of quantum bits.
2. The classification method based on quantum Transformer according to claim 1, characterized in that: The method further comprises: The quantum multi-head attention network is trained using the data to be classified in the training set. The training process uses the quantum gradient descent method and the hybrid quantum classical optimization algorithm to minimize the error function and gradually optimize the parameters of the quantum multi-head attention network. The initial parameters of the quantum multi-head attention network parameters are randomized parameters.
3. The classification method based on quantum Transformer according to claim 1, characterized in that: The quantum wavelet KAN network includes a plurality of stacked quantum wavelet basis function quantum circuits; The wavelet basis function is a continuous wavelet basis function composed of a combination or deformation of a complex sine wave and a Gaussian envelope; The quantum wavelet basis function quantum circuit comprises an amplitude embedding gate and an RX quantum gate; the amplitude embedding gate is used to perform amplitude encoding on the quantum initial state to simulate the Gaussian envelope of the wavelet basis function to obtain an embedded quantum state; the RX quantum gate is used to perform rotation angle encoding on the embedded quantum state to simulate the complex sine wave of the wavelet basis function to obtain the basis function quantum state.
4. The classification method based on quantum Transformer according to claim 1, characterized in that: The quantum self-attention network is expressed as: Among them, the quantum query layer evolves to form a query matrix , the quantum bond layer evolves to form a bond matrix , the quantum value layer evolves to form a value matrix .
5. The classification method based on quantum Transformer according to claim 4, characterized in that: The quantum multi-head attention network is expressed as: ; The quantum multi-head attention network performs h linear projections on Q, K, and V through different linear projections, and executes the quantum self-attention network in parallel. t concatenates the h outputs and uses another learned linear projection Project the output of the join.
6. A classification device based on quantum Transformer, characterized in that: include: A quantum encoding unit is configured to obtain a feature vector of an object to be classified, and to perform quantum encoding on the feature vector to generate a quantum initial state; the object to be classified is text, speech or image; The self-attention unit is configured to respectively construct a quantum query layer, a quantum key layer, and a quantum value layer using quantum variational circuits to form a quantum self-attention network; A multi-head self-attention unit is configured to construct a quantum multi-head attention network based on multiple quantum self-attention networks, and input the quantum initial state into the quantum multi-head attention network; A feedforward unit is configured to use the quantum wavelet KAN network as a feedforward layer to feedforward the output result of the quantum multi-head attention network; The output unit is configured to use a linear layer to perform a linear transformation on the feedforward processing result to obtain a recognition result of the object to be classified; The quantum variation circuit is formed by applying a combined U gate between every two adjacent quantum bits; The combined U gate includes a CNOT gate configured between two adjacent qubits, an RX revolving gate and an RZ revolving gate configured on the first qubit of the two adjacent qubits, and an RY revolving gate configured on the second qubit of the two adjacent qubits; The combined U gate is a forward combined U gate, and the combined U gate is applied between every two adjacent quantum bits, including: arrive For adjacent quantum bits and Apply a forward combination U gate, the forward combination U gate includes configuring a quantum bit and CNOT gates between them, configured on the quantum bit The RX and RZ revolving gates on the qubit RY revolving door on; or, The combined U gate is a reverse combined U gate, and applying the combined U gate between every two adjacent quantum bits includes: arrive For adjacent quantum bits and Applying a reverse combination U gate, the reverse combination U gate includes configuring a quantum bit and CNOT gates between them, configured on the quantum bit The RX and RZ revolving gates on the qubit RY revolving door on; Wherein, the forward combination U gate is expressed as: ; The reverse combination U gate is expressed as: ; represents the unit matrix, and n is the number of quantum bits.
7. A computer storage medium, characterized in that: The computer storage medium stores instructions, and when the instructions are executed, the quantum Transformer-based classification method described in any one of claims 1 to 5 is implemented.
8. A computing device, characterized in that It comprises a processor and a communication interface coupled to the processor; the processor is used to run a computer program or instruction to implement the quantum Transformer-based classification method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Quantum neural network method and system for image recognition and medium
CN112613571A
Text recognition method based on quantum transfer learning
CN119537594A