A Globally Correlated Block Quantum Encoding Method
Patent Information
- Application Number
- CN202610992001.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]本发明的目的是提供一种全局关联分块量子编码方法,以解决现有量子编码方法在有限量子比特和可控量子门数量的约束下,难以同时保留图像局部纹理特征与全局空间关联特征的技术问题
(1)本发明仅用8个量子比特即可编码16×16图像的256个像素值,与角度编码相比量子比特数量减少32倍,显著降低了对量子硬件资源的需求。
Smart Images

Figure CN122675974A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of quantum image classification technology, and in particular to a globally correlated block quantum coding method. Background Technology
[0002] Quantum image classification is an important branch of quantum machine learning, and its core premise lies in efficiently encoding classical image data into quantum states. The encoding method directly determines whether the quantum model can leverage the representational advantages of superposition and entanglement under limited quantum resources. Currently, mainstream quantum encoding algorithms mainly include angle encoding and amplitude encoding.
[0003] Angle encoding uses each pixel value as the angle of a quantum rotation gate, requiring N qubits to encode N pixels. This means that processing a 16×16 grayscale image would require 256 qubits, which is difficult to achieve in the current era of Noisy Intermediate-Scale Quantum (NISQ). Amplitude encoding, while requiring only [log₂N] qubits, requires an exponentially increasing number of quantum gates with N, needing approximately 900 quantum gates to encode the same image size. This results in excessively deep circuitry, severe noise accumulation, and unsatisfactory performance in practice. To balance the overhead of qubits and quantum gates, researchers proposed block encoding, which divides the image into multiple sub-blocks, encodes them independently, and then stitches them together. However, this strategy severs the spatial connections between sub-blocks, losing the overall global correlation features of the image.
[0004] Therefore, existing quantum coding methods struggle to preserve both local texture features and global spatial structure information when dealing with larger images (e.g., 16×16) under constraints of a small number of qubits (e.g., 8) and a controllable number of quantum gates (e.g., hundreds). How to construct a coding algorithm that can efficiently utilize qubits while maintaining local and global correlations is a major technical problem that urgently needs to be solved in the field of quantum image classification. Summary of the Invention
[0005] The purpose of this invention is to provide a globally correlated block quantum coding method to solve the technical problem that existing quantum coding methods, under the constraints of a limited number of qubits and controllable quantum gates, are unable to simultaneously preserve the local texture features and global spatial correlation features of an image.
[0006] To achieve the above objectives, the present invention provides a globally correlated block quantum coding method, comprising the following steps: Step S1: Obtain the classic image data to be encoded, and use the sliding window to divide the classic image data into multiple non-overlapping pixel blocks, each pixel block containing multiple pixel values; Step S2: Treat each pixel block as a local feature unit, and use quantum gate operations to encode multiple pixel values within the pixel block into a quantum bit pair, and establish local entanglement within the block within the quantum bit pair; Step S3: Apply cross-qubit quantum gate operations to different qubit pairs corresponding to different pixel blocks to establish inter-block global entanglement between different qubit pairs, thereby integrating image information from local to global in quantum space.
[0007] Preferably, the classic image data in step S1 is an image with a size of 16×16, and the pixel blocks are 2×2 non-overlapping windows.
[0008] Preferably, in step S2, the process of encoding pixel blocks into qubit pairs specifically includes: The first and second pixel values in the pixel block are encoded into the first and second qubits of the qubit pair using a rotating gate around the X-axis. A controlled NOT gate is applied between the first and second qubits in a qubit pair to establish local entanglement within the block; The third and fourth pixel values in the pixel block are encoded into the first and second qubits of the qubit pair using a rotation gate around the Y-axis.
[0009] Preferably, in step S3, the quantum gate operation across qubits is a controlled NOT gate; the establishment of global entanglement between blocks is achieved by applying a controlled NOT gate between adjacent qubit pairs.
[0010] Preferably, the quantum state obtained after encoding in steps S1 to S3 is represented by 8 qubits.
[0011] Preferably, in step S1, the partitioning method is as follows: the classical image data is divided into K sub-blocks of the same size, and each sub-block contains 2^q pixels; in step S2, each sub-block is encoded onto q qubits using amplitude coding.
[0012] Preferably, the method further includes a step of image classification based on the encoded quantum state: a. The input classical image data is encoded into multiple qubits of a quantum state using steps S1 to S3; b. Input the encoded quantum state into a parameterized quantum circuit for processing, and optimize the characteristic representation of the quantum state by adjusting the tunable parameters in the parameterized quantum circuit; c. Measure the quantum state output by the parameterized quantum circuit to obtain the measurement result; d. Input the measurement results into a classic classifier and output the image classification results.
[0013] Preferably, the parameterized quantum circuit has a depth of 2 layers; the measurement result is an expectation measurement based on the Pauli Z basis, resulting in a feature vector with a dimension of 8; the classical classifier is a linear layer used to map the 8-dimensional feature vector to the classification space.
[0014] Therefore, the present invention employs the above-mentioned globally correlated block quantum coding method, which has the following beneficial effects: (1) The present invention can encode 256 pixel values of a 16×16 image using only 8 qubits, which reduces the number of qubits by 32 times compared with angle encoding, significantly reducing the demand for quantum hardware resources.
[0015] (2) The encoding process of this invention requires only 365 quantum gates, which is about 2.7 times less than amplitude encoding (about 992 quantum gates). The circuit depth is shallower, effectively suppressing the accumulation of quantum noise, and is more suitable for execution on NISQ devices.
[0016] (3) This invention establishes local entanglement within blocks and global entanglement between blocks, so that the encoded quantum state retains the texture details inside the image blocks and captures the spatial structure information between blocks. Experiments show that it achieves an accuracy of 98% and 82.37% in the MNIST and CIFAR-10 image binary classification tasks, respectively, and an accuracy of 80.7% in the coal dataset binary classification task.
[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0018] Figure 1 This is a flowchart of a globally correlated block quantum coding method according to the present invention; Figure 2 This is a schematic diagram of the global correlation block quantum coding method of the present invention to construct the global correlation block coding; wherein, (a) is the division method of 2×2 pixel blocks, (b) is the quantum circuit diagram of encoding a single 2×2 pixel block into a quantum bit pair, and (c) is the overall quantum circuit diagram of constructing the global correlation block coding in three layers; Figure 3 This is a diagram illustrating the overall framework of the globally correlated block quantum coding method for image classification based on encoded quantum states, as described in this invention. Figure 4 This is a schematic diagram illustrating the calculation of the confusion matrix using a globally correlated block quantum coding method according to the present invention; Figure 5 This is a schematic diagram of the parameterized quantum circuit of the global correlation block quantum coding method of the present invention; Figure 6The graph shows the performance metrics of the global association block quantum coding method of the present invention on the binary classification task of MNIST dataset 3-6; where (a) is the accuracy curve, (b) is the precision curve, (c) is the recall curve, and (d) is the F1 score curve. Detailed Implementation
[0019] The following detailed description of embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely illustrates selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0020] Example like Figure 1 As shown, this invention provides a globally correlated block quantum coding method for efficiently encoding large-size classical image data onto a finite number of qubits, while simultaneously preserving the local texture features and global spatial correlation features of the image in the quantum state. This method can be further applied to quantum image classification tasks, and includes the following steps: Step S1: Obtain the classic image data to be encoded, and use a sliding window to divide the classic image data into multiple non-overlapping pixel blocks, each pixel block containing multiple pixel values.
[0021] Classic image data consists of a 16×16 image with non-overlapping 2×2 pixel blocks, such as... Figure 2 As shown in (a), a 2×2 sliding window with a step size of 2 is used to scan the image from left to right and from top to bottom, dividing the entire image into 64 non-overlapping 2×2 pixel blocks, each containing 4 pixel values. Let the input image be... Its pixel value Normalize to the interval [0,1]. Scan the pixels from left to right and top to bottom, numbering them 0 to 63. Record the four pixel values within each block as follows: .
[0022] Step S2: Treat each pixel block as a local feature unit, and use quantum gate operations to encode multiple pixel values within the pixel block into a quantum bit pair, and establish local entanglement within the block within the quantum bit pair.
[0023] This embodiment uses 8 qubits (denoted as 8 qubits). to It is divided into 4 pairs of qubits: Each pair of qubits is responsible for encoding multiple consecutive pixel blocks; specifically, every four pixel blocks form a round, which is then encoded into four pairs of qubits.
[0024] like Figure 2 As shown in (b), the specific process of encoding a 2×2 pixel block into a pair of qubits includes: Use a rotation-X gate (RX gate) to select the first pixel value in the pixel block. Second pixel value The values are encoded into the first and second qubits respectively. The rotation angle of the RX gate is proportional to the pixel value, and is typically set to a value that is... (After normalization).
[0025] A controlled-NOT gate (CNOT gate) is applied between two qubits, with the first qubit as the control bit and the second qubit as the target bit, to establish local entanglement within the block. The truth table of the controlled-NOT gate is: |00>→|00>, |01>→|01>, |10>→|11>, |11>→|10>.
[0026] Use a rotation-Y gate (RY gate) to select the third pixel value in the pixel block. and the fourth pixel value The encoding is applied to the first and second qubits respectively. The rotation angle of the RY gate is also taken as... .
[0027] After the above operations are completed, the four pixel values in a 2×2 pixel block are encoded into a pair of qubits, and the two qubits are entangled due to the action of the controlled NOT gate, so that the correlation information of the pixels in the block can be preserved.
[0028] Step S3: Apply cross-qubit quantum gate operations to different qubit pairs corresponding to different pixel blocks to establish inter-block global entanglement between different qubit pairs, thereby encoding the feature information of the entire image data into the quantum state.
[0029] After completing the local encoding of all pixel blocks, this embodiment further establishes correlations between different pairs of qubits. For example... Figure 1 As shown in (c), a controlled NOT gate is applied between adjacent pairs of qubits. For example, in between, between, A controlled NOT gate is applied between adjacent quantum pairs to allow information to be transferred between them. This is achieved through three layers of such operations (e.g., Figure 1 (c) The three-layer structure can gradually expand the local correlation to the global, so that the features between pixels that are far apart can also influence each other, and realize the integration of image information from local to global in quantum space.
[0030] Finally, the quantum state obtained after encoding in steps S1 to S3 is represented by 8 qubits. This quantum state contains all 256 pixel values of the original 16×16 image, and preserves local texture through intra-block entanglement and captures global spatial structure through inter-block entanglement.
[0031] In another embodiment, step S1 can employ a more general block-based strategy: the image is divided into K sub-blocks of equal size, each containing 2^q pixels; in step S2, each sub-block is encoded onto q qubits using amplitude coding. The basic formula for amplitude coding is: ; in ,and This generalized amplitude encoding block method can also achieve efficient encoding, but the number of quantum gates required will increase accordingly.
[0032] like Figure 3 As shown, it also includes a step of image classification based on the encoded quantum state, and the overall framework is a hybrid quantum-classical model: a. The input classical image data is encoded into 8 qubits of the quantum state using steps S1 to S3.
[0033] b. The encoded quantum state is input into a parameterized quantum circuit for processing. In this embodiment, the parameterized quantum circuit is a strongly entangled circuit, and its structure is as follows: Figure 4 (A strongly entangled circuit is shown.) This circuit contains alternating single-qubit rotation layers and two-qubit entangled layers, with a depth of 2 layers. The characteristic representation of the quantum state is optimized by adjusting the adjustable parameters in the circuit (usually initialized to random values).
[0034] c. Measure the quantum state output by the parameterized quantum circuit. In this embodiment, expectation value measurement based on the Pauli-Z basis is used. The expectation value of the Pauli-Z operator is measured for each qubit, resulting in 8 real numbers that form an 8-dimensional eigenvector. The measurement process can be repeated multiple times (e.g., 1024 times) to statistically analyze the expectation value.
[0035] d. Input the measured 8-dimensional feature vector into a classic classifier. This classic classifier is a linear fully connected layer with an input dimension of 8 and an output dimension equal to the number of categories in the classification task (e.g., 2 for binary classification). The output of the linear layer is converted into a probability distribution using a softmax function, and the category with the highest probability is taken as the final image classification result.
[0036] To verify the technical effectiveness of the method in this embodiment, experiments were conducted on multiple datasets.
[0037] The MNIST dataset is a dataset for recognizing handwritten digits. It contains grayscale images of handwritten digits from 250 different people. The MNIST dataset contains 60,000 training images and 10,000 test images, each image being 28×28 pixels in size.
[0038] The CIFAR-10 dataset is a small dataset containing common objects, with 10 classes of color images: airplanes, cars, birds, cats, deer, dogs, frogs, horses, boats, and trucks. Compared to MNIST, CIFAR-10 not only includes real-world objects, but also features more complex object characteristics. The dataset contains 50,000 training images and 10,000 test images, each image being 32×32 pixels. The selected classes are shown in Table 1.
[0039] Table 1. Categories used for CIFAR-10 classification
[0040] The coal dataset includes five categories: highly destructive coal, destructive coal, non-destructive coal, pulverized coal, and pulverized coal. Since the original sample consisted of only about 1000 images, approximately 2400 images were generated through data augmentation. This dataset contains 836 images of highly destructive coal, 759 images of destructive coal, 665 images of non-destructive coal, 573 images of pulverized coal, and 573 images of pulverized coal. The image size is... This dataset was collected from real-world factory scenarios such as coal mining and storage, naturally presenting challenges such as high levels of dust, dim lighting, and irregular object shapes. Compared to MNIST, which has a simple background and straightforward classification targets, and CIFAR-10, which has no specific interference, the coal dataset is significantly noisy, and coal and impurities are highly similar in color, severely testing the model's resistance to interference and robustness. Furthermore, for quantum models, this also examines the features required by the model, necessitating a highly efficient quantum encoding method that can provide numerous complex features for subsequent processing. The selection of the coal dataset in this embodiment is shown in Table 2.
[0041] Table 2. Categories used for classification in the coal dataset.
[0042] The experiments were conducted on a PC with an AMD Ryzen 9 8945HX CPU and an RTX 5060 GPU, using Python 3.11.13, Pennylane 0.43.1, and PyTorch 2.9.0. Pennylane is a mature framework capable of simulating variational quantum circuits. It integrates seamlessly with machine learning tools such as PyTorch, allowing users to train quantum models like neural networks and enabling automatic differentiation of circuit parameters.
[0043] The original image is transformed using bilinear interpolation. After dimension reduction, the model was trained for 10 epochs using the Adam optimizer with a batch size of 32. The cross-entropy loss function was used as the objective function for the classification task. The MNIST training set size was 1000, and the test set size was 200. The CIFAR-10 training set size was 1000, and the test set size was 400.
[0044] The four commonly used metrics for image classification are: accuracy, precision, recall, and F1 score. Precision, recall, and F1 score are collectively referred to as the PRF metric, which is calculated using a confusion matrix derived from the predicted and ground truth labels. Figure 5 As shown, where, This indicates the number of true positives, true negatives, false positives, and false negatives.
[0045] Accuracy measures the overall predictive performance of a model, and is defined as the proportion of correctly predicted samples out of the total sample size. ; Precision measures the accuracy of a model's predictions for the positive class; it is the proportion of correctly predicted positive samples out of all predicted positive samples. ; Recall measures a model's ability to cover the positive class; it is the proportion of correctly predicted positive samples out of all true positive samples. ; The F1 score comprehensively reflects the overall performance of the model, and is essentially the harmonic mean of precision and recall. A high F1 score is only achieved when both precision and recall are high, thus balancing these two metrics to reflect the model's true performance on imbalanced datasets. .
[0046] This embodiment explores the learning ability of a model under different learning rates with the same parameter settings. A smaller learning rate approaches the global optimum more slowly and may get stuck in local optima. However, a larger learning rate may cause excessively large parameter updates, potentially exceeding the optimal solution of the loss function and preventing convergence. The experimental results for different learning rates are shown in Table 3, where each cell contains two values representing the training set metric and the test set metric (format: training set / test set).
[0047] Table 3 Model metrics at different learning rates in the MNIST dataset
[0048] Taking all factors into account, when the learning rate is 0.05, the model ultimately demonstrates good generalization ability, maintaining an accuracy of approximately 97% on both the training and test sets. As shown in the table above, when the learning rate is set to 0.5, the model achieves an optimal accuracy of 93.5%, maintaining an accuracy of 81.3% on the training set, indicating severe underfitting. Furthermore, when the learning rate is set to 0.003 or 0.1, the model also shows signs of underfitting, possibly due to insufficient training caused by a small test set.
[0049] The random seed is essentially the initial state value of the pseudo-random number generator; therefore, different seeds will generate completely different random sequences. This affects the entire training trajectory of the model, leading to performance differences. Through experiments, learning rates of 0.01 and 0.05 were selected to consider the optimal performance achieved by the model after 10 epochs of training under different seeds, as shown in Table 4. Each cell contains two values, representing the training set metric and the test set metric (format: training set / test set).
[0050] Table 4. Model metrics for the three seeds in the MNIST dataset with learning rates of 0.05 and 0.01.
[0051] Table 4 shows the model's performance with three commonly used seeds. The difference in seed values has a relatively subtle impact on the model. It can be seen that with a learning rate of 0.05, the best performance of the model differs by about 1% among the three seeds. The model achieves its best performance with seed 3407 and a learning rate of 0.05, achieving 98% accuracy on the training set and 97.5% accuracy on the test set.
[0052] Standard deviation (SD) and standard error (SE) are two metrics commonly used to describe the magnitude of sampling error. SD reflects the volatility of data points, while SE reflects the volatility of the mean.
[0053] ; ; in, It is the first in the sample One value, It is the sample mean. That is the sample size.
[0054] Five replicate experiments were conducted with a learning rate of 0.05 and a seed of 3407. The standard deviation and standard error of the experimental data were calculated. Figure 6 , Figure 6 (a)-(d) correspond to accuracy, precision, recall, and F1 score, respectively.
[0055] The initial learning rate was 0.05, and the seed was 3407. Comparative experiments were conducted on block encoding, amplitude encoding, and dense angle encoding methods. The shared Ansatz circuit is a strongly entangled circuit, such as... Figure 3 As shown. The cosine annealing scheduler was used for training, and after 10 epochs, the best-performing experimental data was selected.
[0056] Table 5. Performance of the three encoding methods on binary classification of the MNIST dataset 3-6.
[0057] Table 6. Performance of this embodiment on CIFAR-10
[0058] Table 7. Performance of this embodiment on the coal dataset.
[0059] exist Figure 6 In this embodiment, the average accuracy reached a maximum of 98% on the training set and 97.5% on the test set. Table 5 shows the performance of amplitude coding, dense angle coding, and block coding in 3- and 6-class binary classification on the MNIST dataset. The total number of gates in block coding is reduced by nearly 2.72 times compared to amplitude coding, while the accuracy difference is about 1%. Compared to dense angle coding, the total number of gates in block coding only increases by 1.42 times, but the accuracy increases by 11.5%.
[0060] Furthermore, multiple experiments were conducted on CIFAR-10 with 2, 3, 4, and 5 classification categories, and the results were analyzed using the PRF index, as shown in Table 6. The model maintained a constant total quantum parameter count of 360, achieving an average accuracy of 82.37% for binary classification, 63.2% for triadic classification, 49.6% for quadratic classification, and 41.5% for pentaclassification. Table 7 shows that even with a more complex coal dataset, the model maintained 80% accuracy in binary classification, with a 1% increase in accuracy for triadic classification and a 6.8% increase in accuracy for pentaclassification. However, Table 7 indicates overfitting, which may be due to the limited sample size of the coal dataset. For 256-dimensional image classification, this encoding method successfully converted classical data into quantum states and provided analyzable features for the quantum model. This sufficiently demonstrates that the globally correlated block encoding method proposed in this embodiment can encode large-scale images into a quantum model, providing a powerful encoding tool for subsequent quantum image classification research.
[0061] Therefore, this invention employs the aforementioned globally correlated block quantum coding method, which encodes the 256 pixel values of a 16×16 image into quantum states using only 8 qubits and 365 quantum gates. Compared to amplitude coding, the number of quantum gates is reduced by approximately 2.7 times, and compared to angle coding, the number of qubits is reduced exponentially, while preserving the local texture and global spatial correlation features of the image. It achieves an accuracy of 98% on the MNIST handwritten digit binary classification task and 82.37% on the CIFAR-10 frog and boat binary classification task, making it suitable for image classification on medium-scale quantum devices with noise.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A globally correlated block quantum coding method, characterized in that, Includes the following steps: Step S1: Obtain the classic image data to be encoded, and use the sliding window to divide the classic image data into multiple non-overlapping pixel blocks, each pixel block containing multiple pixel values; Step S2: Treat each pixel block as a local feature unit, and use quantum gate operations to encode multiple pixel values within the pixel block into a quantum bit pair, and establish local entanglement within the block within the quantum bit pair; Step S3: Apply cross-qubit quantum gate operations to different qubit pairs corresponding to different pixel blocks to establish inter-block global entanglement between different qubit pairs, thereby encoding the feature information of the entire classical image data into the quantum state.
2. The globally correlated block quantum coding method according to claim 1, characterized in that: In step S1, the classic image data is a 16×16 image with non-overlapping 2×2 pixel blocks.
3. The globally correlated block quantum coding method according to claim 2, characterized in that, Step S2, the process of encoding pixel blocks into qubit pairs specifically includes: The first and second pixel values in the pixel block are encoded into the first and second qubits of the qubit pair using a rotating gate around the X-axis. A controlled NOT gate is applied between the first and second qubits in a qubit pair to establish local entanglement within the block; The third and fourth pixel values in the pixel block are encoded into the first and second qubits of the qubit pair using a rotation gate around the Y-axis.
4. The globally correlated block quantum coding method according to claim 1, characterized in that: In step S3, the quantum gate operation across qubits is a controlled NOT gate; the establishment of global entanglement between blocks is achieved by applying a controlled NOT gate between adjacent qubit pairs.
5. The globally correlated block quantum coding method according to claim 1, characterized in that: The quantum state obtained after encoding in steps S1 to S3 is represented by 8 qubits.
6. The globally correlated block quantum coding method according to claim 1, characterized in that: In step S1, the partitioning method is as follows: the classical image data is divided into K sub-blocks of the same size, and each sub-block contains 2^q pixels; in step S2, each sub-block is encoded onto q qubits using amplitude coding.
7. The globally correlated block quantum coding method according to claim 1, characterized in that, It also includes the step of image classification based on encoded quantum states: a. The input classical image data is encoded into multiple qubits of a quantum state using steps S1 to S3; b. Input the encoded quantum state into a parameterized quantum circuit for processing, and optimize the characteristic representation of the quantum state by adjusting the adjustable parameters in the parameterized quantum circuit; c. Measure the quantum state output by the parameterized quantum circuit to obtain the measurement result; d. Input the measurement results into a classic classifier and output the image classification results.
8. The globally correlated block quantum coding method according to claim 7, characterized in that: The parameterized quantum circuit has a depth of 2 layers; the measurement result is the expectation value measurement based on the Pauli Z basis, which yields an 8-dimensional feature vector; the classical classifier is a linear layer used to map the 8-dimensional feature vector to the classification space.