A Quantum Multi-Box Object Detection Method

By constructing a quantum classic hybrid multi-frame object detection model, and combining multi-channel quantum convolution neural module with traditional convolution layers, the shortcomings of combining QCNN with deep learning large models are solved, and the effect of significantly improving model performance and efficiency is achieved.

CN118823485BActive Publication Date: 2025-07-01NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411149426.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-07-01
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

In the prior art, the combination of quantum convolutional neural networks (QCNNs) and deep learning models has not been widely studied, and lacks systematic theoretical and practical verification, making it difficult to effectively utilize the advantages of quantum computing.

Method used

A quantum multi-frame object detection method is proposed. By constructing a quantum classic hybrid multi-frame object detection model, replacing the traditional convolution layer with a multi-channel quantum convolution neural module, and training through the improved multi-frame object detection 300 model, the model performance is evaluated using mAP indicators.

Benefits of technology

The effective fusion of quantum computing and multi-frame object detection model is realized, which significantly improves the performance and efficiency of the model. Through ablation experiments, the advantages of quantum computing in data processing are verified, and the computing time and complexity are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118823485B_ABST
    Figure CN118823485B_ABST
Patent Text Reader

Abstract

The present invention discloses a quantum multi-box object detection method, comprising the following steps: (1) obtaining data of the VOC dataset and performing preprocessing; (2) constructing a quantum-classical hybrid multi-box object detection model, namely constructing an improved multi-box object detection 300 model: modifying the fifth additional feature extraction layer of the multi-box object detection model; (3) training the improved multi-box object detection 300 model, and evaluating the performance of the quantum-classical hybrid multi-box object detection model through the mAP metric value; the present invention greatly reduces the computing time and computing complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a quantum multi-box object detection method. Background Art

[0002] In recent years, the research on the integration of quantum computing and classical machine learning models has increased significantly. The research results show that the combination of quantum computing and classical machine learning models has broad prospects. Based on this prospect, many quantum neural network (QNNs) structures have been proposed successively, such as quantum convolutional neural networks (QCNNs), quantum generative adversarial networks (QGANs), and quantum long short-term memory networks (QLSTMs). Although QCNN has shown significant potential in processing quantum computing tasks, current research mainly focuses on small-scale problems and theoretical verification. The combination of QCNN and large deep learning models has not been widely studied. Large deep learning models, such as multi-box object detection models, have powerful capabilities in processing complex visual tasks. If QCNN can be effectively combined with these large models, it is possible to significantly improve the performance and efficiency of the models.

[0003] Currently, the research on the combination of QCNN and large deep learning models is still in its infancy, lacking systematic theoretical and practical verification. Therefore, how to effectively combine QCNN with large deep learning models to give full play to the advantages of quantum computing has become an urgent problem to be solved. Summary of the Invention

[0004] Object of the Invention: The object of the present invention is to provide a quantum multi-box object detection method, which realizes the integration of quantum computing and multi-box object detection models by constructing and training a quantum-classical hybrid multi-box object detection model.

[0005] Technical Solution: A quantum multi-box object detection method according to the present invention includes the following steps:

[0006] (1) Obtain the data of the VOC dataset and perform preprocessing;

[0007] (2) Construct a quantum-classical hybrid multi-box object detection model, that is, construct an improved multi-box object detection 300 model: modify the fifth additional feature extraction layer of the multi-box object detection model;

[0008] (3) Train the improved multi-box object detection 300 model, and evaluate the performance of the quantum-classical hybrid multi-box object detection model through the mAP metric value.

[0009] Further, step (1) is specifically as follows: Obtain a total of 21,503 pictures for the training set and 10,991 pictures for the test set; Preprocess the training set data: Set the image size to 300×300 pixels, randomly adjust the color of the image, randomly flip the image horizontally, update the target bounding box, normalize the image, and match the target object with the default box; Preprocess the test set data: Set the image size to 300×300 pixels and normalize the image.

[0010] Further, step (2) is specifically as follows: Replace the convolutional layer of the multi-box object detection model with a multi-channel quantum convolutional neural module; The fifth additional feature extraction layer includes: 4 convolutional modules and 1 multi-channel quantum convolutional neural module; Among them, each convolutional block consists of a convolutional layer, a BN layer, and a Sigmoid function.

[0011] Further, the multi-channel quantum convolutional neural module consists of quantum convolutional layers; Among them, the quantum convolutional layer includes: initialization of the quantum state, quantum bit action layer, and data decoding layer.

[0012] Further, the initialization of the quantum state is specifically as follows: Use the amplitude encoding method to encode the input data x into a quantum state with n quantum bits , where, , x = ; where, is the element of the classical input data, that is, the image pixel value; i is the index, representing the index position of the data; N is the dimension of the input vector, which determines the total number of quantum states.

[0013] Further, the quantum bit action layer is specifically as follows: Use N variational quantum circuits with the same structure, that is, VQC circuits; Each circuit is responsible for one channel of the convolutional feature map, and the data on each channel is calculated by the same VQC circuit; Among them, the number of channels is determined by the number of VQC circuits; The VQC circuit uses a combination of RY, RZ gates and CNOT gates as the entanglement between quantum bits in the VQC circuit.

[0014] Further, the data decoding layer is used to measure the state of the quantum bits, convert it into the pixel value of the image, and output it.

[0015] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: It can be seen from the comparison of mAP that the model has obvious performance advantages; through ablation experiments, it can be further verified that the quantum-classical hybrid multi-box object detection model can still maintain or even improve the classification accuracy while reducing the number of parameters; the introduction of quantum computing enables the model to utilize the advantages of quantum superposition and parallel computing when processing data, greatly reducing the computing time and computational complexity. Description of the Drawings

[0016] Figure 1 It is an improved multi-box object detection 300 model architecture diagram of the present invention;

[0017] Figure 2 It is a structural diagram of the quantum convolutional layer of the present invention;

[0018] Figure 3 It is a structural diagram of the multi-channel QCNN of the present invention. Specific Embodiments

[0019] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0020] As Figure 1 shown, an embodiment of the present invention provides a quantum multi-box object detection method, including the following steps:

[0021] (1) Obtain the data of the VOC dataset and perform preprocessing; specifically as follows: Obtain a total of 21,503 training set images and 10,991 test set images; perform preprocessing on the training set data: Set the image size to 300×300 pixels, randomly adjust the color of the image, randomly horizontally flip the image, update the target box bounding box, normalize the image, and match the target object with the default Default Box; perform preprocessing on the test set data: Set the image size to 300×300 pixels and normalize the image.

[0022] (2) Construct a quantum-classical hybrid multi-box object detection model, that is, construct an improved multi-box object detection 300 model: Modify the fifth additional feature extraction layer of the multi-box object detection model; specifically as follows: Replace the convolutional layer of the multi-box object detection model with a multi-channel quantum convolutional neural module; The model architecture diagram of the backbone of multi-box object detection 300; Use ResNet50 as the backbone of the multi-box object detection 300 network, and adjust the stride of all convolutional layers in the first residual structure of the 4 additional feature extraction layers to 1, serving as the backbone of the multi-box object detection 300 network. The input image size is 300×300 pixels, and the size of the feature map obtained after the data passes through ResNet50 is 38×38×1024. A total of six feature maps are obtained using 4 additional feature extraction layers and the fifth additional feature extraction layer. Among them, the fifth additional feature extraction layer includes: 4 convolutional modules and 1 multi-channel quantum convolutional neural module; Among them, each convolutional block consists of a convolutional layer, a BN layer, and a Sigmoid function.

[0023] Among them, the multi-channel quantum convolutional neural module consists of quantum convolutional layers; Among them, as Figure 2 shown, the quantum convolutional layer includes: initialization of the quantum state, i.e., the data encoding layer, the quantum bit action layer, and the data decoding layer, i.e., the data measurement layer.

[0024] Among them, the initialization of the quantum state is specifically as follows: Use the amplitude encoding method to encode the input data x into a quantum state with n quantum bits , where, , x = . Among them, is the element of the classical input data, i.e., the image pixel value; i is the index, representing the index position of the data; N is the dimension of the input vector, which determines the total number of quantum states.

[0025] The quantum bit action layer is specifically as follows: Use 8 structurally identical variational quantum circuits, i.e., VQC circuits; Each circuit is responsible for one channel of the convolutional feature map, and the data on each channel is calculated by the same VQC circuit; Among them, the number of channels is determined by the number of VQC circuits; The VQC circuit uses a combination of RY, RZ gates and CNOT gates as the entanglement between quantum bits in the VQC circuit.

[0026] The data decoding layer is used to measure the state of the quantum bits and convert it into the pixel value of the image and output. The specific circuit diagram is as Figure 3 shown, and the specific process is as follows:

[0027] S1 Initial state:

[0028] Embed the input data through amplitude embedding Loaded into a quantum state to produce an initial state:

[0029] ;

[0030] Application of the first SU structure operation in S2: The circuit first applies SU operations to the qubit pairs (0, 1), (2, 3), (4, 5), and these operations are controlled by 15 parameters Control:

[0031] ;

[0032] Among them, the U3 gate is a universal single-qubit gate, which can be controlled by three parameters , , and as defined below:

[0033] ;

[0034] In fact, the U3 gate can be transformed into a combination of and gates, that is:

[0035] ;

[0036] Looking at the overall single structure, it contains 15 trainable parameters, which solves the problem of too few parameters in the above-mentioned QCNN circuit.

[0037] The controlled NOT gate (CNOT) is a two-qubit gate that flips the state of the target qubit (when the control qubit is 1). Its matrix form is:

[0038] ;

[0039] The RY gate is a rotation gate that rotates by an angle around the Y axis. Its matrix form is:

[0040] ;

[0041] The RZ gate is a rotation gate that rotates by an angle around the Z axis. Its matrix form is:

[0042] ;

[0043] The total initial operation can be expressed as:

[0044] ;

[0045] Application of the second SU structure operation in S3: Next, apply the same SU operation to the qubit pairs (1, 2) and (3, 4):

[0046] ;

[0047] Application of the third SU structure operation in S4: Then, apply another SU operation to the qubit pair (2, 3), this time using the second set of parameters :

[0048] ;

[0049] Final state and measurement in S5: Finally, measure the Pauli-Z operator of qubit No. 3 and output the expected value:

[0050] .

[0051] (3) Train the improved multi-box object detection 300 model, and evaluate the performance of the quantum-classical hybrid multi-box object detection model through the mAP metric value. Specifically as follows: In the quantum computing part, use the Pennylane framework to construct a multi-channel QCNN structure; set the BatchSize to 4; set the epoch to 20; set the learning rate to 0.0005; use the SGD optimizer, where the momentum value is 0.9 and the weight decay value is 0.0005; the learning rate dynamically drops to 0.3 times the original after every 5 epochs. Different from the image classification task, the multi-box object detection model needs to predict the target bounding box in addition to predicting the category, so it has two loss functions, which are used to measure respectively: (1) the localization loss, which is used to measure the difference between the target box (bounding box) predicted by the multi-box object detection model and the true target box; (2) the classification loss, which is used to measure the difference between the category predicted by the multi-box object detection model and the true category. Therefore, the L1 loss function (Smooth L1 Loss) is used in the experiment to calculate the localization loss in the quantum-classical hybrid multi-box object detection model, and the cross-entropy loss function is used to calculate the classification loss. Finally, the total loss of the model is calculated by weighted summing the localization loss and the classification loss, and normalizing the total loss according to the number of images with positive samples.

[0052] Table 1 Comparison of mAP between the quantum-classical hybrid multi-box object detection model and some classical object detection models

[0053] ;

[0054] Table 2 Comparison of classification accuracy between the quantum-classical hybrid multi-box object detection model and some classical object detection models

[0055] ;

[0056] As shown in Table 1 and Table 2, the experimental results show that the quantum-classical hybrid multi-box object detection model has the highest mAP index, and in terms of the classification accuracy of 20 categories, the quantum-classical hybrid multi-box object detection model exceeds the other models in 12 categories.

[0057] Table 3 Comparison of ablation experiments

[0058] ;

[0059] As shown in Table 3, the mAP index of the model in the ablation experiment is 0.737, which is lower than 0.742 of the quantum-classical hybrid multi-box object detection model. The accuracy of each category is shown in Table 3. From the experimental results, it can be seen that the results of the ablation experiment are inferior to those of the quantum-classical hybrid multi-box object detection model in terms of accuracy in 14 categories. This shows that the multi-channel QCNN structure can not only be applied to the multi-box object detection model, but also helps to improve the performance of the model and reduce the number of parameters.

Claims

1. A quantum multi-frame target detection method, characterized in that: The following steps are involved: (1) Obtain the data of the VOC dataset and perform preprocessing; specifically, obtain a total of 21,503 images in the training set and 10,991 images in the test set; preprocess the training set data: set the image size to 300×300 pixels, randomly adjust the color of the image, randomly flip the image horizontally, update the target box bounding box, standardize the image, and match the target object with the default box; preprocess the test set data: set the image size to 300×300 pixels and standardize the image; (2) Constructing a quantum classical hybrid multi-frame target detection model, that is, constructing an improved multi-frame target detection 300 model: modifying the fifth additional feature extraction layer of the multi-frame target detection model; specifically, replacing the convolution layer of the multi-frame target detection model with a multi-channel quantum convolution neural module; the fifth additional feature extraction layer includes: 4 convolution modules and 1 multi-channel quantum convolution neural module; each convolution block consists of a convolution layer, a BN layer, and a Sigmoid function; the multi-channel quantum convolution neural module consists of a quantum convolution layer; the quantum convolution layer includes: quantum state initialization, quantum bit action layer, and data decoding layer; The initialization of the quantum state is as follows: the input data x is encoded into a quantum state with n quantum bits using the amplitude coding method. ,in, ,x= ;in, is the element of the classical input data, i.e., the image pixel value; i is the index, representing the index position of the data; N is the dimension of the input vector, which determines the total number of quantum states; The quantum bit action layer is as follows: N variational quantum circuits with the same structure, namely VQC circuits, are used; each circuit is responsible for a channel of the convolution feature map, and the data on each channel is calculated by the same VQC circuit; the number of channels is determined by the number of VQC circuits; the VQC circuit uses a combination of RY, RZ gates and CNOT gates as the entanglement between quantum bits in the VQC circuit; the specific process is as follows: S1 initial state: ; The input data is embedded into Load into the quantum state to generate the initial state: ; S2 Application of the first SU structure operation: The circuit first applies the SU operation to the qubit pairs (0, 1), (2, 3), (4, 5). These operations consist of 15 parameters control: ; The U3 gate is a universal single-qubit gate that can be controlled by three parameters: , , and To control, it is defined as follows: ; The U3 gate can be transformed into and The combination of doors, namely: ; The controlled NOT gate (CNOT) is a two-qubit gate whose matrix form is: ; RY door is a revolving door, rotating around the Y axis at an angle ; Its matrix form is: ; The RZ door is a revolving door that rotates around the Z axis at an angle. ; Its matrix form is: ; The total initial operation is expressed as: ; S3 Application of the second SU structure operation: Apply the same SU operation to the qubit pairs (1, 2) and (3, 4): ; S4 Application of the third SU structure operation: Then apply another SU operation to the qubit pair (2, 3); using the second set of parameters : ; S5 Final state and measurement: Finally, measure the Pauli-Z operator of qubit 3 and output the expected value: ; The data decoding layer is used to convert the state of the measured quantum bit into the pixel value of the image and output it; (3) The improved multi-box target detection 300 model is trained, and the performance of the quantum-classical hybrid multi-box target detection model is evaluated by the mAP indicator value.

Citation Information

Patent Citations

  • Word2vec-QCNN model-based text representation system and method and application of Word2vec-QCNN model-based text representation system and method in construction of lexicon in power field

    CN118036605A