Guided quantum visual image recognition method and device

By introducing a guiding paradigm based on the high and low frequency complementary mechanism of human vision systems in quantum machine learning, combined with classic networks and quantum reload circuits, the problem of qubit limiting of QML in high-resolution image processing is solved, and efficient image classification and accuracy improvement is achieved.

CN120147691AActive Publication Date: 2025-06-13BEIJING INST OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510151358.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-13
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

Existing quantum machine learning (QML) faces the qubit limit in high-resolution image processing, which is time-consuming and difficult to reflect quantum advantages.

Method used

The guidance paradigm based on the high and low frequency complementary mechanism of human vision systems is adopted, and the low frequency information is processed in combination with classic network processing and high frequency information processing by quantum re-upload circuit processing is achieved to achieve efficient fusion of quantum and classic parts.

Benefits of technology

It effectively alleviates the limitations of QML in high-dimensional data processing, improves the efficiency and accuracy of image classification, and gives full play to the advantages of quantum computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147691A_ABST
    Figure CN120147691A_ABST
Patent Text Reader

Abstract

The invention provides a guide type quantum visual image recognition method and device. The method comprises the following steps: designing a guide normal form based on a visual high and low frequency complementary mechanism; and constructing an enhanced quantum classical hybrid model based on the guide normal form for quantum image recognition. The enhanced quantum classical hybrid model comprises an early visual cortex module for extracting low spatial frequency (LSF) information of an input image; the orbital frontal cortex module is used for generating initial prediction based on LSF information and positioning a potential HSF region based on an attention mechanism; the ventral visual pathway module is used for identifying an HSF region based on a high-frequency complexity index of two-dimensional discrete Fourier transform; and the fusiform module is combined with the orbital frontal cortex module and the ventral visual pathway module to obtain the HSF region, and high-spatial-frequency HSF information identification is carried out by utilizing a quantum reupload circuit. According to the invention, the efficiency and accuracy of image classification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of quantum machine learning and computer vision, and particularly to a guided quantum vision image recognition method and device. Background Art

[0002] Quantum Machine Learning (QML) is a discipline that combines quantum computing and machine learning algorithms to utilize the advantages of quantum computing to improve the performance and efficiency of machine learning models. Image classification, as a key issue in classical machine learning, is a core task for verifying the effectiveness of QML. Currently, the research of QML has expanded from simple research examples to more complex real-world tasks. With the increasing enthusiasm for QML research, its application in Computer Vision (CV) tasks has been accelerated.

[0003] Figure 1 Shows the current popular QML paradigms, including the pure quantum paradigm (a), the parallel paradigm (b), and the serial paradigm (c). Early QML research mainly focused on specific quantum algorithms, gradually forming three major directions: Quantum Kernel Methods, which apply quantum kernels to classical kernel machine learning algorithms such as Support Vector Machines (SVM); Quantum Convolutional Neural Networks (QCNN), such as using amplitude embedding and unitary operations to achieve quantum convolution with fewer trainable parameters than classical CNN; Quantum Neural Networks (QNN), based on variable parameter quantum circuits, trained by gradient descent, and representative works include data re-uploading, using multiple uploads of classical data to bypass the no-cloning limit.

[0004] However, these quantum algorithms face limitations in the number of qubits when applied. As Figure 1 shown in, the pure quantum paradigm (a) is usually limited to an input size of 8×8 or smaller. To expand the applicability of QML to high-resolution images, researchers have proposed the parallel paradigm (b), but when the input size exceeds 100x100, the increase in the number of quantum convolution operations leads to a significant increase in time consumption. Another serial paradigm (c), the quantum part cannot directly access the input data, and introduces preprocessing steps such as convolution, pooling, fully connected, and even pre-trained models, making the mechanism of action of quantum advantages unclear, specifically manifested as it is difficult to reflect quantum advantages in ways other than performance metrics. Therefore, exploring new paradigms that can reasonably utilize the advantages of quantum circuits is crucial for QML and even the entire field of quantum computing.

[0005] The human visual system has the characteristic of "Forest Before Trees" (seeing the big picture before the details). Recent research on human brain cognition has found that this characteristic stems from the complementary architecture of high spatial frequency (HSF) and low spatial frequency (LSF) signals in the visual processing process. As Figure 2 shown, visual input is first initially processed in the early visual cortex (EVC) to generate low-frequency signals containing the general outline of the image; subsequently, the orbitofrontal cortex (OFC) receives the low-frequency signals, generates a preliminary prediction and feeds it back to the fusiform gyrus, reducing the subsequent analysis burden; then, the detailed content rich in high-frequency information is transmitted to the fusiform gyrus through the ventral visual pathway (VVS); finally, the fusiform gyrus integrates the output of the OFC and deeply explores high-frequency details such as edges and textures to confirm the target recognition result.

[0006] Inspired by this, the present invention designs a novel and lightweight high-low frequency complementary QML architecture to jointly learn the low-frequency and high-frequency features of images, so as to improve the recognition speed and accuracy.

[0007] In addition, quantum data re-uploading is a QML technology that loads classical data multiple times during the calculation process to improve the expression ability of quantum circuits. Compared with linear quantum models and quantum kernel methods, the number of qubits required for quantum data re-uploading has an exponential advantage. Research shows that quantum data re-uploading circuits can fit Fourier series of any frequency, demonstrating an advantage in high-frequency representation. Subsequently, multiple studies have extended quantum data re-uploading to image tasks, but the experimental datasets are relatively simple and difficult to apply to more complex tasks.

[0008] Therefore, the present invention deeply explores the high-frequency expression ability of the data re-uploading method and combines it with the low-frequency complementarity of classical networks to find suitable practical application scenarios in image classification tasks. Summary of the Invention

[0009] Aiming at the above problems, the purpose of the present invention is to provide a guided quantum vision image recognition method and device, which break through the limitations of the existing QML paradigm in high-dimensional data processing through paradigm innovation, and based on the high-low frequency complementary mechanism of the human visual system, make full use of the high-frequency expression ability of the quantum data re-uploading method and the low-frequency complementarity of classical networks to achieve the efficient integration of the quantum part and the classical part, and improve the efficiency and accuracy of image classification.

[0010] To solve the above technical problems, the present invention provides the following technical solutions:

[0011] On the one hand, a guided quantum vision image recognition method is provided, and the method includes the following steps:

[0012] S1: Design a guiding paradigm based on the visual high-low frequency complementary mechanism;

[0013] S2: Construct an enhanced quantum-classical hybrid model based on the guiding paradigm for quantum image recognition; the enhanced quantum-classical hybrid model includes: an early visual cortex module, an orbitofrontal cortex module, a ventral visual pathway module, and a fusiform gyrus module:

[0014] Among them, the early visual cortex module extracts the low spatial frequency (LSF) information of the input image; the orbitofrontal cortex module generates an initial prediction based on the LSF information and locates potential high spatial frequency (HSF) regions based on the attention mechanism; the ventral visual pathway module identifies the HSF regions based on the high-frequency complexity index of the two-dimensional discrete Fourier transform; the fusiform gyrus module combines the HSF regions obtained by the orbitofrontal cortex module and the ventral visual pathway module and uses a quantum re-upload circuit to identify the high spatial frequency (HSF) information.

[0015] Optionally, the guiding paradigm based on the visual high-low frequency complementary mechanism includes:

[0016] Using a classical convolutional neural network to process the global low spatial frequency (LSF) information and guiding the quantum re-upload circuit to focus on the high spatial frequency (HSF) regions to process the high spatial frequency (HSF) information.

[0017] Optionally, in the early visual cortex module, a classical convolutional neural network structure is adopted, including two convolutional layers and two pooling layers, and the formula is expressed as follows:

[0018] P1 = MaxPool 1 (W 1 *X + b 1 )

[0019] Z Lsr = MaxPool 2 (W 2 *P 1 + b 2 )

[0020] Among them, X is the input image, W 1 is the weight of the first convolutional layer, b 1 is the bias of the first convolutional layer, MaxPool l represents the first pooling layer, P 1 is the output after the first convolutional layer and the first pooling layer, W 2 is the weight of the second convolutional layer, b 2 is the bias of the second convolutional layer, MaxPool 2 represents the second pooling layer, Z LSFis the output of the second convolutional layer and the second pooling layer, that is, the LSF information extracted by the early visual cortex module.

[0021] Optionally, in the orbitofrontal cortex module, the initial prediction is generated as follows:

[0022] h Classical = ReLU(W OFC Z LsF + b OFC )

[0023] where W OFC is the weight function of the initial prediction network, b OFC is the bias function of the initial prediction network, ReLU is the activation function of the initial prediction network, and h Classical is the generated initial prediction result;

[0024] Locate potential HSF regions based on the attention mechanism, expressed as follows:

[0025] A HSF = (W HSF * Z LSF + b HsF )

[0026]

[0027] where W HSF is the weight of the convolutional network based on the attention mechanism, b HSF is the bias of the convolutional network based on the attention mechanism, and A HSF is the attention map output by the convolutional network based on the attention mechanism; (i,j) represents the position coordinates in the attention map A HSF ; the value of each position (i,j) is the attention score of the region where the position is located, corresponding to the degree of human interest in HSF; (i*,j*) is the coordinate corresponding to the maximum value in A HSF ; (i orig , j orig ) represents the position coordinates in the original image corresponding to the HSF region; X represents the original input image, which is a square grayscale image with side length H, and h is the side length of the square HSF region extracted by the attention mechanism of the current module; represents the high-spatial-frequency HSF region located by the orbitofrontal cortex module.

[0028] Optionally, in the ventral visual pathway module, the high-frequency complexity index based on the two-dimensional discrete Fourier transform is calculated as follows:

[0029]

[0030] where R i,jRepresents non - overlapping regions of size h*h segmented from the original input image X, where i and j index the regions in the height and width directions respectively; (u, v) represents the point position coordinates in each region R i,j in, C u,v represents the two - dimensional discrete Fourier transform of R i,j to obtain the frequency coefficients, f u,v represents the calculated frequency amplitude of this region; E High , E Low respectively represent the sum of the energies of the high - frequency parts and the sum of the energies of the low - frequency parts at all positions in R i,j , r Ri,j represents the high - frequency complexity index;

[0031] The HSF region is selected as follows:

[0032]

[0033] Among them, represents the HSF region recognized by the ventral visual pathway module.

[0034] Optionally, the processing process of the fusiform gyrus module is as follows:

[0035] Sobel operator edge enhancement:

[0036]

[0037] Among them, K x , K y respectively represent the horizontal and vertical convolution kernels of the Sobel operator, respectively represent the horizontal and vertical gradients after anthropomorphic attention localization based on the orbitofrontal cortex module, used to capture the key features of objects in vision; respectively represent the edge gradients of the image in the horizontal and vertical directions after applying the Sobel operator based on the HSF region; E Humans and E Metrics respectively represent the edge intensity maps of the processed image, used to enhance the contours and details of objects;

[0038] Feature connection and flattening:

[0039] E Combined = Concat(E Human , E Metrie )

[0040] z HSF = Flatten(E Combined )

[0041] Among them, Concat represents feature connection, and Flatten represents feature flattening; ECobined Represents the concatenated feature, Z HSF Represents the HSF information;

[0042] Quantum reload circuit processing:

[0043] h Quantum = f Quantum ((z HSF )

[0044]

[0045] where f Quantum Represents a quantum reload circuit with multiple layers of parameterized quantum gates, h Quantum Represents the output vector of the circuit; L represents the number of stacked non-linear quantum blocks in the circuit, and l represents the serial number of each specific non-linear quantum block; Represents the classical feature vector of the nth input data in the lth layer. In the quantum circuit, these feature vectors are encoded as the rotation angles of the quantum state through the R z rotation gate; θ represents the set of all trainable parameters in the quantum circuit, Represents the parameter of the kth parameterized quantum gate module U SE in the lth layer, Represents the parameter of the additional parameterized quantum gate module at the end of the lth layer, used to adjust the quantum state before output; U (l) Represents the overall quantum operation of the lth layer, including L repeated modular operations and the final entanglement operation before output; U SE Represents a parameterized strong entanglement layer with a CZ entanglement gate, R z Represents a single-qubit rotation gate around the Z axis; z (l) Represents the classical output vector of the lth layer, Z j Represents the Pauli-Z operator acting on the jth qubit, d represents the number of qubits in the quantum circuit, i.e., the system dimension, ψ (l) Represents the quantum state processed by the lth layer, f M Represents the Mth classical post-processing function, responsible for further mapping the z (l) output by the quantum circuit to the final result h Quantum .

[0046] On the other hand, a guided quantum vision image recognition device is provided for implementing the method described in any one of the above. The device includes an enhanced quantum-classical hybrid model constructed based on a guiding paradigm, and the guiding paradigm is a guiding paradigm based on the visual high-low frequency complementary mechanism;

[0047] The enhanced quantum-classical hybrid model includes: an early visual cortex module, an orbitofrontal cortex module, a ventral visual pathway module, and a fusiform gyrus module;

[0048] Among them, the early visual cortex module is used to extract low spatial frequency (LSF) information of the input image; the orbitofrontal cortex module is used to generate an initial prediction based on the LSF information and locate potential high spatial frequency (HSF) regions based on the attention mechanism; the ventral visual pathway module is used to identify the HSF regions based on the high-frequency complexity index of the two-dimensional discrete Fourier transform; the fusiform gyrus module is used to combine the HSF regions obtained by the orbitofrontal cortex module and the ventral visual pathway module, and use the quantum reload circuit to identify the high spatial frequency (HSF) information.

[0049] On the other hand, an electronic device is provided, and the electronic device includes:

[0050] A processor;

[0051] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are loaded and executed by the processor, the steps of the above-mentioned guided quantum vision image recognition method are implemented.

[0052] On the other hand, a computer-readable storage medium is provided, and program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the steps of the above-mentioned guided quantum vision image recognition method.

[0053] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0054] (1) The present invention proposes a quantum-classical hybrid paradigm inspired by human cognition, namely the guided paradigm. This guided paradigm processes the complete low-frequency information of the image through a classical network, while guiding the quantum reload circuit to focus on the high-frequency complex regions in the image, thus effectively alleviating the limitations of the existing QML paradigm in high-dimensional data processing. The unique design of the guided paradigm realizes the efficient cooperation between the classical and quantum parts, ensuring that the advantages of quantum computing can be fully exerted.

[0055] (2) The present invention designs an architecture with complementary high- and low-frequency dual channels. The classical channel is responsible for processing the low-frequency global structure of the image, while the quantum channel focuses on modeling the high-frequency details. This design not only reduces the processing burden of the quantum part but also makes up for the deficiency of the classical part in high-frequency expression, realizing fast and accurate image classification.

[0056] (3) The present invention proposes a high-frequency complexity index based on the two-dimensional discrete Fourier transform (2D-DFT). This index is used to identify the complex high-frequency regions in the image that are difficult to effectively express by the classical network, guiding the quantum resources to concentrate on processing these key regions. The introduction of this high-frequency complexity index enhances the ability of the quantum part in high-frequency feature representation and further improves the classification performance of the overall model.

[0057] (4) The present invention realizes the integrated design of the early visual cortex module (EVC), orbitofrontal cortex module (OFC), ventral visual pathway module (VVS), and fusiform gyrus module (Fusiform). These modules are respectively responsible for low-frequency information extraction, initial prediction generation and high-frequency region localization, identifying high-frequency regions based on the high-frequency complexity index of the image, and detailed identification of high-frequency information through the quantum re-upload circuit. The integration method and collaborative working mechanism in the present invention further enhance the performance and interpretability of the overall system.

[0058] (5) The present invention realizes the efficient integration of the quantum and classical parts through paradigm innovation, high-low frequency dual-channel complementary design, and innovation of the high-frequency complexity index, improving the efficiency and accuracy of image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0060] Figure 1 is a schematic diagram of the currently popular QML paradigm;

[0061] Figure 2 is a schematic diagram of the high-low frequency complementary mechanism of human vision;

[0062] Figure 3 is a flowchart of a guided quantum vision image recognition method provided by an embodiment of the present invention;

[0063] Figure 4 is a schematic diagram of the guided paradigm provided by an embodiment of the present invention;

[0064] Figure 5 is a schematic diagram of the structure of the EQC model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present invention with reference to the drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0066] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner.

[0067] Embodiments of the present invention provide a guided quantum visual image recognition method, which can be implemented by an electronic device, and the electronic device can be a terminal or a server. As Figure 3 shown, the processing flow of this method may include the following steps:

[0068] S1: Design a guidance paradigm based on the visual high-low frequency complementary mechanism.

[0069] Humans tend to first perceive the overall contour (forest) of a scene and then focus on local details (trees) during the visual processing process, that is, the "Forest Before Trees" effect. This mechanism can not only effectively reduce the cognitive load, but also ensure the perception accuracy while maintaining the processing efficiency.

[0070] Inspired by this, the present invention proposes an innovative guidance paradigm, as Figure 4 shown. The classical network processes the complete forest and guides the quantum circuit to the complex trees. This design has two advantages: the classical network undertakes the feature extraction of the complete input, breaking through the input dimension limit; the quantum circuit directly processes the original input, ensuring that the unique advantages of quantum computing are fully exerted.

[0071] Specifically, in the present invention, the guidance paradigm based on the visual high-low frequency complementary mechanism includes:

[0072] Using a classical convolutional neural network to process global low spatial frequency (LSF) information and guiding the quantum re-upload circuit to focus on the high spatial frequency (HSF) region to process the high spatial frequency (HSF) information.

[0073] The visual high-low frequency complementary mechanism has the following advantages:

[0074] 1) Quick preliminary judgment: Make a preliminary judgment by quickly transmitting LSF information, saving time for subsequent processing. 2) Reduce cognitive load: Concentrate on processing relevant information based on the preliminary prediction, reducing the cognitive load. 3) Improve recognition accuracy and speed: Integrate HSF information and LSF information to improve the accuracy and speed of object recognition.

[0075] S2: Build an enhanced quantum-classical hybrid (EQC) model based on the guidance paradigm for quantum image recognition. As Figure 5As shown, the enhanced quantum-classical hybrid model (EQC model) includes: an early visual cortex module (EVC), an orbitofrontal cortex module (OFC), a ventral visual pathway module (VVS), and a fusiform gyrus module (Fusiform).

[0076] Among them, the early visual cortex module extracts the low spatial frequency (LSF) information of the input image; the orbitofrontal cortex module generates an initial prediction based on the LSF information and locates potential high spatial frequency (HSF) regions based on the attention mechanism; the ventral visual pathway module identifies the HSF regions based on the high-frequency complexity index of the two-dimensional discrete Fourier transform; the fusiform gyrus module combines the HSF regions obtained by the orbitofrontal cortex module and the ventral visual pathway module and uses a quantum reload circuit to identify the high spatial frequency (HSF) information.

[0077] The EQC model enhances the interpretability of the model through a module design that aligns with each stage of human vision, and obtains inspiration from biological processes to more naturally solve image classification tasks. The EQC model fully utilizes the advantages of both paradigms by thoroughly integrating classical and quantum channels to capture rich information from images.

[0078] Specifically, the classical channel is responsible for low-frequency processing of images and discovery of high-frequency regions, while the quantum channel focuses on modeling local high frequencies. This complementary design of high- and low-frequency dual channels reduces the processing burden on the quantum part while making up for the deficiency of the classical part in high-frequency expression, enabling fast and accurate image classification and alleviating the limitation of QML in input dimensions.

[0079] Furthermore, in the early visual cortex module (EVC), a classical convolutional neural network structure is adopted, including two convolutional layers and two pooling layers, aiming to quickly process the input image and extract coarse-grained low-frequency information.

[0080] The formula is expressed as follows:

[0081] P1 = MaxPool l (W 1 *X + b 1 )

[0082] Z LsF = MaxPool 2 (W 2 *P 1 + b 2 )

[0083] Among them, X is the input image, W 1 is the weight of the first convolutional layer, b 1 is the bias of the first convolutional layer, MaxPool 1 represents the first pooling layer, P 1is the output after the first convolutional layer and the first pooling layer, W 2 is the weight of the second convolutional layer, b 2 is the bias of the second convolutional layer, MaxPool 2 represents the second pooling layer, Z LSF is the output after the second convolutional layer and the second pooling layer, that is, the LSF information extracted by the early visual cortex module.

[0084] Furthermore, in the orbitofrontal cortex module (OFC), the initial prediction is generated as follows:

[0085] h Classical = ReLU(W OFC Z LSF + b OFC )

[0086] where, W OFC is the weight function of the initial prediction network, b OFC is the bias function of the initial prediction network, ReLU is the activation function of the initial prediction network, h Classical is the generated initial prediction result;

[0087] Based on the attention mechanism, the potential HSF region is located, which is expressed as follows:

[0088] A HsF = (W HsF * Z LsF + b HsF )

[0089]

[0090] where, W HSF is the weight of the convolutional network based on the attention mechanism, b HSF is the bias of the convolutional network based on the attention mechanism, A HSF is the attention map output by the convolutional network based on the attention mechanism; (i, j) represents the position coordinates in the attention map A HSF ; the value of each position (i, j) is the attention score of the region where the position is located, corresponding to the degree of human interest in HSF; (i*, j*) is the coordinate corresponding to the maximum value in A HSF ; (i orig , j orig ) represents the position coordinates in the original image corresponding to the HSF region; X represents the original input image, which is a square grayscale image with side length H, and h is the side length of the square HSF region extracted by the attention mechanism of the current module; represents the high spatial frequency HSF region located by the orbitofrontal cortex module.

[0091] Further, in the ventral visual pathway module (VVS), the high-frequency complexity index based on the two-dimensional discrete Fourier transform (2D-DFT) is calculated as follows:

[0092]

[0093] where R i,j represents a non-overlapping region of size h*h segmented from the original input image X, where i and j index the regions in the height and width directions respectively; (u, v) represents the point position coordinates in each region R i,j ; C u,v represents the frequency coefficient obtained by performing a two-dimensional discrete Fourier transform on R i,j ; f u,v represents the calculated frequency amplitude of this region; E High , E Low respectively represent the sum of the energies of the high-frequency part and the sum of the energies of the low-frequency part at all positions of R i,j ; r Ri,j represents the high-frequency complexity index;

[0094] The HSF region is selected as follows:

[0095]

[0096] where represents the HSF region recognized by the ventral visual pathway module.

[0097] Further, the processing process of the fusiform gyrus module (Fusiform) is as follows:

[0098] Sobel operator edge enhancement:

[0099]

[0100] where K x , K y respectively represent the horizontal and vertical convolution kernels of the Sobel operator, respectively represent the horizontal and vertical gradients after anthropomorphic attention localization based on the orbitofrontal cortex module, used to capture the key features of objects in vision; respectively represent the edge gradients of the image in the horizontal and vertical directions after applying the Sobel operator based on the HSF region; E Humans and E Metrics respectively represent the edge intensity maps of the processed image, used to enhance the contours and details of objects;

[0101] Feature connection and flattening:

[0102] E Combined = Concat(EHuman , E Metric )

[0103] z HSF = Flatten(E Combined )

[0104] where Concat represents feature concatenation, and Flatten represents feature flattening; E Cobined represents the concatenated features, and Z HSF represents the HSF information.

[0105] Quantum re-upload circuit processing:

[0106] h Quantum = f Quantum ((z HSF )

[0107]

[0108] where f Quantum represents a quantum re-upload circuit with multiple layers of parameterized quantum gates, and h Quantum represents the output vector of the circuit; L represents the number of stacked non-linear quantum blocks in the circuit, and l represents the serial number of each specific non-linear quantum block; represents the classical feature vector of the nth input data in the l-th layer. In the quantum circuit, these feature vectors are encoded as the rotation angles of the quantum states through the R z rotation gate; θ represents the set of all trainable parameters in the quantum circuit, represents the parameter of the k-th parameterized quantum gate module U SE in the l-th layer, represents the parameter of the additional parameterized quantum gate module at the end of the l-th layer, which is used to adjust the quantum state before output; U (l) represents the overall quantum operation of the l-th layer, including L repeated modular operations and the final entanglement operation before output; U SE represents a parameterized strong entanglement layer with CZ entanglement gates, and R z represents a single-qubit rotation gate around the Z axis; z (l) represents the classical output vector of the l-th layer, and Z j represents the Pauli-Z operator acting on the j-th qubit, d represents the number of qubits in the quantum circuit, that is, the system dimension, and ψ (l) represents the quantum state processed by the l-th layer, and f M represents the M-th classical post-processing function, which is responsible for further mapping the z (l) output by the quantum circuit to the final result h Quantum .

[0109] Through the collaborative work of the above-mentioned modules, the EQC model can effectively distinguish and integrate rough and detailed information, achieving efficient and accurate image classification.

[0110] The guided quantum vision image recognition method proposed in the present invention brings beneficial effects in many aspects through innovative technical solutions, including but not limited to the following points:

[0111] 1. Break through the limitations of the existing QML paradigm;

[0112] Optimize the combination paradigm of quantum and classical algorithms: In the current noisy intermediate-scale quantum (NISQ) era, the guided paradigm proposed in the present invention effectively optimizes the combination of quantum algorithms and classical algorithms. By processing global low-frequency information through a classical network and guiding the quantum circuit to focus on high-frequency complex regions, the application scope of quantum algorithms in high-dimensional image data processing is significantly expanded. This method not only alleviates the problem of the number of quantum feature extraction operations expanding with the input dimension, but also improves the utilization efficiency of quantum computing resources, enabling the QML model to process larger-scale and higher-resolution image data.

[0113] 2. Improve the applicability of quantum algorithms in image classification tasks;

[0114] Optimization of classification tasks combined with high-frequency Fourier analysis: The present invention accurately identifies high-frequency regions suitable for quantum processing by introducing a high-frequency complexity index based on two-dimensional discrete Fourier transform (2D-DFT). This method of combining high-frequency Fourier analysis enables quantum algorithms to focus on capturing fine texture and edge information in images, thus showing significant performance improvement in various complex image classification tasks (such as medical image classification and texture classification). Experimental results show that the EQC architecture based on the present invention outperforms existing classical and quantum baseline models on multiple high-resolution and high-frequency complexity datasets, verifying its superiority in practical applications.

[0115] 3. Innovation in the QML paradigm;

[0116] Proposing the guided paradigm: The present invention first proposes a quantum-classical hybrid paradigm inspired by human cognition - the guided paradigm. This paradigm guides the quantum circuit to the most suitable high-frequency complex regions for processing through a classical network, effectively alleviating the dimensionality limitation problem of QML when processing high-dimensional data. This innovative paradigm design realizes the efficient cooperation between the classical part and the quantum part, ensuring the full play of the advantages of quantum computing, thereby improving the image recognition performance of the overall model.

[0117] 4. Complementary architecture design;

[0118] High - frequency and low - frequency dual - channel complementary design: Drawing on the high - frequency and low - frequency complementary mechanism in the human visual system, the present invention designs a dual - channel architecture. The classical channel is responsible for processing the low - frequency global structure of the image, while the quantum channel focuses on modeling high - frequency details. This design not only reduces the processing burden of the quantum part but also makes up for the deficiency of the classical part in high - frequency expression, achieving fast and accurate image classification. The high - frequency and low - frequency dual - channel complementary design becomes the core innovation point of the present invention in the quantum - classical hybrid network architecture, significantly improving the classification efficiency and accuracy of the model.

[0119] 5. High - frequency region discovery strategy;

[0120] Innovative high - frequency region discovery strategy: The present invention proposes two efficient and novel high - frequency region discovery strategies, including the HSF localization algorithm that mimics human attention and the HSF localization algorithm driven by the high - frequency complexity index based on 2D - DFT. These strategies can accurately identify complex high - frequency regions in the image that are difficult to effectively express by the classical network, guiding the quantum resources to centrally process these key regions, and further enhancing the ability of the quantum part in high - frequency feature representation. Through these high - frequency region discovery strategies, the present invention expands the application scope of the quantum - classical hybrid algorithm, enabling it to adapt to more diverse and complex image classification tasks.

[0121] 6. Enhancing the interpretability and robustness of the model;

[0122] Modular design and biological inspiration: The EQC model, through a modular design aligned with each stage of human vision, not only enhances the interpretability of the model but also improves its robustness in different noise environments. The early visual cortex module (EVC) is responsible for low - frequency information extraction, the orbitofrontal cortex module (OFC) generates initial predictions and locates high - frequency regions, the ventral visual pathway module (VVS) identifies high - frequency regions based on the Fourier high - frequency complexity index, and the fusiform gyrus module (Fusiform) conducts detailed identification of high - frequency information through the quantum re - upload circuit. The collaborative working mechanism of each module enables the EQC model to maintain efficient training and inference performance when processing high - resolution and high - frequency complexity images, avoiding training difficulties such as gradient disappearance.

[0123] Correspondingly, an embodiment of the present invention further provides a guided quantum vision image recognition device, and the device includes an enhanced quantum - classical hybrid model (EQC model) constructed based on a guiding paradigm, and the guiding paradigm is a guiding paradigm based on the visual high - frequency and low - frequency complementary mechanism.

[0124] The enhanced quantum - classical hybrid model includes: an early visual cortex module, an orbitofrontal cortex module, a ventral visual pathway module, and a fusiform gyrus module.

[0125] Among them, the early visual cortex module is used to extract the low spatial frequency (LSF) information of the input image; the orbitofrontal cortex module is used to generate an initial prediction based on the LSF information and locate potential high spatial frequency (HSF) regions based on the attention mechanism; the ventral visual pathway module is used to identify the HSF regions based on the high-frequency complexity index of the two-dimensional discrete Fourier transform; the fusiform gyrus module is used to combine the HSF regions obtained by the orbitofrontal cortex module and the ventral visual pathway module, and use the quantum re-upload circuit to identify the high spatial frequency (HSF) information.

[0126] The device in this embodiment can be used to execute the technical solutions of the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0127] In summary, through paradigm innovation, high-low frequency dual-channel complementary design, and innovation of high-frequency complexity index, the present invention realizes a significant performance improvement of the quantum-classical hybrid network in image classification tasks. Its breakthrough technical solution not only expands the application scope of QML in high-dimensional data processing, but also improves the classification accuracy and training efficiency of the model.

[0128] In an exemplary embodiment, the present invention also provides an electronic device, which includes:

[0129] A processor;

[0130] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the above-guided quantum vision image recognition method are implemented.

[0131] In an exemplary embodiment, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by the processor to implement the steps of the above-guided quantum vision image recognition method. For example, the computer-readable storage medium can be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0132] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal device including the element.

[0133] The specification mentions "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc., indicating that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when combining a specific feature, structure, or characteristic with an embodiment, implementing such feature, structure, or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0134] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0135] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0136] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0137] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.

[0138] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0139] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, may exist physically alone for each unit, or two or more units may be integrated in one unit.

[0140] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0141] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the essence and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components, and circuits, etc., are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0142] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A guided quantum visual image recognition method, characterized in that: The following steps are involved: S1: Design a guidance paradigm based on the complementary mechanism of high and low frequency vision; S2: Construct an enhanced quantum-classical hybrid model based on the guided paradigm for quantum image recognition; The enhanced quantum classical hybrid model includes: an early visual cortex module, an orbitofrontal cortex module, a ventral visual pathway module and a fusiform gyrus module; Among them, the early visual cortex module extracts the low spatial frequency LSF information of the input image; the orbitofrontal cortex module generates an initial prediction based on the LSF information, and locates the potential high spatial frequency HSF area based on the attention mechanism; the ventral visual pathway module identifies the HSF area based on the high frequency complexity index of the two-dimensional discrete Fourier transform; the fusiform gyrus module combines the HSF areas obtained by the orbitofrontal cortex module and the ventral visual pathway module, and uses the quantum reloading circuit to identify the high spatial frequency HSF information.

2. The guided quantum visual image recognition method according to claim 1, characterized in that: The guidance paradigm based on the visual high-low frequency complementary mechanism includes: The classical convolutional neural network is used to process the global low spatial frequency LSF information, and the quantum re-upload circuit is guided to focus on the high spatial frequency HSF region to process the high spatial frequency HSF information.

3. The guided quantum visual image recognition method according to claim 1, characterized in that: In the early visual cortex module, a classic convolutional neural network structure is adopted, including two convolutional layers and two pooling layers, and the formula is as follows: P1=MaxPool1(W1*X+b1) <h2 style=";text-align:left;direction:ltr">Z<h2 style=";text-align:left;direction:ltr"> LSF <h2 style=";text-align:left;direction:ltr"> =MaxPool2(W2*P1+b2) Among them, X is the input image, W1 is the weight of the first convolutional layer, b1 is the bias of the first convolutional layer, MaxPool1 represents the first pooling layer, P1 is the output of the first convolutional layer and the first pooling layer, W2 is the weight of the second convolutional layer, b2 is the bias of the second convolutional layer, MaxPool2 represents the second pooling layer, and Z LSF It is the output of the second convolutional layer and the second pooling layer, that is, the LSF information extracted by the early visual cortex module.

4. The guided quantum visual image recognition method according to claim 3, characterized in that: In the orbitofrontal cortex module, the initial prediction is generated as follows: h Classical =ReLU(W OFC WITH LSF +b OFC ) Among them, W OFC is the weight function of the initial prediction network, b OFC is the bias function of the initial prediction network, ReLU is the activation function of the initial prediction network, and h Classical is the initial prediction result generated; Based on the attention mechanism, potential HSF areas are located, which are expressed as follows: A HSF =(W HSF *Z LSF +b HSF ) Among them, W HSF is the weight of the convolutional network based on the attention mechanism, b HSF is the bias of the attention-based convolutional network, A HSF It is the attention map output by the convolutional network based on the attention mechanism; (i, j) represents the attention map A HSF The value of each position (i, j) is the attention score of the area where the position is located, corresponding to the degree of human interest in HSF; (i*, j*) is A HSF The coordinates corresponding to the maximum value in (i orig , j orig ) represents the position coordinates corresponding to the HSF region in the original image; X represents the original input image, which is a square grayscale image with a side length of H, and h is the side length of the square HSF region extracted by the attention mechanism of the current module; High spatial frequency (HSF) regions indicating module localization in orbitofrontal cortex.

5. The guided quantum visual image recognition method according to claim 4, characterized in that: In the ventral visual pathway module, the high-frequency complexity index based on two-dimensional discrete Fourier transform is calculated as follows: Among them, R i,j represents the non-overlapping regions of size h*h that are segmented in the original input image X, where i and j index the regions in height and width respectively; (u, v) represents each region R i,j The point position coordinates in C u,v Indicates R i,j The frequency coefficients obtained by performing a two-dimensional discrete Fourier transform, f u,v Indicates the calculated frequency amplitude of the area; E High , E Low Respectively represent R i,j The energy of the high frequency part and the energy of the low frequency part at all positions, r Ri,j represents the high-frequency complexity index; HSF area selection is as follows: in, Represents the HSF region identified by the ventral visual pathway module.

6. The guided quantum visual image recognition method according to claim 5, characterized in that: The processing process of the fusiform gyrus module is as follows: Sobel operator edge enhancement: Among them, K x , K y Respectively represent the horizontal and vertical convolution kernels of the Sobel operator, They represent the horizontal and vertical gradients after human-like attention localization based on the orbitofrontal cortex module, which is used to capture the key features of objects in vision; They represent the edge gradients of the image in the horizontal and vertical directions after applying the Sobel operator based on the HSF region; E Humans and E Metrics They represent the edge intensity maps of the processed images, which are used to enhance the contours and details of the objects; Feature connection and flattening: AND Combined =Concat(E Human ,AND Metric ) z HSF =Flatten(E Combined ) Among them, Concat means feature connection, Flatten means feature flattening; E Cobined Represents the connected features, Z HSF Indicates HSF information; Quantum re-upload circuit processing: h Quantum =f Quantum (z HSF ) Among them, f Quantum represents a quantum reloading circuit with multiple layers of parameterized quantum gates, h Quantum represents the output vector of the circuit; L represents the number of stacked nonlinear quantum blocks in the circuit, and l represents the specific serial number of each nonlinear quantum block; Represents the classical eigenvector of the nth input data in the lth layer. In the quantum circuit, these eigenvectors are represented by R z The revolving gate is encoded as the rotation angle of the quantum state; θ represents the set of all trainable parameters in the quantum circuit, represents the kth parameterized quantum gate module U in the lth layer SE Parameters, represents the parameters of the additional parameterized quantum gate module at the end of the lth layer, which is used to adjust the quantum state before output; U (l) represents the overall quantum operation of the lth layer, including L repeated modular operations and the final entanglement operation before the output at the end; U SE represents a parameterized strongly entangled layer with CZ entanglement gates, R z represents a single-qubit rotating gate around the Z axis; z (l) represents the classic output vector of layer l, Z j represents the Pauli-Z operator acting on the jth qubit, d represents the number of qubits in the quantum circuit, i.e., the system dimension, ψ (l) represents the quantum state after the lth layer of processing, f M represents the Mth classical post-processing function, which is responsible for converting the z output by the quantum circuit (l) Further mapped to the final result h Quantum .

7. A guided quantum visual image recognition device, the device being used to implement the method according to any one of claims 1 to 6, characterized in that: The device includes an enhanced quantum classical hybrid model constructed based on a guided paradigm, wherein the guided paradigm is a guided paradigm based on a visual high-low frequency complementary mechanism; The enhanced quantum classical hybrid model includes: an early visual cortex module, an orbitofrontal cortex module, a ventral visual pathway module and a fusiform gyrus module; Among them, the early visual cortex module is used to extract low spatial frequency LSF information of the input image; the orbitofrontal cortex module is used to generate initial predictions based on the LSF information, and locate potential high spatial frequency HSF areas based on the attention mechanism; the ventral visual pathway module is used to identify the HSF area based on the high frequency complexity index of the two-dimensional discrete Fourier transform; the fusiform gyrus module is used to combine the HSF areas obtained by the orbitofrontal cortex module and the ventral visual pathway module, and use the quantum reloading circuit to identify high spatial frequency HSF information.

8. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are loaded and executed by the processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Medical CT image super-resolution algorithm based on dual-frequency-domain denoising enhancement

    CN116777749A

  • Fine-grained representation remote sensing image interpretation method and system based on Shearlet priori knowledge

    CN119229284A

  • Few-shot urban remote sensing image information extraction method based on meta learning and attention

    US20230215166A1