Image Feature Extraction Method, Device, Image Classification, Object Detection, and Image Segmentation Method Improved Based on Quantum Computing Operations
By introducing quantum computing operations into the ViT model, including quantum state feature enhancement, quantum self-attention mechanism optimization and quantum linear condition random field improvement, the limitations of the ViT model in image feature extraction are solved, higher feature extraction accuracy and effectiveness are achieved, and the performance of downstream tasks is improved.
Patent Information
- Application Number
- CN202510292966.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing ViT model has limitations in image feature extraction, including difficulty in digging deep into image spatial relationships, high information loss, insufficient feature accuracy and representativeness of complex texture and structural images, and low performance stability when processing images of different resolutions, scenes and types.
Image feature extraction methods based on quantum computing operations are adopted, including quantum state feature enhancement, quantum self-attention mechanism optimization and quantum linear condition random field improvement, and the feature extraction capability of ViT model is enhanced through these technical means.
It improves the accuracy and effectiveness of image feature extraction, enhances the model's understanding of image content, reduces information loss, and improves the performance of downstream image analysis tasks.
Smart Images

Figure CN119810463B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to an image feature extraction method, device, image classification, object detection, and image segmentation method improved based on quantum computing operations. Background Art
[0002] Image feature extraction is a key task in computer vision and plays a decisive role in the performance of various downstream tasks such as subsequent image classification, object detection, and image segmentation. Traditional image feature extraction methods are mainly based on handcrafted features such as edge detection, texture analysis, and color histograms. Although these methods can describe the characteristics of images to a certain extent, they rely on manually designed feature descriptors, have great limitations, and are difficult to handle complex image content and scene changes.
[0003] With the development of deep learning, convolutional neural networks (CNNs) have gradually become the mainstream method for image feature extraction. However, CNNs have limitations in dealing with long-range dependencies and variable-length sequences. To address these problems, the Transformer architecture has been introduced into the field of computer vision. Its core is the self-attention mechanism, which enables the model to focus on the dependencies between different positions in the input sequence and has powerful expressive and parallel computing capabilities. The Vision Transformer (ViT) model is a successful attempt to apply the Transformer architecture to image tasks. It divides the image into small patches and converts them into sequences, and processes these sequences through the self-attention mechanism to extract features.
[0004] Although the ViT model has achieved remarkable results, there are still some deficiencies. First, in terms of position information processing, it mainly relies on a simple position encoding strategy, and the capture of the local and global spatial relationships of the image is superficial, making it difficult to deeply explore the complex associations therein, resulting in a certain limitation in the depth of the model's understanding of the image spatial structure. Second, during the feature extraction process, there is a relatively large amount of information loss. Especially when faced with images with complex textures and structures, the extracted features are still insufficient in terms of accuracy and representativeness, resulting in an impact on the performance of downstream tasks. In addition, when processing images with different resolutions, scenes, and types, the performance stability of the traditional ViT model is relatively low, and its adaptability and generalization ability need to be further strengthened to meet the diverse visual task requirements.
[0005] Quantum computing technology has developed rapidly in recent years. Its unique advantages in processing complex computing tasks have brought new inspirations to deep learning algorithms in the field of computer vision. The characteristics of quantum computing, such as parallelism and superposition, contain great potential and are expected to reshape the architecture of image feature extraction models, optimize the feature processing process, and make up for the deficiencies of existing ViT models. However, quantum computing technology is still in the early stage of development. How to apply quantum computing to the field of computer vision and improve its computing performance is an important issue that needs to be solved in the interdisciplinary research of quantum computing and computer vision. Summary of the Invention
[0006] The purpose of the present invention is to propose an image feature extraction method and device based on improved quantum computing operations in view of the problems existing in the prior art, so as to overcome the limitations of the ViT model in image feature extraction, improve the accuracy and effectiveness of feature extraction, and further improve the performance of downstream image analysis tasks such as image classification, object detection, and image segmentation.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] An image feature extraction method based on improved quantum computing operations, the method includes:
[0009] Construct a basic architecture of the ViT model adapted to quantum computing operations, and improve the basic architecture of the ViT model in the following ways:
[0010] Quantum state feature enhancement: converting conventional features into quantum state representations based on quantum random characteristics;
[0011] Optimization of the quantum self-attention mechanism: incorporating quantum state calculations into the self-attention mechanism of the ViT model by using the influence of quantum conjugate calculation and quantum state superposition on information interaction;
[0012] Improvement of the quantum linear conditional random field: introducing a quantum linear conditional random field to improve the ViT model;
[0013] The above-mentioned quantum state feature enhancement, optimization of the quantum self-attention mechanism, and improvement of the quantum linear conditional random field act together on the image feature extraction process of the ViT model.
[0014] In the above-mentioned image feature extraction method based on improved quantum computing operations, constructing a basic architecture of the ViT model adapted to quantum computing operations includes:
[0015] Construct a Vision Transformer (ViT) model stacked by multiple Transformer encoder layers, each Transformer encoder layer includes a multi-head self-attention mechanism and a feed-forward neural network, and a quantum operation interface is reserved;
[0016] The calculation formula of the multi-head self-attention mechanism is as follows:
[0017]
[0018] Among them, represents the query matrix, represents the key matrix, represents the value matrix, represents the dimension of the key matrix;
[0019] Introduce the quantum parallel factor to the attention weight, and the adjusted attention weight is calculated as follows:
[0020]
[0021] Among them , which is initially determined by the parallel resources and the complexity of the image processing task, and will be dynamically adjusted according to the model training situation. If the computing device has rich parallel resources and the image processing task is complex (such as high-resolution image classification, multi-class object detection, etc.), can be set to be close to 1; otherwise, can be set lower, close to 0. During the training process, will be dynamically updated according to the change of the validation set loss value and the GPU idle rate.
[0022] In the above image feature extraction method improved based on quantum computing operations, the quantum state feature enhancement specifically includes:
[0023] Generate a phase factor based on the quantum random property:
[0024] Generate a random phase according to the principle of quantum random number generation , , represents the feature dimension of the feature vector;
[0025] Calculate the phase factor ;
[0026] Multiply the conventional feature by the phase factor to obtain the quantum state feature
[0027]
[0028] Among them represents element-wise multiplication, and the real part is taken as the initial quantum state enhanced feature.
[0029] In the above image feature extraction method improved based on quantum computing operations, the optimization of the quantization self-attention mechanism includes:
[0030] Calculating the conjugate of the query vector according to the principle of quantum state conjugation of ;
[0031] Using to perform matrix multiplication with the key vector , and then multiplying by the scaling factor to obtain the attention score matrix :
[0032]
[0033] where is the dimension of each head;
[0034] Obtaining the attention weight matrix through the Softmax function, and the formula is:
[0035]
[0036] where represents the row index, represents the column index, is the number of columns of the matrix. This process simulates the influence of quantum state superposition on information interaction, aiming to improve the accuracy and effectiveness of the self-attention mechanism, enabling the model to focus more precisely on key information when processing image features.
[0037] In the above image feature extraction method improved based on quantum computing operations, the training process of the quantization linear conditional random field includes:
[0038] Constructing a conditional probability model:
[0039]
[0040] where, represents the input image, represents the extracted features, is the normalization factor, is the feature function, used to describe the local feature relationship;
[0041] Optimizing the conditional probability model using a loss function, using a custom loss function to optimize the conditional probability model, making the extracted features more in line with the actual feature distribution of the image, and additional regularization terms, such as L1 or L2 regularization, can be added to the loss function according to the actual situation. The adjusted loss function can be expressed as:
[0042]
[0043] Among them, is the regularization parameter, representing the weight vector of the conditional random field.
[0044] During the training process, for the weights of the linear layer perform the following quantum rotation operation
[0045]
[0046] Among them, is the learnable quantum rotation angle, which is continuously adjusted through training. This operation is analogous to the rotation of a quantum state and can change the distribution of the weights, thereby introducing the unique characteristics of quantum computing and bringing a new optimization dimension to the model.
[0047] Use the rotated weights and the input features to perform matrix multiplication and add the bias vector , to obtain the quantized output: .
[0048] Through the above optimization of the loss function and quantum rotation operation, according to the local and global context information of the input image, adjust parameters such as the weights of the feature function and the quantum rotation angle, so that the model can more accurately optimize the feature extraction process, make full use of the characteristics related to quantum computing to improve the linear conditional random field, reduce the feature extraction loss in the feature space, and thus improve the performance of the model in the image feature extraction task.
[0049] An image classification method for an image feature extraction method improved based on quantum computing operations, the method comprising: constructing a ViT model improved based on quantum computing operations through the method, and adding a classification module to obtain an image classification model;
[0050] Use an image classification data set to train the image classification model so that it has the corresponding image classification ability.
[0051] Specifically, add a classification head after the "quantized linear conditional random field improvement" module. The classification head can adopt a fully connected layer structure, and its input is the quantized output features processed by the quantized linear conditional random field. Through the fully connected layer, learn and distinguish the image feature patterns of different classes, and finally output the class prediction corresponding to the image.
[0052] An object detection method for an image feature extraction method improved based on quantum computing operations, the method comprising: constructing a ViT model improved based on quantum computing operations through the method, and adding a detection module to obtain an object detection model;
[0053] Train the object detection model using an object detection dataset to enable it to have the corresponding object detection ability.
[0054] Specifically, add an object detection head after the "Improved Quantized Linear Conditional Random Field" module. This detection head can adopt a convolutional neural network structure, which includes a convolutional layer, a pooling layer, and a fully connected layer. The quantized output features are first passed through the convolutional layer to extract multi-scale semantic features using different convolutional kernels, then downsampled by the pooling layer to enhance robustness, and finally the fully connected layer learns the mapping relationship to output the object position and class information.
[0055] An image segmentation method based on an improved image feature extraction method using quantum computing operations, the method comprising: constructing a ViT model improved based on quantum computing operations through the method, and adding a segmentation module to obtain an image segmentation model;
[0056] Train the image segmentation model using an image segmentation dataset to enable it to have the corresponding image segmentation ability.
[0057] Specifically, add a segmentation head after the "Improved Quantized Linear Conditional Random Field" module. This segmentation head can adopt a combination of a convolutional layer and an upsampling layer. The convolutional layer extracts local features, and the upsampling layer restores the feature map size to be close to the original image size. Its input is the quantized output features processed by the quantized linear conditional random field, and the output is the segmentation mask of the image, which assigns class labels to the original image pixels to distinguish different target regions.
[0058] An image feature extraction device improved based on quantum computing operations, the device comprising a model architecture construction module, a self-attention optimization module, a feature enhancement module, and a quantized linear conditional random field module;
[0059] The model architecture construction module is used to construct the basic architecture of the ViT model adapted to quantum computing operations;
[0060] The self-attention optimization module is used to incorporate quantum state calculations into the self-attention mechanism of the ViT model by the influence of quantum conjugate calculation and quantum state superposition on information interaction;
[0061] The feature enhancement module is used to convert conventional features into quantum state representations based on quantum random characteristics;
[0062] The quantized linear conditional random field module is used to introduce a quantized linear conditional random field into the ViT model.
[0063] In the above-mentioned image feature extraction device improved based on quantum computing operations, the training process of the quantized linear conditional random field includes:
[0064] Construct a conditional probability model:
[0065]
[0066] Among them, represents the input image, represents the extracted features, is the normalization factor, is the feature function, which is used to describe the local feature relationship;
[0067] The conditional probability model is optimized using the loss function;
[0068] During the training process, for the weights of the linear layer perform the following quantum rotation operation
[0069]
[0070] Among them, is the learnable quantum rotation angle, which is continuously adjusted through training;
[0071] Use the rotated weights and the input features to perform matrix multiplication and add the bias vector , to obtain the quantized output: .
[0072] The advantages of the present invention are as follows:
[0073] 1. By constructing a ViT model infrastructure adapted to quantum computing, this solution provides a solid framework for the integration of quantum computing and the ViT model, giving full play to the advantages of quantum computing, fundamentally enhancing the performance potential of the ViT model, and enabling it to better handle complex image feature extraction tasks;
[0074] 2. Through the quantum state feature enhancement operation, this solution greatly enriches the expression dimension of features, enabling the model to capture subtle feature information that is difficult to detect by traditional methods, effectively enhancing the model's ability to understand image content, and providing a richer and more valuable data basis for subsequent analysis and processing;
[0075] 3. The optimization of the quantized self-attention mechanism proposed in this solution significantly improves the information focusing ability and interaction accuracy of the model when processing image features, reduces the interference of irrelevant information, improves the quality and efficiency of feature extraction, and makes the extracted features more discriminative and representative;
[0076] 4. The improvement of the quantized linear conditional random field proposed in this solution effectively utilizes operations such as quantum rotation to optimize the feature extraction process, reduces information loss, enhances the model's ability to model the feature space, improves the generalization performance of the model, and enables it to perform more stably and excellently in different datasets and task scenarios;
[0077] 5. This solution realizes an innovative integration of the fields of quantum computing and computer vision, introducing the unique advantages of quantum computing into the task of image feature extraction, which helps to promote the progress of related cross-research. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 It is a flowchart of the method for improving image feature extraction based on quantum computing operations of the present invention;
[0079] Figure 2 It is a flowchart of applying the method for improving image feature extraction based on quantum computing operations of the present invention to downstream tasks of image classification;
[0080] Figure 3 It is the training process of the method for improving image feature extraction based on quantum computing operations of the present invention on the CIFAR100 dataset;
[0081] Figure 4 It is a comparison chart of the accuracy rates of the method for improving image feature extraction based on quantum computing operations of the present invention and other feature extraction methods in the control group on the CIFAR100 dataset;
[0082] Figure 5 It is a comparison chart of the computational consumption of the method for improving image feature extraction based on quantum computing operations of the present invention and other feature extraction methods in the control group on the CIFAR100 dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0083] This solution provides a method for improving image feature extraction based on quantum computing operations, which can be used for downstream tasks such as image classification, object detection, and image segmentation. This embodiment takes the image classification task as an example to introduce this solution. As Figure 1 shown, first, an image feature extraction module is built based on the proposed method, then an image classification model is established by introducing a classification module. Finally, a comparative analysis is carried out with several traditional feature extraction methods. Among them, the image feature extraction module of this embodiment is common for image classification, object detection, and image segmentation tasks. For different downstream tasks, only the corresponding dataset and the subsequent module need to be replaced, which will not be elaborated here.
[0084] The specific steps of this embodiment are as follows:
[0085] S1. Select CIFAR100 as the standard dataset. This dataset consists of 60,000 color images, which are divided into 100 categories, and each category contains 600 images.
[0086] S2. Select some feature extraction methods for comparison with the method proposed in this solution: In addition to the traditional feature extraction methods SIFT and HOG, it also includes a classic feature extraction method VGG16 based on neural network. Add the above three methods to the control group. At the same time, the difference between the improved ViT and the unimproved ViT based on the proposed method is also investigated.
[0087] Among them, the SIFT algorithm can detect local features in the image and describe the texture information near the feature points;
[0088] HOG constructs features by calculating and statistically analyzing the gradient direction histogram of local regions of the image. Both use SVM for subsequent classification.
[0089] VGG16 is a deep convolutional neural network structure composed of multiple convolutional layers and pooling layers stacked alternately. The network depth is 16, and a fully connected layer is used for classification.
[0090] S3. Build the basic architecture of the ViT model adapted to quantum computing operations. The basic ViT model built in this embodiment is stacked by 6 Transformer encoder layers, and each encoder layer contains a multi-head self-attention mechanism and a feed-forward neural network.
[0091] In the multi-head self-attention mechanism, the dimension of the key matrix .
[0092] When calculating the attention weights, for the CIFAR100 dataset, considering that the complexity of this dataset is at a moderate level and the parallel computing resources used in this embodiment are relatively abundant, the quantum parallel factor is initially set to 0.5. After every 10 training batches, calculate the average loss value on the validation set. If the validation set loss value no longer decreases for 3 consecutive times and the current GPU idle rate is greater than 30%, then increase the quantum parallel factor by a fixed value of 0.05 to try to use more parallel resources to improve the model performance; otherwise, decrease it by 0.05 to reduce the computational burden and stabilize the model training.
[0093] S4. As Figure 2 shown, optimize the ViT model built in S3 based on the proposed method.
[0094] Quantum state feature enhancement:
[0095] For the feature vector after the preliminary processing of the input image, its feature dimension .
[0096] According to the principle of quantum random number generation, generate a random phase for each dimension , . Calculate the phase factor , and then the conventional features Multiply element - by - element with the phase factor and take the real part as the initial quantum state enhancement feature. Through this quantum state feature enhancement method, the diversity and expressiveness of features can be significantly enriched, enabling the model to capture more complex and subtle image feature information and providing reliable support for subsequent feature processing and analysis.
[0097] Optimization of the quantized self - attention mechanism:
[0098] For the query vector , calculate its conjugate according to the principle of quantum state conjugation ; for the dimension of each head, use to perform matrix multiplication with the key vector , and then multiply by the scaling factor to obtain the attention score matrix , and then obtain the attention weight matrix through the function.
[0099] Improvement of the quantized linear conditional random field:
[0100] Construct the conditional probability model . In the CIFAR100 dataset, the image size is . For the feature function , considering the characteristic that adjacent pixels in CIFAR100 images usually have a high correlation, and combining the rule that the correlation of diagonal adjacent pixels is slightly weaker and the correlation of non - adjacent pixels is even weaker, set a weight relationship that can reflect the pixel - based position: the initial weight of the feature relationship for horizontal or vertical adjacent pixels is set to 0.8, the weight of the feature relationship for diagonal adjacent pixels is set to 0.6, and the weight of the feature relationship for non - adjacent pixels is set to 0.3.
[0101] Then, use the custom loss function to optimize the conditional probability model, and the regularization parameter . During the training process, for the linear layer weight of the quantized linear conditional random field, the initial quantum rotation angle is set to 0.1 and will be continuously adjusted during the training process. Multiply the rotated weight with the input feature , and add the bias vector to obtain the quantized output: .
[0102] Introduction of the classification module:
[0103] After establishing the feature extraction method and obtaining the quantized output, connect the overall ViT structure with the classification head module. In this embodiment, since the CIFAR100 dataset has 100 categories, when setting the fully connected layer of the classification head, map the feature dimension of the quantized output to 100 dimensions, and use the above loss function to train and optimize the entire model including feature extraction and the classification head to complete the classification task.
[0104] S5. During the training process, set the learning rate to 0.001, the batch size to 64, and the number of training epochs to 50, and conduct performance evaluation on the CIFAR100 dataset. In a specific experiment, after each round of training, the performance of the model on the test set will be evaluated, and performance metrics such as accuracy and recall will be statistically analyzed. As Figure 3 can be seen from the results, after 20 rounds of training, the accuracy has increased by approximately 0.3.
[0105] S6. Conduct a comparative experiment on the proposed optimized ViT model with feature extraction methods such as VGG16, SIFT, HOG, and traditional ViT. Keep the settings in S5 unchanged and retrain to obtain the accuracy, time consumption, and feature representation distribution in the feature space of each model.
[0106] From Figure 3 the training performance graph of the optimized ViT model shown, it can be seen that as the number of training epochs increases, the training loss gradually decreases and the validation accuracy gradually increases, indicating that the model is continuously optimized.
[0107] As Figure 4 can be seen, the optimized ViT model based on the proposed method finally reached an accuracy of 59.32% after 50 epochs of training. Compared with the accuracy of the traditional ViT model of 56.45%, there is an improvement of about 2% - 3%, which is quite remarkable on the CIFAR100 standard dataset. It proves the feasibility of the ViT model improved based on quantum computing in image classification tasks.
[0108] As Figure 5 can be seen, comparing the computing time consumption of the method of the present invention with each control group on the CIFAR100 dataset, the proposed method is superior to traditional methods such as VGG16, HOG, and SIFT; and there is no excessive additional resource consumption on the basis of the traditional ViT, indicating that the method of the present invention has achieved good performance within a reasonable range of computing resource consumption.
[0109] It can be seen that through the collaborative optimization of quantum computing and linear conditional random fields, this solution effectively enhances the diversity and expressiveness of image features, optimizes the information interaction efficiency in the self - attention mechanism, reduces the information loss in the feature extraction process, improves the accuracy of feature extraction, and ultimately improves the performance of the model in image analysis tasks.
[0110] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains may make various modifications or supplements to the described specific embodiments or use similar means for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. An image feature extraction method based on quantum computing operation improvement, characterized in that: The method includes: Build a ViT model infrastructure that adapts to quantum computing operations and improve the ViT model infrastructure in the following ways: Quantum state feature enhancement: converting conventional features into quantum state representation based on quantum random properties, including generating phase factors based on quantum random properties: generating random phases based on the principle of quantum random number generation , , Represents the feature dimension of the feature vector; Calculate the phase factor ; The general features Multiplying with the phase factor gives the quantum state characteristics , in It represents element-by-element multiplication, and the real part is taken as the initial quantum state enhancement feature; Optimization of the quantized self-attention mechanism: Using quantum conjugation computing and the influence of quantum state superposition on information interaction, quantum state computing is integrated into the self-attention mechanism of the ViT model, including calculating the query vector based on the quantum state conjugation principle. Conjugation ; use With key vector Perform matrix multiplication and then add the scaling factor Multiply them together to get the attention score matrix : in is the dimension of each head; The attention weight matrix is obtained through the Softmax function , introducing quantum parallelism factor for attention weights , the adjusted attention weight The calculation is as follows: in , updated with training; Quantized linear conditional random field improvement: The quantized linear conditional random field is introduced to improve the ViT model, including the construction of a conditional probability model: in, represents the input image, represents the extracted features, is the normalization factor, is a characteristic function, used to describe the local feature relationship; Use loss function to optimize the conditional probability model; During training, the linear layer weights Perform the following quantum rotation operation in, It is a learnable quantum rotation angle that is continuously adjusted through training; Use the rotated weights With input features Perform matrix multiplication and add the bias vector , and get the quantized output: ; The quantum state feature enhancement, quantized self-attention mechanism optimization and quantized linear conditional random field improvement work together in the feature extraction process of the ViT model for the image.
2. The image feature extraction method based on quantum computing operation improvement according to claim 1 is characterized in that: The ViT model infrastructure for quantum computing operations includes: Build a ViT model consisting of multiple stacked Transformer encoder layers, each of which contains a multi-head self-attention mechanism and a feedforward neural network, and has a reserved quantum operation interface; The calculation formula of the multi-head self-attention mechanism is as follows: in, represents the query matrix, represents the key matrix, represents the value matrix, Indicates the dimensions of the key matrix.
3. The image feature extraction method based on quantum computing operation improvement according to claim 1 is characterized in that: The attention weight matrix is obtained through the Softmax function The formula is: in Represents the row index, Represents the column index, yes The number of columns in the matrix.
4. An image classification method based on an improved image feature extraction method using quantum computing operations, characterized in that: The method includes: Constructing a ViT model improved based on quantum computing operations by the method described in any one of claims 1 to 3, and adding a classification module to obtain an image classification model; The image classification model is trained using an image classification dataset so that it has corresponding image classification capabilities.
5. An object detection method based on an improved image feature extraction method using quantum computing operations, characterized in that: The method includes: Constructing a ViT model improved based on quantum computing operations by the method described in any one of claims 1 to 3, and adding a detection module to obtain an object detection model; The object detection model is trained using an object detection dataset so that it has corresponding object detection capabilities.
6. An image segmentation method based on an improved image feature extraction method using quantum computing operations, characterized in that: The method includes: Constructing a ViT model improved based on quantum computing operations by the method described in any one of claims 1 to 3, and adding a segmentation module to obtain an image segmentation model; The image segmentation model is trained using an image segmentation dataset so that it has corresponding image segmentation capabilities.
7. An image feature extraction device improved based on quantum computing operation, characterized in that: The device includes a model architecture building module, a self-attention optimization module, a feature enhancement module, and a quantized linear conditional random field module; Model architecture building module, used to build the ViT model infrastructure adapted to quantum computing operations; Self-attention optimization module, Using quantum conjugation computing and the influence of quantum state superposition on information interaction, quantum state computing is integrated into the self-attention mechanism of the ViT model, including calculating the query vector based on the quantum state conjugation principle. Conjugation ; use With key vector Perform matrix multiplication and then add the scaling factor Multiply them together to get the attention score matrix : in is the dimension of each head; The attention weight matrix is obtained through the Softmax function , introducing quantum parallelism factor for attention weights , the adjusted attention weight The calculation is as follows: in , updated with training; Feature enhancement module for converting conventional features into quantum state representations based on quantum random properties, including the generation of phase factors based on quantum random properties: Generate random phases based on the principle of quantum random number generation , , Represents the feature dimension of the feature vector; Calculate the phase factor ; The general features Multiplying with the phase factor gives the quantum state characteristics , in It represents element-by-element multiplication, and the real part is taken as the initial quantum state enhancement feature; The quantized linear conditional random field module is used to introduce quantized linear conditional random fields into the ViT model, including building a conditional probability model: in, represents the input image, represents the extracted features, is the normalization factor, is a characteristic function, used to describe the local feature relationship; Use loss function to optimize the conditional probability model; During training, the linear layer weights Perform the following quantum rotation operation in, It is a learnable quantum rotation angle that is continuously adjusted through training; Use the rotated weights With input features Perform matrix multiplication and add the bias vector , and get the quantized output: .
Citation Information
Patent Citations
Molecular generation method based on quantum Transform model
CN118380072A
Quantum computer-implemented method for solving a partial differential equation
US20240296201A1