A quantum-enhanced multi-modal data alignment method
By combining quantum-enhanced multimodal alignment with classical feature representation, the problem of capturing deep nonlinear relationships in multimodal data processing is solved. This achieves more efficient multimodal data alignment and semantic understanding, and improves the scalability and robustness of multimodal association modeling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIAXING UNIV
- Filing Date
- 2026-05-09
- Publication Date
- 2026-06-05
Smart Images

Figure CN122153812A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology; and more particularly to a quantum-enhanced multimodal data alignment method. Background Technology
[0002] Multimodal data includes heterogeneous data such as images and text. Multimodal alignment refers to establishing semantic relationships between different modalities, ensuring that the representations of each modality are aligned in a common space, and integrating multimodal information into a unified prediction. This leverages the strengths of each modality to improve overall model performance and enhance its ability to understand complex real-world scenarios, playing a crucial role in future intelligent information processing, content understanding, and knowledge discovery. To extract effective information and deep semantics from massive amounts of multimodal data, it is necessary to process heterogeneous data of various forms simultaneously to achieve high-precision and high-efficiency cross-modal semantic alignment. Pre-training methods based on contrastive learning learn to map images and text to the same semantic space by training on large-scale image-text pairs, achieving multimodal data alignment. However, their expressive power is limited when dealing with complex relationships with multiple relationships and levels, making it difficult to capture deep nonlinear relationships and complex semantic associations. Summary of the Invention
[0003] Based on the above analysis, this invention aims to disclose a quantum-enhanced multimodal data alignment method; overcoming the limitations of traditional methods in handling the curse of dimensionality and multimodal correlation modeling, further improving the scalability and robustness of multimodal alignment, and providing more powerful semantic understanding capabilities for processing massive multimodal information.
[0004] This invention discloses a quantum-enhanced multimodal data alignment method, comprising: S1. Receive multimodal raw data including text and images, and sequentially perform modality type definition, relation modeling and synchronous loading processing to obtain a multimodal dataset with alignment relationships; S2. Construct the quantum-enhanced multimodal contrastive learning model QCLIP and perform end-to-end training using a multimodal dataset; The model extracts a unified classical feature representation across modalities from multimodal data, and then processes it sequentially through quantum encoding, parameterized quantum circuit transformation, and measurement to obtain quantum-enhanced features. The quantum-enhanced features are then fused with the classical feature representation to form a quantum-classical hybrid feature representation. Cross-modal fusion similarity is calculated through a relation-aware attention mechanism, and feature alignment is performed by combining the joint optimization of multiple loss functions. A contrastive learning framework is used to simultaneously optimize model parameters by minimizing alignment loss, resulting in a fully trained model. S3. Receive actual multimodal data, call the quantum-enhanced multimodal contrastive learning model for processing, and output cross-modal semantic alignment results.
[0005] Further, S1 includes: S1-1. Define the modality type and intermodal relationships for multimodal raw data, including text and images, to obtain multimodal data with defined relationships. S1-2: Receive multimodal data with defined relationships, and obtain multimodal data with hierarchical storage through hierarchical organization and metadata association processing; S1-3: Receive multimodal data stored in a hierarchical manner, construct and map the retrieval path through an inverted index, and output a multimodal data index structure that supports fast retrieval; S1-4. Receive a multimodal data index structure that supports fast retrieval, process it by synchronously loading and aligning multimodal samples, and output a multimodal dataset with alignment relationships.
[0006] Furthermore, the quantum-enhanced multimodal contrastive learning model includes: a multimodal data encoding module, a quantum feature enhancement module, and a multi-relation alignment module; The multimodal data encoding module is used to encode each modal data using a pre-trained CLIP basic encoder, and extract cross-modal unified features through projection transformation and residual enhancement processing to obtain classical feature representations. The quantum feature enhancement module is used to project the classical feature representation onto the qubit space through the feature adaptation layer, encode it into a quantum state using an angle encoding method, perform quantum entanglement transformation through a parameterized quantum circuit, measure the quantum enhanced feature, and adaptively fuse it with the classical feature representation to form a quantum-classical hybrid feature representation. The multi-relation alignment module is used to calculate cross-modal fusion similarity based on the quantum-classical hybrid feature representation through a relation-aware attention mechanism, and perform feature alignment by combining contrastive learning loss, quantum fidelity loss and consistency constraint loss, and output aligned features.
[0007] Furthermore, the processing steps of the multimodal data encoding module include: 1) Receive multimodal data, process it through the CLIP basic encoder, input each type of modal data into the corresponding CLIP basic encoder, and output a modality-specific encoded feature vector; 2) Map the modality-specific encoded feature vectors to a unified feature space, so that the features of different modalities are projected onto a unified dimensional space to obtain unified dimensional features; 3) The unified dimensional features are subjected to residual enhancement processing through a gated residual network architecture to obtain a unified feature representation with enhanced feature representation capabilities; 4) Introduce feature alignment constraints to align features of different modalities in a unified space.
[0008] Furthermore, modal encoders ; “ " is a function compound operator; in, , For modality The original data space, For samples in the original data space, The output dimension of the basic encoder; For modality The data is input into the corresponding CLIP basic encoder to obtain the output feature vector; Projection function ; in, The projection matrix is used to realize the transformation from mode Basic encoder output dimension To a unified dimension Linear transformation; It is the bias vector; The representation layer normalization operation standardizes the input features; Residual enhancement processing function ; in, For the sigmoid function, Let Gaussian error be the activation function of the linear unit. This is the gate layer weight matrix. This is the gate layer bias vector. This is the transformation layer weight matrix. This is the bias vector for the transformation layer.
[0009] Furthermore, the feature alignment constraint is as follows: ; in, For distance measurement function, Representing modes sample Passed by modal encoder Processed feature vectors and modes sample Passed by modal encoder The differences between the processed feature vectors; To align the set of sample pairs; For regularization terms, This is the regularization coefficient.
[0010] Furthermore, the processing procedure of the quantum feature enhancement module includes: 1) Perform feature adaptation by projecting the classical feature representation onto the qubit space to obtain the qubit space features; 2) Perform angle encoding to map the spatial characteristics of the qubits to quantum rotation angles and output the angle-encoded quantum state; 3) Perform quantum transformation to entangle the angle-encoded quantum state and output the entangled quantum state; 4) Perform quantum measurement: measure the entangled quantum state through the Pauli z operator measurement layer, calculate the expected value of each qubit, and convert the probabilistic quantum measurement into a deterministic value to obtain the quantum enhancement feature; 5) Perform adaptive fusion: map the quantum enhanced features back to the original dimension through a classical post-processing network, and perform adaptive weighted fusion with the classical feature representation through learnable fusion coefficients to output a quantum-classical hybrid feature representation.
[0011] Furthermore, quantum-classical hybrid features Represented as: ; in, For learnable fusion coefficients, For quantum enhancement features, For classic post-processing networks, For modality The data is input into the modal encoder to obtain the basic encoded features.
[0012] Furthermore, the cross-modal attention weights determined using the relation-aware attention mechanism in the multi-relation alignment module are as follows: ; in, For modality For modes Cross-modal attention weights; For modality For modes Cross-modal fusion similarity; ; in, These are learnable equilibrium parameters; For modality For modes Quantum state similarity; For modality For modes Classic similarity.
[0013] Furthermore, the total loss function combining contrastive learning loss, quantum fidelity loss, and consistency constraint loss is: ; in, To compare learning loss; For quantum fidelity loss; Loss due to consistency constraints; , These are the weighting coefficients; gradient of total loss Calculated via backpropagation: ; in, Represents all trainable parameters; Comparative learning loss ; in, ; ; For cross-modal fusion of similarity No. Line number The elements of the column represent modalities. No. Individual samples and modalities No. The similarity of individual samples; For cross-modal fusion of similarity No. Line number The elements of the column represent modalities. No. Individual samples and modalities No. The similarity of individual samples; the similarity matrix For temperature parameters, Batch size; Quantum fidelity loss ; in, , For the sample , The density matrix; For the sample The positive sample density matrix; Quantum fidelity is used to measure the degree of overlap between two quantum states. It takes values [0,1], with values closer to 1 indicating greater similarity.
[0014] Consistency constraint loss ; in, For modality No. Quantum-classical hybrid features resulting from quantum enhancement and adaptive fusion of classical features from a single sample. For modality No. The basic encoded features extracted from each sample by the CLIP basic encoder. For modality No. Quantum-classical hybrid features resulting from quantum enhancement and adaptive fusion of classical features from a single sample. For modality No. The basic encoded features extracted from each sample by the CLIP basic encoder.
[0015] Compared with the prior art, the present invention has at least the following advantages: This invention addresses the need for accurate alignment and efficient utilization of massive multimodal data. Considering the significant heterogeneity between modalities and the complex and diverse semantic relationships, it utilizes the properties of quantum entanglement to construct a quantum-classical hybrid neural network. This network enables efficient joint representation and alignment of massive multimodal data. By learning complex cross-modal mapping relationships through variable quantum circuits, it overcomes the limitations of traditional methods in handling the curse of dimensionality and multimodal correlation modeling, further enhancing the scalability and robustness of multimodal alignment and providing stronger semantic understanding capabilities for processing massive multimodal information. Furthermore, its framework is general, training is stable, and it supports adaptive optimization, exhibiting good scalability. Attached Figure Description
[0016] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Figure 1 This is a flowchart of the quantum-enhanced multimodal data alignment method in an embodiment of the present invention; Figure 2 This is a structural diagram of the quantum-enhanced multimodal contrastive learning model in an embodiment of the present invention. Detailed Implementation
[0017] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and, together with the embodiments of the present invention, serve to illustrate the principles of the present invention.
[0018] One embodiment of the present invention discloses a quantum-enhanced multimodal data alignment method, such as... Figure 1 As shown, it includes: S1. Receive multimodal raw data including text and images, and sequentially perform modality type definition, relation modeling and synchronous loading processing to obtain a multimodal dataset with alignment relationships; S2. Construct the quantum-enhanced multimodal contrastive learning model QCLIP and perform end-to-end training using a multimodal dataset; The model extracts a unified classical feature representation across modalities from multimodal data, and then processes it sequentially through quantum encoding, parameterized quantum circuit transformation, and measurement to obtain quantum-enhanced features. The quantum-enhanced features are then fused with the classical feature representation to form a quantum-classical hybrid feature representation. Cross-modal fusion similarity is calculated through a relation-aware attention mechanism, and feature alignment is performed by combining the joint optimization of multiple loss functions. A contrastive learning framework is used to simultaneously optimize model parameters by minimizing alignment loss, resulting in a fully trained model. S3. Receive actual multimodal data, call the quantum-enhanced multimodal contrastive learning model for processing, and output cross-modal semantic alignment results.
[0019] Specifically, S1 includes: S1-1. Define the modality type and intermodal relationships for multimodal raw data, including text and images, to obtain multimodal data with defined relationships. Multimodal data involves different types of data modalities such as text and images. Each modality of data has specific data formats and processing requirements. For example, image modalities use raster image formats such as JPEG and PNG to store pixel matrices and color space information, while text modalities use text formats such as TXT and JSON to store character sequences and semantic tags.
[0020] Based on this, this embodiment establishes a multimodal data type definition mechanism for declaring and managing collections of multiple data modal types. ,in Indicates the number of supported modalities, and for each modality It has specific data formats and processing requirements.
[0021] Then, a multimodal relationship type definition mechanism is established to define the set of association relationship types between modalities. ,in , This represents the semantic connection between different modalities. Relationship types include: one-to-one mapping, indicating a strict bijective correspondence between two modal samples; one-to-many association, indicating that one modal sample corresponds to multiple samples of another modality; and many-to-many connection, indicating a complex correspondence network between two sets of modal samples.
[0022] S1-2: Receive multimodal data with defined relationships, and obtain multimodal data with hierarchical storage through hierarchical organization and metadata association processing; Design a multimodal data storage structure that adopts a hierarchical organization: the first layer is the dataset identifier, which stores dataset metadata; the second layer is the modality type identifier, which stores modality configuration information; the third layer is the data sample identifier, which stores basic sample information; and the fourth layer is the relation metadata, which stores the association information between modalities.
[0023] Multimodal data stored in a hierarchical manner is represented as: quadruple ; in Represents a dataset collection. , Number of datasets; Represents a set of modal types. ; Number of modal types; Represents the sample set, , The number of samples; Represents a set of relations. ,in, For a set of relation types, For the weight range; The quadruple is linked by the relational adjacency tensor. This indicates the strength of multiple types of relationships between samples; where Indicates in the dataset medium sample and Interval The strength of class relationships; among which, Represents the dataset medium sample and Interval The strength of class relationships.
[0024] S1-3: Receive multimodal data stored in a hierarchical manner, construct and map the retrieval path through an inverted index, and output a multimodal data index structure that supports fast retrieval; Building an inverted index Quickly retrieve data samples with specific relationships; among which, Indicates the dataset identifier, Indicates the modal type identifier. Indicates the data sample identifier. Indicates the relation type identifier. Indicates the relationship strength weight.
[0025] S1-4. Receive a multimodal data index structure that supports fast retrieval, process it by synchronously loading and aligning multimodal samples, and output a multimodal dataset with alignment relationships.
[0026] Establish a multimodal data synchronous loading mechanism to ensure that related modal data are loaded simultaneously and maintained in alignment. For each data sample... Synchronously load the representations of the samples in different modalities ;in, For the number of modes, For the sample In modality The features are represented and the alignment relationship between samples is established. Output synchronously loaded multimodal sample data; among which, Indices representing different modalities Represents data samples In modality and modality Alignment type between them; Defined based on configuration relationships. Automatically inferring relationships between samples, where Indicates the modal type, Indicates the relation type, Define the relation inference rules; output multimodal data with a multi-relation graph structure; Processed by batch data generator, in batches Output data, where, For batch Mid-mode eigenmatrix For the sample relationship tensor within the batch; For batch size, Given the number of relation types, output multimodal training data with batch results.
[0027] In step S2, a multimodal dataset is received, cross-modal unified features are extracted using a pre-trained CLIP basic encoder, the unified features are mapped to the quantum state space through angle encoding, nonlinear transformation is performed through a parameterized quantum circuit with entanglement gate, quantum enhancement features are obtained through Pauli measurement, and adaptive fusion of quantum enhancement features and classical features is achieved based on learnable coefficients to form a quantum-classical hybrid feature representation. like Figure 2 As shown, the quantum-enhanced multimodal contrastive learning model QCLIP includes: a multimodal data encoding module, a quantum feature enhancement module, and a multi-relation alignment module; The multimodal data encoding module is used to encode each modal data using a pre-trained CLIP basic encoder, and extract cross-modal unified features through projection transformation and residual enhancement processing to obtain classical feature representations. The quantum feature enhancement module is used to project the classical feature representation onto the qubit space through the feature adaptation layer, encode it into a quantum state using an angle encoding method, perform quantum entanglement transformation through a parameterized quantum circuit, measure the quantum enhanced feature, and adaptively fuse it with the classical feature representation to form a quantum-classical hybrid feature representation. The multi-relation alignment module is used to calculate cross-modal fusion similarity based on the quantum-classical hybrid feature representation through a relation-aware attention mechanism, and perform feature alignment by combining contrastive learning loss, quantum fidelity loss and consistency constraint loss, and output aligned features.
[0028] Specifically, the processing steps of the multimodal data encoding module include: 1) Receive multimodal data, process it through the CLIP basic encoder, input each type of modal data into the corresponding CLIP basic encoder, and output a modality-specific encoded feature vector; Specifically, each modal data Input to the corresponding CLIP base encoder ,in For modality The original data space, The output dimension of the basic encoder; For modality The data is input into the corresponding CLIP basic encoder to obtain the output feature vector; The CLIP basic encoder is designed according to the principle of modality specificity, that is, the most suitable coding architecture is selected or designed according to the data characteristics and information structure of each modality.
[0029] 2) Map the modality-specific encoded feature vectors to a unified feature space. This allows features from different modalities to be projected onto a unified model. In a dimensional space, unified dimensional features are obtained; The projection function is: ; in, For projection matrix, realize the projection from a specific dimension of the mode. To a unified dimension Linear transformation; By introducing translational degrees of freedom as the bias vector, the expressive power of the model is enhanced. The representation layer normalization operation standardizes the input features. express eigenvectors; 3) The unified dimensional features are subjected to residual enhancement processing through a gated residual network architecture to obtain a unified feature representation with enhanced feature representation capabilities; Residual enhancement module for residual enhancement processing The mathematical expression is: ; in, For the input feature vector, for Residual enhancement output, For the sigmoid function, Let Gaussian error be the activation function of the linear unit. This is the gate layer weight matrix. This is the gate layer bias vector. This is the transformation layer weight matrix. This is the bias vector for the transformation layer.
[0030] After steps 1)-3), a complete modal encoder is formed in the multimodal data encoding module. ,satisfy This encoder achieves an end-to-end mapping from the original modal space to a unified feature space, where " " is a function compound operator; For modality The data is input into the modal encoder to obtain the basic encoded features.
[0031] Modal encoders are cascaded combinations of multi-layer networks: ; in, express In the Layer transformation, total number of layers Ensure sufficient modeling capabilities.
[0032] 4) Introduce feature alignment constraints to align features of different modalities in a unified space.
[0033] To ensure that the features output by different modal encoders are aligned in a unified space, feature alignment constraints are introduced: ; in, For distance measurement function, Representing modes sample After the encoder Processed feature vectors and modes sample After the encoder The differences between the processed feature vectors; To align the set of sample pairs, This is a regularization term that controls encoder complexity and prevents overfitting. This is the regularization coefficient, which balances the alignment loss and model complexity.
[0034] Specifically, the processing procedure of the quantum feature enhancement module includes: 1) Perform feature adaptation by projecting the classical feature representation onto the qubit space to obtain the qubit space features; CLIP features typically have high dimensionality. In this embodiment, a feature adaptation layer is designed to project these high-dimensional features onto the quantum bit space. The nonlinear mapping from the CLIP feature space to the quantum space is established as follows:
[0035] in, For learnable weight matrix, For bias vectors, Number of qubits; modes Basic coding features The dimension is Then the eigenvectors mapped to quantum space Dimensions .
[0036] 2) Perform angle encoding to map the spatial characteristics of the qubits to quantum rotation angles and output the angle-encoded quantum state; Specifically, for each quantum bit Application of revolving doors Encoding classical features into quantum states Output angle-encoded quantum state; 3) Perform quantum transformation to entangle the angle-encoded quantum state and output the entangled quantum state; Specifically, quantum circuits use parameterized quantum gate sequences to transform quantum characteristics: ; in, It is a set of trainable parameters; For parameterized quantum circuits; For the number of quantum circuit layers, For entangled connections; To calculate the tensor product of qubits, To bypass Axis revolving door, indicating the first Layer 1 qubit Rotation of the axis, angle is ; To bypass Axis revolving door, indicating the first Layer 1 qubit Rotation of the axis, angle is ; For controlled NOT gate, it means that the first gate is the gate with the second gate as the control. The first qubit is a control bit, the second... A two-qubit entanglement gate with one qubit as the target bit.
[0037] The final entangled quantum state is: ; Entangled states cannot be decomposed into the product of two independent quantum states. In other words, before the final measurement, this correlation is global and not a simple linear combination, giving the model an exponential ability to express combinations of features.
[0038] 4) Perform quantum measurement: measure the entangled quantum state through the Pauli z operator measurement layer, calculate the expected value of each qubit, and convert the probabilistic quantum measurement into a deterministic value to obtain the quantum enhancement feature; Pauli measurement of each qubit Operator expectation: ; The final quantum state; For the first Pauli qubits matrix; Transforming probabilistic quantum measurements into deterministic numerical values yields quantum enhancement features. .
[0039] 5) Perform adaptive fusion: map the quantum enhanced features back to the original dimension through a classical post-processing network, and perform adaptive weighted fusion with the classical feature representation through learnable fusion coefficients to output a quantum-classical hybrid feature representation.
[0040] Among them, quantum-classical hybrid features Represented as: ; in, For learnable fusion coefficients, For quantum-classical integrated networks; This refers to the classical post-processing network (Multi-Layer Perceptron post-processing), whose function is to map the low-dimensional quantum-enhanced features obtained from quantum measurements back to the original feature dimensions so that they can be fused with classical features.
[0041] The complete encoding process of the quantum feature enhancement module can be represented as: ; in, For quantum-classical integrated networks, This represents the mapping process from fundamental characteristics to quantum measurements, which is executed sequentially: , , , , For modality Sample data CLIP encoding.
[0042] Specifically, the multi-relation alignment module leverages the quantum superposition property to enhance relation modeling capabilities and handle complex and diverse relation patterns in multimodal data.
[0043] In this module, the multi-relation alignment module receives basic encoded features, obtains quantum state representations through a quantum enhancement layer, and calculates quantum state similarity. It then adaptively fuses classical cosine similarity and quantum state similarity to construct cross-modal attention weights. include: 1) For two modes Basic coding features , First, their quantum representations are obtained through a quantum enhancement layer. Then calculate the quantum state similarity: ; To simultaneously utilize classical similarity and quantum correlation, cross-modal fusion similarity is defined as the basis for cross-modal attention weights. Classical similarity can be expressed as: ; The cross-modal fusion similarity can then be expressed as: ; 2) Cross-modal attention weights It can be represented as: ; in, These are learnable equilibrium parameters.
[0044] Cross-modal attention weights are used for cross-modal feature enhancement, enhancing the modality. Features and modes The feature can be represented as: ; ; Enhanced mode and modality Features can be used for downstream tasks.
[0045] The multi-relation alignment module combines contrastive learning loss, quantum fidelity loss, and consistency constraint loss for joint optimization to perform feature alignment; among them, 1) Based on the aforementioned hybrid similarity metric, a symmetric InfoNCE loss form is used for modality analysis. To mode and modality To mode Bidirectional contrastive learning, designing contrastive learning loss To maintain modality To mode Alignment capability; Comparative learning loss ; Modal To mode Contrast loss ; Modal To mode Contrast loss ; in For cross-modal fusion of similarity No. Line number The elements of a column, i.e., the diagonal elements, represent modalities. No. Individual samples and modalities No. Similarity of individual samples (similarity of positive sample pairs); For cross-modal fusion of similarity No. Line number The elements of a column, i.e., the off-diagonal elements, represent modalities. No. Individual samples and modalities No. The similarity of individual samples (similarity of negative sample pairs). The similarity matrix... For temperature parameters, For batch size, optimize modality. To mode and modality To mode Alignment ensures bidirectional consistency.
[0046] 2) Using quantum state fidelity as a measure of positive sample similarity: Designing quantum fidelity loss to enhance the model's quantum feature alignment capability; Quantum fidelity loss ; in, To compare the loss in quantum fidelity, , For the sample , The density matrix; For the sample The positive sample density matrix; Quantum fidelity is used to measure the degree of overlap between two quantum states. It takes values [0,1], with values closer to 1 indicating greater similarity.
[0047] That is, in the formula, Quantum fidelity score for positive sample pairs; The sum of quantum fidelity scores for all sample pairs.
[0048] 3) Receive quantum-classical hybrid features and CLIP original features as input, design consistency loss, and maintain semantic consistency; Consistency loss ; in, For modality No. Quantum-classical hybrid features resulting from quantum enhancement and adaptive fusion of classical features from a single sample. For modality No. The basic encoded features extracted from each sample by the CLIP basic encoder; For modality No. Quantum-classical hybrid features resulting from quantum enhancement and adaptive fusion of classical features from a single sample. For modality No. The basic encoded features extracted from each sample by the CLIP basic encoder.
[0049] 4) The total loss function combining contrastive learning loss, quantum fidelity loss, and consistency constraint loss is: ; , The weight coefficients are adjusted based on the actual training results to keep the total loss numerically stable and prevent any single loss from dominating the entire optimization process. By employing a contrastive learning framework, the parameters of the multimodal data encoding module, quantum feature enhancement module, and multi-relation alignment module are simultaneously optimized by minimizing the alignment loss, resulting in a fully trained model. The gradient of the total loss is calculated through backpropagation: ; in, This represents all trainable parameters.
[0050] Specifically, in step S3, the pre-trained quantum-enhanced multimodal contrastive learning model (QCLIP) is integrated into the actual multimodal data processing system. The system receives the actual multimodal data and uses the pre-trained QCLIP to align the data. Based on the modal combination and distribution characteristics of the input data, the system calls a quantum-encoded multimodal alignment model to calculate cross-modal semantic associations. Based on the actual alignment results, the system continuously updates the model parameters and alignment strategy to ensure continuous improvement in cross-modal semantic understanding alignment performance.
[0051] The processing procedure includes: 1) Multimodal data alignment; A standard interface is set up to receive multimodal data from different sources. The system automatically identifies the modality and distribution characteristics of the input data and dynamically calls the corresponding preprocessing modules based on the modal characteristics to transform the raw data into a unified format suitable for QCLIP model processing. The multimodal data encoding module uses the CLIP model encoder to unify the feature representation of different modalities. The quantum feature enhancement module uses quantum entanglement to enhance the model's feature representation capability. The multi-relation alignment module uses quantum entanglement to improve the model's ability to model complex relational patterns.
[0052] 2) Continuous optimization; For each batch of multimodal data input, the system executes a complete alignment process and outputs the alignment results. The system calculates and detects accuracy metrics in real time to quantify the alignment effect. Based on the accuracy of downstream tasks, it adjusts the variable quantum circuit parameters in the model and continuously adjusts the quantum encoding method to better capture complex correlations in the data.
[0053] This invention tests the performance of the QCLIP algorithm on six public datasets: CIFAR-10, CIFAR-100, MNIST, Rendered SST2, STL10, and Country211, for multimodal data alignment tasks. As shown in Table 1, the QCLIP method improves performance on the CIFAR-10, CIFAR-100, SST, and Country211 datasets, with a significant performance improvement on the Country211 dataset. The Country211 dataset contains 211 images. The QCLIP model can accurately model the complex relationships between data, learn complex mapping relationships between cross-modal data, and has strong feature representation capabilities, thus achieving a significant performance improvement on the Country211 dataset.
[0054] Table 1 Model Performance Comparison
[0055] In summary, the quantum-enhanced multimodal data alignment method of this embodiment is based on the need for accurate alignment and efficient utilization of massive multimodal data. Considering the real challenges of significant heterogeneity between modes and complex and diverse semantic relationships, a quantum-enhanced connection quantum-enhanced multimodal data alignment method is proposed, which improves the accuracy and robustness of cross-modal retrieval and understanding. Furthermore, due to its general framework, stable training, and support for adaptive optimization, it has good scalability.
[0056] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A quantum-enhanced multimodal data alignment method, characterized in that, include: S1. Receive multimodal raw data including text and images, and sequentially perform modality type definition, relation modeling and synchronous loading processing to obtain a multimodal dataset with alignment relationships; S2. Construct a quantum-enhanced multimodal contrastive learning model and perform end-to-end training using a multimodal dataset; The model extracts a unified classical feature representation across modalities from multimodal data, and then processes it sequentially through quantum encoding, parameterized quantum circuit transformation, and measurement to obtain quantum-enhanced features. The quantum-enhanced features are then fused with the classical feature representation to form a quantum-classical hybrid feature representation. Cross-modal fusion similarity is calculated through a relation-aware attention mechanism, and feature alignment is performed by combining the joint optimization of multiple loss functions. A contrastive learning framework is used to simultaneously optimize model parameters by minimizing alignment loss, resulting in a fully trained model. S3. Receive actual multimodal data, call the quantum-enhanced multimodal contrastive learning model for processing, and output cross-modal semantic alignment results.
2. The quantum-enhanced multimodal data alignment method according to claim 1, characterized in that, S1 includes: S1-1. Define the modality type and intermodal relationships for multimodal raw data, including text and images, to obtain multimodal data with defined relationships. S1-2: Receive multimodal data with defined relationships, and obtain multimodal data with hierarchical storage through hierarchical organization and metadata association processing; S1-3: Receive multimodal data stored in a hierarchical manner, construct and map the retrieval path through an inverted index, and output a multimodal data index structure that supports fast retrieval; S1-4. Receive a multimodal data index structure that supports fast retrieval, process it by synchronously loading and aligning multimodal samples, and output a multimodal dataset with alignment relationships.
3. The quantum-enhanced multimodal data alignment method according to claim 1, characterized in that, The quantum-enhanced multimodal contrastive learning model includes: a multimodal data encoding module, a quantum feature enhancement module, and a multi-relation alignment module; The multimodal data encoding module is used to encode each modal data using a pre-trained CLIP basic encoder, and extract cross-modal unified features through projection transformation and residual enhancement processing to obtain classical feature representations. The quantum feature enhancement module is used to project the classical feature representation onto the qubit space through the feature adaptation layer, encode it into a quantum state using an angle encoding method, perform quantum entanglement transformation through a parameterized quantum circuit, measure the quantum enhanced feature, and adaptively fuse it with the classical feature representation to form a quantum-classical hybrid feature representation. The multi-relation alignment module is used to calculate cross-modal fusion similarity based on the quantum-classical hybrid feature representation through a relation-aware attention mechanism, and perform feature alignment by combining contrastive learning loss, quantum fidelity loss and consistency constraint loss, and output aligned features.
4. The quantum-enhanced multimodal data alignment method according to claim 3, characterized in that, The processing steps of the multimodal data encoding module include: 1) Receive multimodal data, process it through the CLIP basic encoder, input each type of modal data into the corresponding CLIP basic encoder, and output a modality-specific encoded feature vector; 2) Map the modality-specific encoded feature vectors to a unified feature space, so that the features of different modalities are projected onto a unified dimensional space to obtain unified dimensional features; 3) The unified dimensional features are subjected to residual enhancement processing through a gated residual network architecture to obtain a unified feature representation with enhanced feature representation capabilities; 4) Introduce feature alignment constraints to align features of different modalities in a unified space.
5. The quantum-enhanced multimodal data alignment method according to claim 4, characterized in that, Modal encoder ; " " is a function compound operator; in, , For modality The original data space, For samples in the original data space, The output dimension of the basic encoder; For modality The data is input into the corresponding CLIP basic encoder to obtain the output feature vector; Projection function ; in, The projection matrix is used to realize the transformation from mode Basic encoder output dimension To a unified dimension Linear transformation; It is the bias vector; The representation layer normalization operation standardizes the input features; Residual enhancement processing function ; in, For the sigmoid function, Let Gaussian error be the activation function of the linear unit. This is the gate layer weight matrix. This is the gate layer bias vector. This is the transformation layer weight matrix. This is the bias vector for the transformation layer.
6. The quantum-enhanced multimodal data alignment method according to claim 5, characterized in that, The feature alignment constraint is: ; in, For distance measurement function, Representing modes sample Passed by modal encoder Processed feature vectors and modes sample Passed by modal encoder The differences between the processed feature vectors; To align the set of sample pairs; For regularization terms, is the regularization coefficient.
7. The quantum-enhanced multimodal data alignment method according to claim 3, characterized in that, The processing procedure of the quantum feature enhancement module includes: 1) Perform feature adaptation by projecting the classical feature representation onto the qubit space to obtain the qubit space features; 2) Perform angle encoding to map the spatial characteristics of the qubits to quantum rotation angles and output the angle-encoded quantum state; 3) Perform quantum transformation to entangle the angle-encoded quantum state and output the entangled quantum state; 4) Perform quantum measurement: measure the entangled quantum state through the Pauli z operator measurement layer, calculate the expected value of each qubit, and convert the probabilistic quantum measurement into a deterministic value to obtain the quantum enhancement feature; 5) Perform adaptive fusion: map the quantum enhanced features back to the original dimension through a classical post-processing network, and perform adaptive weighted fusion with the classical feature representation through learnable fusion coefficients to output a quantum-classical hybrid feature representation.
8. The quantum-enhanced multimodal data alignment method according to claim 7, characterized in that, Quantum-classical hybrid characteristics Represented as: ; in, For learnable fusion coefficients, For quantum enhancement features, For classic post-processing networks, For modality The data is input into the modal encoder to obtain the basic encoded features.
9. The quantum-enhanced multimodal data alignment method according to claim 3, characterized in that, The cross-modal attention weights determined using the relation-aware attention mechanism in the multi-relation alignment module are as follows: ; in, For modality For modes Cross-modal attention weights; For modality For modes Cross-modal fusion similarity; ; in, These are learnable equilibrium parameters; For modality For modes Quantum state similarity; For modality For modes Classic similarity.
10. The quantum-enhanced multimodal data alignment method according to claim 9, characterized in that, The total loss function, combining contrastive learning loss, quantum fidelity loss, and consistency constraint loss, is: ; in, To compare learning loss; For quantum fidelity loss; Loss due to consistency constraints; , These are the weighting coefficients; gradient of total loss Calculated via backpropagation: ; in, Represents all trainable parameters; Comparative learning loss ; in, ; ; For cross-modal fusion of similarity No. Line number The elements of the column represent modalities. No. Individual samples and modalities No. The similarity of individual samples; For cross-modal fusion of similarity No. Line number The elements of the column represent modalities. No. Individual samples and modalities No. The similarity of individual samples; the similarity matrix For temperature parameters, Batch size; Quantum fidelity loss ; in, , For the sample , The density matrix; For the sample The positive sample density matrix; Quantum fidelity is used to measure the degree of overlap between two quantum states, taking values [0,1], with values closer to 1 indicating greater similarity. Consistency constraint loss ; in, For modality No. Quantum-classical hybrid features resulting from quantum enhancement and adaptive fusion of classical features from a single sample. For modality No. The basic encoded features extracted from each sample by the CLIP basic encoder; For modality No. Quantum-classical hybrid features resulting from quantum enhancement and adaptive fusion of classical features from a single sample. For modality No. The basic encoded features extracted from each sample by the CLIP basic encoder.
Citation Information
Patent Citations
Multi-modal large model training data acquisition method and system
CN121388387A
Multi-modal sentiment analysis method based on large language model and quantum computing
CN121479556A