Multi-source heterogeneous data alignment method and device
Through the Transformer module of multi-head self-attention and the comparison learning method, the multi-source heterogeneous data is characterized by encoding and mapping, which solves the problem of inaccurate alignment of modal data in traditional methods, and improves the decision-making effect of multi-modal data fusion.
Patent Information
- Application Number
- CN202510489601.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The traditional single-modal method separates the semantic connections between different modal data, resulting in the multimodal data being unable to be accurately aligned before input into the fusion model, affecting the model's wrong cross-modal relationships and reducing the decision-making effect of the fusion model.
The Transformer module and contrast learning method of multi-head self-attention are used to extract the characteristics of multi-source heterogeneous data through the multi-head attention mechanism, and feature encoding and mapping are used for feature encoding and mapping to achieve accurate semantic alignment of cross-modal data.
The precise semantic alignment of multi-source heterogeneous data is realized, which lays a good foundation for subsequent multi-modal data fusion and improves the decision-making ability of multi-modal fusion decisions.
Smart Images

Figure CN120541366A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a method and device for aligning multi-source heterogeneous data. Background Art
[0002] Multi-source heterogeneous data refers to data with different sources and structures. Because multi-source heterogeneous data can more comprehensively mine its value, it is widely used in numerous fields, such as the Industrial Internet, smart cities, and finance. Taking the Industrial Internet as an example, it has driven the transformation from single-node digitization to integrated overall systems, from design, manufacturing, operations, to service. During this process, massive amounts of multi-source heterogeneous data, including text, images, audio, and video, originating from various network nodes are collected and transmitted to the Industrial Internet platform, laying the foundation for subsequent in-depth analysis and applications (such as the Industrial Internet Knowledge Question and Answer Big Model).
[0003] The focus of multi-source heterogeneous data aggregation is model analysis, that is, to deeply explore the value of massive data and promote the realization of scientific decision-making and intelligent applications driven by data. Traditional unimodal methods extract deep features from a single modality (such as text, images, and audio) for recognition. The significant drawback of unimodal methods is that they sever the semantic connection between different modal data, making these data functionally isolated from each other. For example, reasoning based on data from one modality is usually ineffective. In contrast, multimodal methods achieve better performance by extracting and fusing features from multiple modalities. However, if multimodal data is not accurately aligned before being input into the fusion model, the model will learn incorrect cross-modal relationships, which will greatly affect the decision-making effect of the fusion model and significantly reduce model performance. Summary of the Invention
[0004] The present disclosure aims to solve one of the technical problems in the related art at least to a certain extent.
[0005] To this end, a first embodiment of the present disclosure proposes a method for aligning multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data and first modality data. The method includes:
[0006] extracting a first text feature from the text data and a first modal feature from the first modal data respectively;
[0007] Encoding the first text feature and the first modal feature using a first shared encoder to obtain a second text feature and a second modal feature;
[0008] Inputting the second text feature and the second modality feature into a pre-trained first Transformer module, and obtaining a first text cross feature and a first modality cross feature through a multi-head attention mechanism;
[0009] The first text cross-feature and the first modal cross-feature are mapped to a first vector space through a pre-trained first contrastive learning module for feature alignment to obtain a vector representation of the first text cross-feature and a vector representation of the first modal cross-feature.
[0010] In some embodiments of the present disclosure, the first Transformer module is pre-trained through the following steps: obtaining text feature samples and modal feature samples; inputting the text feature samples and the modal feature samples into the first Transformer module, obtaining text cross-feature samples and first modal cross-feature samples through a multi-head attention mechanism, and calculating an attention score; based on the text cross-feature samples and the first modal cross-feature samples, the first Transformer module is trained with the goal of minimizing the maximum mean difference loss function based on an adaptive Gaussian kernel of an attention score, and the bandwidth of the Gaussian kernel function in the maximum mean difference loss function is determined by the attention score.
[0011] In some embodiments of the present disclosure, the maximum mean difference loss function is expressed as follows:
[0012]
[0013] Among them, l AS_MMD (X, Y) is the maximum mean difference loss function, X is the text cross feature sample set, x i is the i-th text cross-feature sample in the text cross-feature sample set, x j is the jth text cross-feature sample in the text cross-feature sample set, n is the number of text cross-feature samples in the text cross-feature sample set, Y is the first modal cross-feature sample set, y i is the i-th first modality cross-feature sample in the first modality cross-feature sample set, y j is the jth first modality cross-feature sample in the first modality cross-feature sample set, m is the number of first modality cross-feature samples in the first modality cross-feature sample set, k(x i ,y j ) is the Gaussian kernel function, s is the attention score, δ is the standard deviation of the Gaussian kernel function, and δ / s is the bandwidth of the Gaussian kernel function.
[0014] In some embodiments of the present disclosure, the first contrastive learning module is pre-trained by the following steps: obtaining multiple groups of cross-feature sample pairs, each group of the cross-feature sample pairs including a text cross-feature sample and a first modality cross-feature sample; calculating the cosine similarity of each group of the cross-feature sample pairs; dividing the multiple groups of cross-feature sample pairs into positive sample pairs and negative sample pairs based on the cosine similarity; calculating the contrast loss function based on the positive sample pairs and the negative sample pairs, adjusting the parameters in the first contrastive learning module by minimizing the contrast loss function, and training the first contrastive learning module.
[0015] In some embodiments of the present disclosure, the contrast loss function is expressed as follows:
[0016]
[0017] Among them, loss CLM is the contrast loss function, P(i) is the set of positive sample pairs, |P(i)| is the number of positive sample pairs, is the pth positive sample pair, is the i-th group of cross-feature sample pairs, and n is the number of the cross-feature sample pairs.
[0018] In some embodiments of the present disclosure, the multi-source heterogeneous data also includes second modality data; the data modality of the second modality data is different from that of the first modality data and the text data; the method also includes: extracting third modality features from the second modality data; encoding the first text features and the third modality features using a second shared encoder to obtain third text features and fourth modality features; inputting the third text features and the fourth modality features into a pre-trained second Transformer module, and obtaining second text cross-features and second modality cross-features through a multi-head attention mechanism; the second Transformer module has the same structure as the first Transformer module; mapping the second text cross-features and the second modality cross-features to a second vector space through a pre-trained second contrastive learning module for feature alignment, obtaining a vector representation of the second text cross-feature and a vector representation of the second modality cross-feature; and aligning the first modality cross-feature and the second modality cross-feature with reference to the vector representation of the first text cross-feature and the vector representation of the second text cross-feature.
[0019] A second embodiment of the present disclosure provides a multi-source heterogeneous data alignment device, wherein the multi-source heterogeneous data includes text data and first modality data, and the device includes:
[0020] a feature extraction module, configured to extract first text features from the text data and first modal features from the first modal data respectively;
[0021] an encoding module, configured to encode the first text feature and the first modal feature using a first shared encoder to obtain a second text feature and a second modal feature;
[0022] a coarse alignment module, configured to input the second text feature and the second modality feature into a pre-trained first Transformer module, and obtain a first text cross feature and a first modality cross feature through a multi-head attention mechanism;
[0023] A fine alignment module is used to map the first text cross-feature and the first modal cross-feature to a first vector space for feature alignment through a pre-trained first contrastive learning module to obtain a vector representation of the first text cross-feature and a vector representation of the first modal cross-feature.
[0024] In some embodiments of the present disclosure, a first training module is also included; wherein the first training module is used to: obtain text feature samples and modal feature samples; input the text feature samples and the modal feature samples into the first Transformer module, obtain text cross-feature samples and first modal cross-feature samples through a multi-head attention mechanism, and calculate the attention score; based on the text cross-feature samples and the first modal cross-feature samples, the first Transformer module is trained with the goal of minimizing the maximum mean difference loss function based on the attention score adaptive Gaussian kernel, and the bandwidth of the Gaussian kernel function in the maximum mean difference loss function is determined by the attention score.
[0025] In some embodiments of the present disclosure, a second training module is further included; wherein the second training module is used to: obtain multiple groups of cross-feature sample pairs, each group of the cross-feature sample pairs includes a text cross-feature sample and a first modality cross-feature sample; calculate the cosine similarity of each group of the cross-feature sample pairs; divide the multiple groups of cross-feature sample pairs into positive sample pairs and negative sample pairs based on the cosine similarity; calculate the contrast loss function based on the positive sample pairs and the negative sample pairs, adjust the parameters in the first contrast learning module by minimizing the contrast loss function, and train the first contrast learning module.
[0026] A third embodiment of the present disclosure provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0027] The memory stores computer-executable instructions;
[0028] The processor executes the computer-executable instructions stored in the memory to implement the method described in the first aspect.
[0029] The multi-source heterogeneous data alignment method provided by the present disclosure utilizes a Transformer module containing multi-head self-attention to process text data and first modality data, realizes coarse cross-modal alignment of first modality data and text data, and utilizes a contrastive learning method to realize fine-grained alignment between cross-modal data. It can realize accurate semantic alignment of multi-source heterogeneous data, lay a good foundation for subsequent multimodal data fusion, and improve the decision-making ability of multimodal fusion decision-making.
[0030] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0032] Figure 1 A flowchart of a multi-source heterogeneous data alignment method provided by an embodiment of the present disclosure;
[0033] Figure 2 An architectural diagram of a multi-source heterogeneous data alignment method provided by an embodiment of the present disclosure;
[0034] Figure 3 A schematic diagram of a multi-source heterogeneous data alignment device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0036] Specifically, the multi-source heterogeneous data alignment method and apparatus according to an embodiment of the present disclosure will be described below with reference to the accompanying drawings.
[0037] Figure 1 The following is a flow chart of a multi-source heterogeneous data alignment method provided by an embodiment of the present disclosure. The multi-source heterogeneous data includes text data and first modality data. Optionally, the first modality data may be audio data or image data, or may be other modality data, which is not limited by the present disclosure. Figure 1 As shown, the multi-source heterogeneous data alignment method may include the following steps:
[0038] Step 101: extract first text features from text data and first modal features from first modal data respectively.
[0039] Figure 2 This is an architectural diagram of a multi-source heterogeneous data alignment method provided by an embodiment of the present disclosure. Figure 2 As shown, in some embodiments of the present disclosure, a sentence-level BERT model can be used to extract the first text feature in the text data, which is represented by v t When the first modality data is image data, the pre-trained Detectron2 target detection toolkit can be used to extract image features (i.e., first modality features) from the image data, expressed as v p When the first modal data is audio data, the audio sequence can be converted into a fixed-length raw waveform with a sampling rate of 48kHz, and zero-padded for shorter audio. The pre-trained Wav2Vec2.0 model is used to extract audio features (i.e., first modal features) from the input audio data, which is expressed as v s .
[0040] Step 102: Encode the first text feature and the first modal feature using a first shared encoder to obtain a second text feature and a second modal feature.
[0041] The first shared encoder is used to encode the first text feature and the first modal feature into the same modal subspace. The second text feature v in the subspace t and the second modal feature (when the first modal data is image data, the second modal feature is represented by v pp ; When the first modal data is audio data, the second modal feature is represented by v ps ) retains the unique information of each modality (such as the syntax of text data, the texture of image data, and the intonation of audio data), enriching the common feature information independent of the modality.
[0042] Step 103: Input the second text feature and the second modal feature into the pre-trained first Transformer module, and obtain the first text cross feature and the first modal cross feature through the multi-head attention mechanism.
[0043] The second text feature and the second modal feature are input into the first Transformer module, and each modality is made aware of the other's features through the attention mechanism, and the first text cross feature v is extracted. cpt and the first modality cross feature (when the first modality data is image data, the first modality cross feature is represented by v cpp ; When the first modal data is audio data, the first modal cross feature is represented by v cps), achieving coarse semantic alignment of cross-modal features in a shared subspace.
[0044] In order to improve the convergence speed and model accuracy of the first Transformer module, in some embodiments of the present disclosure, the first Transformer module can be trained by the maximum mean difference loss function based on the attention score adaptive Gaussian kernel. In one implementation method, text data samples and first modal data samples for training can be obtained, and text feature samples and modal feature samples therein can be extracted respectively. The text feature samples and modal feature samples are input into the first Transformer module, and the text cross-feature samples and the first modal cross-feature samples are obtained through the multi-head attention mechanism, and the attention score is calculated. According to the text cross-feature samples and the first modal cross-feature samples, the first Transformer module is trained with the maximum mean difference loss function (AS-MMD loss function) based on the attention score adaptive Gaussian kernel as the training goal. Among them, the bandwidth of the Gaussian kernel function in the maximum mean difference loss function is determined by the attention score.
[0045] The bandwidth of the Gaussian kernel function is adaptively adjusted using the attention score. A larger bandwidth results in a flatter Gaussian kernel curve and a larger local influence range; a smaller bandwidth results in a steeper Gaussian kernel curve and a smaller local influence range. By minimizing the AS-MMD loss function to calculate the difference in probability distribution between text cross-feature samples and first-modality cross-feature samples, the unbiased estimate of sample embedding in the RKHS space is more accurate. This, in turn, leads to faster convergence of the Transformer network model during training, higher coarse alignment accuracy, and alignment between cross-features within the shared subspace.
[0046] Among them, the derivation process of the maximum mean difference loss function based on the attention score adaptive Gaussian kernel can be referred to as follows:
[0047]
[0048] Among them, formula (1) is the maximum mean difference loss function, X is the text cross feature sample set, x i is the i-th text cross-feature sample in the text cross-feature sample set, n is the number of text cross-feature samples in the text cross-feature sample set, Y is the first modal cross-feature sample set, y j is the jth first modal cross-feature sample in the first modal cross-feature sample set, m is the number of first modal cross-feature samples in the first modal cross-feature sample set, Φ is a mapping function, which maps the sample to a reproducing kernel Hilbert space (RKHS), and ‖·‖ represents the L2 norm.
[0049] RKHS is a complete high-dimensional inner product space, Φ(x i ) and Φ(y j ) can be adaptively performed by using the attention score-based Gaussian kernel function k(x i ,y i ) calculation, that is:
[0050]
[0051] Where s is the attention score and δ is the standard deviation of the Gaussian kernel function.
[0052] Based on formula (1) and formula (2), the maximum mean difference loss function based on the attention score adaptive Gaussian kernel proposed in this disclosure is obtained:
[0053]
[0054] Among them, l AS_MMD (X, Y) is the maximum mean difference loss function, X is the text cross feature sample set, x i is the i-th text cross-feature sample in the text cross-feature sample set, x j is the jth text cross-feature sample in the text cross-feature sample set, n is the number of text cross-feature samples in the text cross-feature sample set, Y is the first modal cross-feature sample set, y i is the i-th first modality cross-feature sample in the first modality cross-feature sample set, y j is the jth first modality cross-feature sample in the first modality cross-feature sample set, m is the number of first modality cross-feature samples in the first modality cross-feature sample set, k(x i ,y j ) is the Gaussian kernel function, s is the attention score, δ is the standard deviation of the Gaussian kernel function, and δ / s is the bandwidth of the Gaussian kernel function.
[0055] In step 104 , the first text cross-feature and the first modality cross-feature are mapped to a first vector space through a pre-trained first contrastive learning module for feature alignment to obtain a vector representation of the first text cross-feature and a vector representation of the first modality cross-feature.
[0056] The first contrastive learning module maps the first text cross-features and the first modal cross-features to the first vector space for fine-grained alignment between the features. Through pre-training, the first contrastive learning module can optimize the distribution distance of the first text cross-features and the first modal cross-features in the vector space based on the correlation between the cross-features, so that the distance between semantically related cross-features in the vector space is closer, and the distance between semantically unrelated cross-features in the vector space is farther.
[0057] In some embodiments of the present disclosure, multiple groups of cross-feature sample pairs can be obtained, each group of cross-feature sample pairs including a text cross-feature sample and a first modality cross-feature sample. The cosine similarity of each group of cross-feature sample pairs is calculated to represent the semantic similarity between the features. The multiple groups of cross-feature sample pairs are divided into positive sample pairs and negative sample pairs based on the cosine similarity. When the cosine similarity is equal to or higher than a preset threshold, the cross-feature sample pair is considered to be semantically related and recorded as a positive sample pair; when the cosine similarity is lower than the preset threshold, the cross-feature sample pair is considered to be semantically unrelated and recorded as a negative sample pair. The contrast loss function is calculated based on the positive sample pairs and the negative sample pairs, and the parameters in the first contrast learning module are adjusted by minimizing the contrast loss function. The first contrast learning module is trained to further improve the data alignment accuracy.
[0058] The contrast loss function is expressed as follows:
[0059]
[0060] Where, the numerator is the sum of the cosine similarities of all cross-feature samples and positive sample pairs, and the denominator is the sum of the cosine similarities of all cross-feature sample pairs (including positive samples). CLM is the contrast loss function, P(i) is the set of positive sample pairs, |P(i)| is the L1 norm of P(i), that is, the number of positive sample pairs, For the i-th group of cross-feature sample pairs (including a text cross-feature sample v cpt and a first modality cross-feature sample), is the pth positive sample pair, and n is the number of cross-feature sample pairs. When the first modality data is image data, the cross-feature sample pairs Expressed as When the first modality data is audio data, the cross feature sample pair Expressed as
[0061] By implementing the embodiments of the present disclosure, the text data and the first modality data are processed using a Transformer module containing multi-head self-attention, and coarse cross-modal alignment of the first modality data and the text data is achieved. The contrastive learning method is used to achieve fine-grained alignment between cross-modal data, which can achieve accurate semantic alignment of multi-source heterogeneous data, lay a good foundation for subsequent multimodal data fusion, and improve the decision-making ability of multimodal fusion decisions.
[0062] In addition to achieving alignment between text data and first modal data in the above-mentioned embodiments, data alignment between two or more modal data can also be achieved. Optionally, in some embodiments of the present disclosure, the multi-source heterogeneous data may also include second modal data, where the data modality of the second modal data is different from that of the first modal data and the text data. Taking the example of the first modal data being image data, the second modal data may be audio data or data of another modality, which is not limited in this disclosure. The multi-source heterogeneous data alignment method may also include: extracting third modal features from the second modal data. Using a second shared encoder, the first text features and the third modal features are encoded to obtain third text features and fourth modal features. The third text features and the fourth modal features are input into a pre-trained second Transformer module, and a multi-head attention mechanism is used to obtain second text cross-features and second modal cross-features. The second Transformer module has the same structure as the first Transformer module. The second text cross-features and the second modal cross-features are mapped to a second vector space by a pre-trained second contrastive learning module for feature alignment, thereby obtaining a vector representation of the second text cross-features and a vector representation of the second modal cross-features. With the vector representation of the first text cross-feature and the vector representation of the second text cross-feature as reference, feature alignment is performed on the first modal cross-feature and the second modal cross-feature, thereby achieving data alignment among the text data, the first modal data and the second modal data.
[0063] The training methods of the second Transformer module and the second contrastive learning module may refer to the training methods of the first Transformer module and the first contrastive learning module in the aforementioned embodiments, and will not be described in detail here.
[0064] Figure 3 This is a schematic diagram of a multi-source heterogeneous data alignment device provided by an embodiment of the present disclosure. Figure 3 As shown, the multi-source heterogeneous data alignment device may include: a feature extraction module 301 , an encoding module 302 , a coarse alignment module 303 and a fine alignment module 304 .
[0065] The feature extraction module 301 is used to extract first text features from the text data and first modal features from the first modal data respectively.
[0066] The encoding module 302 is configured to encode the first text feature and the first modal feature using a first shared encoder to obtain a second text feature and a second modal feature.
[0067] The coarse alignment module 303 is used to input the second text feature and the second modality feature into the pre-trained first Transformer module, and obtain the first text cross feature and the first modality cross feature through the multi-head attention mechanism.
[0068] The fine alignment module 304 is used to map the first text cross-features and the first modality cross-features to the first vector space for feature alignment through the pre-trained first contrastive learning module to obtain the vector representation of the first text cross-features and the vector representation of the first modality cross-features.
[0069] In some embodiments of the present disclosure, Figure 3 On the basis of the illustrated embodiment, the multi-source heterogeneous data alignment device may further include a first training module. The first training module is used to obtain text feature samples and modal feature samples. The text feature samples and modal feature samples are input into the first Transformer module, and text cross-feature samples and first modal cross-feature samples are obtained through a multi-head attention mechanism, and an attention score is calculated. Based on the text cross-feature samples and the first modal cross-feature samples, the first Transformer module is trained with the goal of minimizing the maximum mean difference loss function based on the attention score adaptive Gaussian kernel. The bandwidth of the Gaussian kernel function in the maximum mean difference loss function is determined by the attention score.
[0070] In some embodiments of the present disclosure, Figure 3 Based on the illustrated embodiment, the multi-source heterogeneous data alignment device may further include a second training module. The second training module is configured to: obtain multiple groups of cross-feature sample pairs, each group of cross-feature sample pairs including a text cross-feature sample and a first modality cross-feature sample. Calculate the cosine similarity of each group of cross-feature sample pairs. Divide the multiple groups of cross-feature sample pairs into positive sample pairs and negative sample pairs based on the cosine similarity. Calculate the contrast loss function based on the positive sample pairs and the negative sample pairs, adjust the parameters in the first contrast learning module by minimizing the contrast loss function, and train the first contrast learning module.
[0071] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0072] In order to implement the above embodiments, the present disclosure also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.
[0073] In order to implement the above embodiments, the present disclosure further proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0074] In order to implement the above embodiments, the present disclosure further provides a computer program product, including a computer program, which implements the methods provided in the above embodiments when executed by a processor.
[0075] In the descriptions of the aforementioned embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually inconsistent.
[0076] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0077] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0078] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0079] It should be understood that various parts of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0080] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0081] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0082] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. A person of ordinary skill in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for aligning multi-source heterogeneous data, wherein the multi-source heterogeneous data includes text data and first modality data, characterized in that: The method comprises the following steps: extracting a first text feature from the text data and a first modal feature from the first modal data respectively; Encoding the first text feature and the first modal feature using a first shared encoder to obtain a second text feature and a second modal feature; Inputting the second text feature and the second modality feature into a pre-trained first Transformer module, and obtaining a first text cross feature and a first modality cross feature through a multi-head attention mechanism; The first text cross-feature and the first modal cross-feature are mapped to a first vector space through a pre-trained first contrastive learning module for feature alignment to obtain a vector representation of the first text cross-feature and a vector representation of the first modal cross-feature.
2. The method according to claim 1, wherein The first Transformer module is pre-trained by the following steps: Obtain text feature samples and modal feature samples; Inputting the text feature sample and the modality feature sample into the first Transformer module, obtaining text cross-feature samples and first modality cross-feature samples through a multi-head attention mechanism, and calculating an attention score; According to the text cross-feature samples and the first modality cross-feature samples, the first Transformer module is trained with the goal of minimizing the maximum mean difference loss function based on the attention score adaptive Gaussian kernel, and the bandwidth of the Gaussian kernel function in the maximum mean difference loss function is determined by the attention score.
3. The method according to claim 2, wherein The maximum mean difference loss function is expressed as follows: Among them, L AS_MMD (X, Y) is the maximum mean difference loss function, X is the text cross feature sample set, x i is the i-th text cross-feature sample in the text cross-feature sample set, x j is the jth text cross-feature sample in the text cross-feature sample set, n is the number of text cross-feature samples in the text cross-feature sample set, Y is the first modal cross-feature sample set, y i is the i-th first modality cross-feature sample in the first modality cross-feature sample set, y j is the jth first modality cross-feature sample in the first modality cross-feature sample set, m is the number of first modality cross-feature samples in the first modality cross-feature sample set, k(x i ,y j ) is the Gaussian kernel function, s is the attention score, δ is the standard deviation of the Gaussian kernel function, and δ / s is the bandwidth of the Gaussian kernel function.
4. The method according to claim 1, wherein The first contrastive learning module is pre-trained by the following steps: Acquire multiple groups of cross-feature sample pairs, each group of cross-feature sample pairs including a text cross-feature sample and a first modality cross-feature sample; Calculating the cosine similarity of each group of cross-feature sample pairs; Dividing the plurality of groups of cross-feature sample pairs into positive sample pairs and negative sample pairs based on the cosine similarity; A contrast loss function is calculated based on the positive sample pairs and the negative sample pairs, and parameters in the first contrast learning module are adjusted by minimizing the contrast loss function to train the first contrast learning module.
5. The method according to claim 4, wherein The contrast loss function is expressed as follows: Among them, loss CLM is the contrast loss function, P(i) is the set of positive sample pairs, |P(i)| is the number of positive sample pairs, is the pth positive sample pair, is the i-th group of cross-feature sample pairs, and n is the number of the cross-feature sample pairs.
6. The method according to any one of claims 1 to 5, wherein The multi-source heterogeneous data further includes second modality data; the second modality data is different from the first modality data and the text data in data modality; the method further includes: extracting third modality features from the second modality data; Encoding the first text feature and the third modal feature using a second shared encoder to obtain a third text feature and a fourth modal feature; Inputting the third text feature and the fourth modality feature into a pre-trained second Transformer module, and obtaining a second text cross feature and a second modality cross feature through a multi-head attention mechanism; the second Transformer module has the same structure as the first Transformer module; Mapping the second text cross-feature and the second modality cross-feature to a second vector space for feature alignment using a pre-trained second contrastive learning module to obtain a vector representation of the second text cross-feature and a vector representation of the second modality cross-feature; The first modality cross-feature and the second modality cross-feature are feature aligned with each other using the vector representation of the first text cross-feature and the vector representation of the second text cross-feature as references.
7. A multi-source heterogeneous data alignment device, wherein the multi-source heterogeneous data includes text data and first modality data, characterized in that: include: a feature extraction module, configured to extract first text features from the text data and first modal features from the first modal data respectively; an encoding module, configured to encode the first text feature and the first modal feature using a first shared encoder to obtain a second text feature and a second modal feature; a coarse alignment module, configured to input the second text feature and the second modality feature into a pre-trained first Transformer module, and obtain a first text cross feature and a first modality cross feature through a multi-head attention mechanism; A fine alignment module is used to map the first text cross-feature and the first modal cross-feature to a first vector space for feature alignment through a pre-trained first contrastive learning module to obtain a vector representation of the first text cross-feature and a vector representation of the first modal cross-feature.
8. The device according to claim 7, wherein Also included is a first training module; wherein the first training module is used to: Obtain text feature samples and modal feature samples; Inputting the text feature sample and the modality feature sample into the first Transformer module, obtaining text cross-feature samples and first modality cross-feature samples through a multi-head attention mechanism, and calculating an attention score; According to the text cross-feature samples and the first modality cross-feature samples, the first Transformer module is trained with the goal of minimizing the maximum mean difference loss function based on the attention score adaptive Gaussian kernel, and the bandwidth of the Gaussian kernel function in the maximum mean difference loss function is determined by the attention score.
9. The device according to claim 7, wherein Also included is a second training module; wherein the second training module is used to: Acquire multiple groups of cross-feature sample pairs, each group of cross-feature sample pairs including a text cross-feature sample and a first modality cross-feature sample; Calculating the cosine similarity of each group of cross-feature sample pairs; Dividing the plurality of groups of cross-feature sample pairs into positive sample pairs and negative sample pairs based on the cosine similarity; A contrast loss function is calculated based on the positive sample pairs and the negative sample pairs, and parameters in the first contrast learning module are adjusted by minimizing the contrast loss function to train the first contrast learning module.
10. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-modal retrieval method based on multilayer semantic alignment
CN112966127A
Dam defect image text cross-modal retrieval method and model
CN113220919A
Multi-modal representation learning method based on text guide image block screening
CN117421591A
Intelligent monitoring image defogging method and system based on coring attention
CN119273581A
Cross-modal retrieval method and device, electronic equipment and storage medium
CN119597939A