Training method and device for fusion representation model of multi-modal data

CN120654193BActive Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510802645.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2026-09-25
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

然而目前的深度学习模型过于关注精度,忽略了不同模态数据之间的关联,导致输出结果可信度低

Benefits of technology

[0032]根据本说明书实施例的第九方面,提供了一种计算机程序产品,包括计算机程序或指令,该计算机程序或指令被处理器执行时实现上述针对多模态数据的融合表示模型的训练方法、针对多模态数据的融合表示方法、政务数据处理方法的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654193B_ABST
    Figure CN120654193B_ABST
Patent Text Reader

Abstract

The embodiment of the specification provides a training method and device for a fusion representation model of multi-modal data, wherein the method comprises: determining text word embedding and non-text embedding of target data, and inputting the text word embedding and the non-text embedding into a pre-training model; generating first alignment embedding and text sentence embedding based on the text word embedding and the non-text embedding through a first encoder, generating second alignment embedding and text segment embedding based on the first alignment embedding and the text sentence embedding through a second encoder; generating a sentence positive-negative sample pair according to the first alignment embedding and the text sentence embedding, and generating a segment positive-negative sample pair according to the second alignment embedding and the text segment embedding; performing a fusion representation training task for the pre-training model through the sentence positive-negative sample pair and the segment positive-negative sample pair to obtain a fusion representation model. The fusion representation model has the ability to input a unified representation of fusion non-text modal information, and the interpretability of data representation is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of natural language processing technology, and in particular to a training method and apparatus for a fusion representation model of multimodal data. Background Technology

[0002] With the rapid development of the internet and digital media, a vast amount of multimodal data has been created and shared. Information from different modalities can provide complementary perspectives and rich semantic information. By comprehensively utilizing multimodal information, more comprehensive and accurate search results can be obtained, meeting diverse user needs. Currently, data in many industries is transforming from unimodal to multimodal data; for example, in the government sector, report data has shifted from traditional text to multimodal data including text, images, and videos. With the widespread application of deep learning technology, deep learning models are used to integrate multimodal data in order to better process it. However, current deep learning models focus too much on accuracy, neglecting the correlation between different modalities, resulting in low reliability of output results. Therefore, improving the reliability of models when processing multimodal data is a pressing issue that needs to be addressed. Summary of the Invention

[0003] In view of this, embodiments of this specification provide a training method for a fusion representation model of multimodal data, a fusion representation method for multimodal data, and a government data processing method. One or more embodiments of this specification also relate to a training device for a fusion representation model of multimodal data, a fusion representation device for multimodal data, a government data processing device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, a method for training a fusion representation model for multimodal data is provided, comprising:

[0005] Determine the text word embeddings and non-text embeddings of the target data, and input the text word embeddings and non-text embeddings into a pre-trained model, the pre-trained model including a first encoder pair and a second encoder pair;

[0006] The first encoder generates a first alignment embedding and a text sentence embedding based on the text word embedding and non-text embedding. The second encoder generates a second alignment embedding and a text fragment embedding based on the first alignment embedding and the text sentence embedding. The alignment dimensions of the first alignment embedding and the text sentence embedding are different from those of the second alignment embedding and the text fragment embedding.

[0007] Sentence positive and negative sample pairs are generated based on the first alignment embedding and the text sentence embedding, and fragment positive and negative sample pairs are generated based on the second alignment embedding and the text fragment embedding;

[0008] By performing a fusion representation training task on the pre-trained model using the sentence positive and negative sample pairs and the segment positive and negative sample pairs, a fusion representation model is obtained.

[0009] According to a second aspect of the embodiments of this specification, a method for fusion representation of multimodal data is provided, comprising:

[0010] Multimodal data is identified and input into a fusion representation model, wherein the fusion representation model is trained using the training method described above for a fusion representation model of multimodal data.

[0011] Obtain the target fusion representation corresponding to the multimodal data output by the fusion representation model, wherein the target fusion representation is used to perform data processing tasks on the multimodal data.

[0012] According to a third aspect of the embodiments of this specification, a government data processing method is provided, comprising:

[0013] The target government data for the data analysis task is determined, wherein the target government data includes government data in at least two modalities;

[0014] The target government data is input into the fusion representation model to obtain the target fusion representation corresponding to the target government data output by the fusion representation model. The fusion representation model is trained by the above-mentioned training method for the fusion representation model of multimodal data.

[0015] The data analysis task is performed based on the target fusion representation to obtain the data analysis results corresponding to the target government data.

[0016] According to a fourth aspect of the embodiments of this specification, a training apparatus for a fusion representation model of multimodal data is provided, comprising:

[0017] An input module is configured to determine the text word embeddings and non-text embeddings of the target data, and input the text word embeddings and the non-text embeddings into a pre-trained model, the pre-trained model including a first encoder pair and a second encoder pair;

[0018] The first generation module is configured to generate a first alignment embedding and a text sentence embedding based on the text word embedding and non-text embedding using the first encoder, and to generate a second alignment embedding and a text fragment embedding based on the first alignment embedding and the text sentence embedding using the second encoder, wherein the alignment dimensions of the first alignment embedding and the text sentence embedding are different from those of the second alignment embedding and the text fragment embedding.

[0019] The second generation module is configured to generate sentence positive and negative sample pairs based on the first alignment embedding and the text sentence embedding, and to generate segment positive and negative sample pairs based on the second alignment embedding and the text segment embedding;

[0020] The execution module is configured to perform a fusion representation training task on the pre-trained model using the sentence positive and negative sample pairs and the segment positive and negative sample pairs, to obtain the fusion representation model.

[0021] According to a fifth aspect of the embodiments of this specification, a fusion representation apparatus for multimodal data is provided, comprising:

[0022] The input module is configured to determine multimodal data and input the multimodal data into the fusion representation model, wherein the fusion representation model is trained using the above-described training method for the fusion representation model of multimodal data.

[0023] The acquisition module is configured to acquire a target fusion representation corresponding to the multimodal data output by the fusion representation model, wherein the target fusion representation is used to perform data processing tasks on the multimodal data.

[0024] According to a sixth aspect of the embodiments of this specification, a government data processing apparatus is provided, comprising:

[0025] The determination module is configured to determine the target government data for the data analysis task, wherein the target government data includes government data in at least two modalities;

[0026] The input module is configured to input the target government data into the fusion representation model to obtain the target fusion representation of the target government data output by the fusion representation model, wherein the fusion representation model is trained by the above-mentioned training method for the fusion representation model of multimodal data;

[0027] The execution module is configured to perform the data analysis task based on the target fusion representation and obtain the data analysis results corresponding to the target government data.

[0028] According to a seventh aspect of the embodiments of this specification, a computing device is provided, comprising:

[0029] Memory and processor;

[0030] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the above-mentioned training method for the fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method.

[0031] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the training method for a fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method described above.

[0032] According to a ninth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described training method for a fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method.

[0033] One embodiment of this specification implements feature extraction for both textual and non-textual modal data of target data, determining the text word embeddings and non-textual embeddings of the target data. A first encoder of a pre-trained model generates first aligned embeddings and text sentence embeddings based on text word embeddings and non-textual embeddings along the text sentence alignment dimension. A second encoder generates second aligned embeddings and text segment embeddings based on the first aligned embeddings and text sentence embeddings along the text segment alignment dimension. This allows for the fusion of non-textual modal information using two different alignment dimensions, enabling its combination with corresponding text sentences and segments, thus enhancing the semantic association between different modalities. Subsequently, positive and negative sample pairs are generated based on the first aligned embeddings and text sentence embeddings, and positive and negative sample pairs are generated based on the second aligned embeddings and text segment embeddings. These positive and negative sample pairs are then used for training a fusion representation task. This allows the fusion representation model trained through contrastive learning to output interpretable multimodal data fusion representations, facilitating the execution of downstream tasks based on the model's output fusion representations, improving the reliability of the model output and the accuracy of downstream tasks. Attached Figure Description

[0034] Figure 1 A schematic diagram of the framework of a training method for a fusion representation model of multimodal data according to an embodiment of this specification is shown;

[0035] Figure 2 A flowchart is shown illustrating a training method for a fusion representation model of multimodal data according to an embodiment of this specification;

[0036] Figure 3 This specification shows a schematic diagram of a self-attention mask layer provided in one embodiment;

[0037] Figure 4 A flowchart illustrating the processing steps of a training method for a fusion representation model of multimodal data, provided in one embodiment of this specification, is shown.

[0038] Figure 5 This specification shows a schematic diagram of the structure of a training device for a fusion representation model of multimodal data according to one embodiment of the present specification;

[0039] Figure 6 A flowchart is shown illustrating a method for fusing and representing multimodal data according to an embodiment of this specification;

[0040] Figure 7 This specification shows a schematic diagram of the structure of a fusion representation apparatus for multimodal data provided in one embodiment;

[0041] Figure 8 A flowchart of a government data processing method according to an embodiment of this specification is shown;

[0042] Figure 9 This specification shows a schematic diagram of the structure of a government data processing device according to one embodiment;

[0043] Figure 10 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0044] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0045] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0046] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0047] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0048] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0049] Multimodal data refers to data from different modalities or sources. This data typically has different forms and characteristics, and each modality can provide unique information. By combining data from multiple modalities, a scene or task can be described more comprehensively. Multimodal data can include various forms of data such as text modality, image modality, audio modality, video modality, sensor data, and structured data.

[0050] Contrastive learning is a technique widely used in unsupervised or self-supervised learning, particularly in multimodal learning, representation learning, and generative models. Its core idea is to learn useful representations of data by comparing positive sample pairs (similar samples) and negative sample pairs (dissimilar samples). This method effectively captures the inherent structure of data without relying on explicit label information.

[0051] Currently, with the increasing abundance of industry data across various sectors, industry data is gradually shifting from unimodal to multimodal data. For example, in the field of autonomous driving, vehicles need to process data from various sensors (such as radar, lidar, GPS, and cameras) in real time to obtain comprehensive information about the surrounding environment. In the smart home field, systems may need to process multimodal data including audio (such as voice commands), video (such as images from surveillance cameras), and text (such as user text commands). In the government sector, systems may need to process multimodal data including text, images, audio, and video to better understand and meet user needs. Therefore, the integration and application of massive amounts of multimodal data has become an urgent problem to be solved.

[0052] Current technical solutions often employ machine learning or deep learning models to process multimodal data. Multimodal fusion, a key technology in artificial intelligence, aims to deeply integrate data from different modalities to construct a more comprehensive and accurate information representation. However, current multimodal data fusion methods often use feature-level concatenation or decision-level voting, achieving modal information superposition solely through weight allocation. This leads to semantic gaps caused by heterogeneity between different modalities, resulting in logical breaks between text data and other modal data. Consequently, the final fused representation lacks interpretability, making it difficult to process subsequent downstream tasks and resulting in unsatisfactory execution results.

[0053] Based on this, this specification provides a training method for a fusion representation model of multimodal data, a fusion representation method for multimodal data, and a government data processing method. This specification also provides a training device for a fusion representation model of multimodal data, a fusion representation device for multimodal data, a government data processing device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0054] See Figure 1 , Figure 1This diagram illustrates a framework of a training method for a fusion representation model of multimodal data according to an embodiment of this specification. The training method framework can be divided into four processing parts: modality feature extraction, text-level semantic alignment, model training, and unified representation generation. Modality feature extraction can be understood as extracting features from target data for different modalities, including extracting text lexical embeddings for text data and extracting initial non-textual embeddings for non-text data. Document-level semantic alignment can be understood as performing semantic alignment at different levels, such as sentences and fragments. Sentence-level semantic alignment yields lexical-aligned non-textual embeddings and multimodal fusion sentence-level text embeddings; fragment-level semantic alignment yields sentence-aligned non-textual embeddings and multimodal fusion fragment-level text embeddings. Model training can be understood as training the model using the embeddings obtained from semantic alignment, including sentence-level comparative learning using lexical-aligned non-textual embeddings and multimodal fusion sentence-level text embeddings, and fragment-level comparative learning using sentence-aligned non-textual embeddings and multimodal fusion fragment-level text embeddings. The fusion model obtained through two different levels of comparative learning tasks can generate a multimodal fusion representation based on modal semantic alignment in unified representation generation. This fusion representation combines semantic information between different modalities, making the multimodal fusion representation interpretable. This facilitates the execution of downstream tasks based on the fusion representation output by the model, improving the reliability of the model output and the accuracy of downstream tasks.

[0055] See Figure 2 , Figure 2 A flowchart is shown of a training method for a fusion representation model of multimodal data according to an embodiment of this specification, specifically including the following steps.

[0056] Step 202: Determine the text word embeddings and non-text embeddings of the target data, and input the text word embeddings and non-text embeddings into the pre-trained model, the pre-trained model including a first encoder pair and a second encoder pair.

[0057] In this context, target data can be understood as data used for model training. Target data includes various modalities, such as text, image, and audio modal data. The target data also varies across different industry sectors. For example, in government affairs, target data might be government documents, which could include text and images of service guides. In the smart home sector, target data could be smart home processing data, including images and video data captured by cameras, voice commands received by voice control devices, and text commands manually entered by users. Across different industry sectors, the target data can be categorized into text and non-text modal data. By training the multimodal data fusion representation model provided in this manual, text and non-text modal data are semantically fused, enabling the trained model to output interpretable fusion representations, thus improving the reliability of the model output and the accuracy of downstream tasks.

[0058] In practical applications, since the target data includes both textual and non-textual modal data, feature extraction can be performed separately for each modality. Accordingly, feature extraction from textual modal data yields textual word embeddings, while feature extraction from non-textual modal data yields non-textual embeddings. These textual and non-textual embeddings are then input into a pre-trained model for training.

[0059] In practice, text word embeddings can be obtained by extracting features from text modal data in the target data, while non-text embeddings can be obtained by extracting features from non-text modal data in the target data. After obtaining the text word embeddings and non-text embeddings, the first encoder pair and the second encoder pair in the pre-trained model can be used to combine the features of the input text word embeddings and non-text embeddings, thereby achieving data integration of multiple modalities. The first encoder pair and the second encoder pair can be understood as encoders in the encoding layer of the pre-trained model. Each of the first encoder pair and the second encoder pair contains two encoders, which encode the input embeddings respectively. The pre-trained model can be a model that has been pre-trained on data. Pre-trained models refer to deep learning models that have been pre-trained on large-scale datasets. These models can be designed specifically for specific types of data (such as text, images, etc.) or built to handle multimodal data. The main advantage of pre-trained models is that they can capture general features in the data, which makes them perform well in various tasks and can be quickly adapted to new tasks or domains through fine-tuning.

[0060] Furthermore, in order to determine the text word embeddings and non-text embeddings of the target data, it is necessary to distinguish the data of different modalities in the target data. Specifically, determining the text word embeddings and non-text embeddings of the target data includes: extracting the text data and non-text data from the target data; encoding the text data and non-text data to obtain the text word embeddings of the text data and the non-text embeddings corresponding to the non-text data.

[0061] Text data can be understood as text modal data in the target data, such as the text in a document. Non-text data can be understood as modal data in the target data other than text modal data, such as images in a document. By extracting text and non-text data from the target data, the target data can be distinguished by text and non-text modalities. Subsequently, text data in the text modal can be encoded to obtain text word embeddings, and non-text data in the non-text modal can be encoded to obtain non-text embeddings.

[0062] In practical applications, natural language processing techniques can be used to extract text from target data. For example, extracting all text content from a PDF document or scraping text information from a webpage. For non-text data, such as images, audio, and video, appropriate methods are needed to extract them. For example, extracting all images from an ebook with illustrations or separating audio and video frames from a video with subtitles.

[0063] In practice, text data typically undergoes preprocessing (such as word segmentation and stop word removal) and is then encoded using pre-trained language models like BERT (Bidirectional Encoder Representations from Transformers) and RoBERTa (Robustly optimized BERT approach) to obtain text word embeddings. These embeddings capture the semantic information within the text. Non-text data is encoded using different methods depending on its type. For images, Convolutional Neural Networks (CNNs) can be used to generate image feature vectors; for audio, Mel-frequency cepstral coefficients or other acoustic feature representations can be used; and for video, a combination of visual and auditory features may be necessary. By encoding both text and non-text data, text word embeddings for the text data and non-text embeddings for the corresponding non-text data can be obtained.

[0064] In one specific embodiment of this specification, the target data is government information documents. Text data of the text modality is extracted from these documents. This text data is the primary data source for analyzing government information documents. The pre-trained weights of FinBERT (a natural language processing model based on the BERT architecture specifically designed for the financial field) can be used to initialize the lexical encoder. The initialized lexical encoder is then used to encode the text data, obtaining text word embeddings containing contextual information. For other modalities of non-text data in government information documents, encoders such as VIT (Vision Transformer, an image classification model) and Whisper (an automatic speech recognition model) can be used to encode the non-text data and obtain non-text embeddings.

[0065] Based on this, by distinguishing different modalities in the target data and encoding textual and non-textual modal data separately, textual word embeddings and non-textual embeddings can be obtained. Subsequently, multimodal fusion is performed based on the textual word embeddings and non-textual embeddings to achieve a more comprehensive data understanding and analysis.

[0066] Step 204: Generate a first alignment embedding and a text sentence embedding based on the text word embedding and non-text embedding using the first encoder; generate a second alignment embedding and a text fragment embedding based on the first alignment embedding and the text sentence embedding using the second encoder, wherein the alignment dimensions of the first alignment embedding, the text sentence embedding and the second alignment embedding and the text fragment embedding are different.

[0067] After obtaining the text word embeddings and non-text embeddings, the text word embeddings and non-text embeddings can be input into the pre-trained model. Semantic alignment is performed through the alignment of the pre-trained model, which facilitates the fusion of information from different modalities.

[0068] In practical applications, semantic gaps exist between different modalities of data. To effectively integrate information from different modalities within the target data, semantic alignment of embeddings for different modalities can be achieved using the first and second encoder pairs in a pre-trained model. The pre-trained model utilizes encoder pairs with different alignment dimensions to achieve semantic alignment between non-textual and textual modalities at multiple document structure levels.

[0069] In practice, a first encoder processes text word embeddings and non-text embeddings to generate a first alignment embedding and a text sentence embedding; a second encoder processes the first alignment embedding and the text sentence embedding to generate a second alignment embedding and a text fragment embedding. Since the first alignment embedding and the text sentence embedding are semantic alignments at the sentence level, while the second alignment embedding and the text fragment embedding are semantic alignments at the fragment level, the alignment dimensions of the first alignment embedding and the text sentence embedding differ from those of the second alignment embedding and the text fragment embedding.

[0070] Furthermore, each encoder pair in the pre-trained model contains two Transformer encoders. These two encoders form an encoder pair, and the two-level encoder pair achieves semantic alignment at different document levels. Specifically, at the sentence level, the first encoder pair includes a first alignment encoder and a sentence encoder. Generating a first aligned embedding and a text sentence embedding based on the text word embeddings and non-text embeddings using the first encoder pair includes: aligning the text word embeddings and non-text embeddings using the first alignment encoder to obtain the first aligned embedding; and encoding the text word embeddings using the sentence encoder to obtain the text sentence embedding.

[0071] The first encoder pair includes a first alignment encoder and a sentence encoder. The first alignment encoder performs semantic alignment processing on the text word embeddings and non-text embeddings to obtain the first aligned embedding. The sentence encoder maps the text word embeddings to sentence-level text sentence embeddings. The first aligned embedding can be understood as the aligned non-text modal embedding, and the text sentence embedding can be understood as the sentence-level embedding obtained by mapping the text word embeddings.

[0072] In practical applications, since data from different modalities may have different feature representations and distributions, direct fusion may lead to information loss or misunderstanding. Therefore, a first alignment encoder is used to align data from different modalities (text word embeddings and non-text embeddings) at the semantic level to obtain first aligned embeddings. Specifically, this can be achieved by using attention-based methods such as cross-attention to enhance the correlation between different modalities. This allows the encoder model to learn how to adjust non-text embeddings to make them more consistent with text word embeddings in the same semantic space. The sentence encoder then transforms the input text word embeddings into sentence-level embeddings. This means that the sentence encoder must consider not only the meaning of individual words but also the context and overall meaning of the entire sentence. Specifically, this can be achieved by using a self-attention mechanism to capture the dependencies between words and further refining sentence-level features through multi-layer stacking, ultimately outputting a fixed-length vector, the text sentence embedding, to represent the semantic information of the entire sentence.

[0073] In practice, a switchable masked self-attention layer is used between the first alignment encoder and the sentence encoder for information exchange. See [link to relevant documentation]. Figure 3 , Figure 3 A schematic diagram of a self-attention masking layer provided in one embodiment of this specification is shown, wherein the self-attention masking layer can switch between a single-modal self-attention masking mode and a multi-modal causal attention masking mode. During semantic alignment processing, the masking self-attention layer between the two encoders switches to the single-modal self-attention masking mode, which ensures that there is no information leakage between the two encoders.

[0074] In a specific embodiment of this specification, the text word embeddings and non-text embeddings are aligned using a first encoder. The specific alignment steps are as follows:

[0075] Step 1: Construct an aligned cross-attention layer to extract semantic features from non-textual embeddings, see formula (1).

[0076]

[0077] in, It is the query vector, derived from the hidden state h of the text modal data. f K is obtained after linear transformation. x =Linear K (x), V x =Linear V (x) are the key vector and value vector of the non-text modal data, respectively, which are obtained by linear transformation of the non-text modal input x, and f represents the semantic alignment level.

[0078] Step 2: Align the semantic features of the non-text embeddings with the information of the text word embeddings through the alignment cross-attention layer to obtain the aligned non-text embeddings, i.e., the first aligned embeddings. Where k is the length of the embedding.

[0079] The alignment method described above can be used to align text word embeddings with non-text embeddings, resulting in a first aligned embedding. The input text word embeddings are then processed by a sentence encoder. Text sentence embeddings mapped to the sentence level

[0080] Based on this, by constructing a switchable masked self-attention layer and an aligned cross-attention layer, semantic alignment between non-textual and textual modalities can be effectively achieved. This method fully utilizes the powerful modeling capabilities of attention mechanisms, enabling the capture of complex relationships across modalities, and is widely applied in multimodal tasks such as image-text retrieval, image-text generation, and sentiment analysis. In the embodiments of this specification, the first alignment encoder and sentence encoder in the first encoder pair achieve semantic correspondence and sentence embedding mapping at the sentence level for the input text word embeddings and non-text embeddings, obtaining the first aligned embedding and text sentence embedding. Subsequently, the first aligned embedding and text sentence embedding can be input into the second encoder pair for further processing.

[0081] Furthermore, to achieve semantic alignment at the fragment level, the first alignment embedding and the text sentence embedding obtained from sentence-level processing need to be input into a second encoder pair for further processing. Specifically, the second encoder pair includes a second alignment encoder and a fragment encoder; generating a second alignment embedding and a text fragment embedding based on the first alignment embedding and the text sentence embedding using the second encoder pair includes: aligning the first alignment embedding and the text sentence embedding using the second alignment encoder to generate the second alignment embedding; and encoding the text sentence embedding using the fragment encoder to obtain the text fragment embedding.

[0082] Correspondingly, the second encoder pair also includes two encoders: a second alignment encoder and a fragment encoder. The second alignment encoder is used to semantically align the first alignment embedding and the text sentence embedding to generate the second alignment embedding. The fragment encoder is used to map the text sentence embedding to fragment-level text fragment embeddings.

[0083] In practical applications, the first encoder achieves semantic alignment between text and non-text modalities at the sentence level. To maintain semantic consistency between text and non-text modalities at the fragment level, corresponding semantic alignment processing is required at the fragment level. The second alignment encoder further aligns the first alignment embedding and the text sentence embedding at the fragment level, aiming to extend the sentence-level alignment to a larger contextual scope, i.e., the fragment level, thereby generating a second alignment embedding to capture more complex cross-modal relationships. The fragment encoder maps the sentence-level text sentence embedding to the fragment-level embedding representation, considering the relationships between multiple sentences during the mapping process, thus generating a higher-level fragment-level embedding, i.e., the text fragment embedding.

[0084] In specific implementation, the semantic alignment processing of the second encoder pair can be the same as that of the first encoder pair. The difference is that the input object of the second encoder pair is the output of the first encoder pair. That is, the second encoder pair processes the first alignment embedding and text sentence embedding output by the first encoder pair, and the output of the second encoder pair is the second alignment embedding and text fragment embedding.

[0085] In one specific embodiment of this specification, the second alignment code in the second encoder pair embeds the first alignment. With text sentence embedding Perform semantic alignment to obtain the aligned non-textual modal embedding, i.e., the second aligned embedding. The segment encoder in the second encoder pair embeds the input text sentence. Text fragment embedding mapped to fragment level

[0086] In summary, by employing a semantic hierarchical design, progressively expanding from the sentence level to the fragment level, the model is allowed to gradually refine information at different granularities, thereby better capturing local and global semantic relationships. Semantic alignment is achieved at each level, enabling non-textual modal embeddings and textual modal embeddings to interact within the same semantic space, ensuring effective fusion of information from both modalities and providing powerful representational capabilities and flexibility for complex tasks.

[0087] Step 206: Generate sentence positive and negative sample pairs based on the first alignment embedding and the text sentence embedding, and generate fragment positive and negative sample pairs based on the second alignment embedding and the text fragment embedding.

[0088] Since the semantic gap between different modalities can lead to poor interpretability of the fused features, this embodiment employs a contrastive learning method to train the model, thereby achieving semantic alignment between non-textual and textual modal information and improving the interpretability of the fused representation. Because the semantic alignment stage is divided into sentence-level and segment-level tasks, the contrastive learning training task is also divided into two training tasks: sentence-level and segment-level.

[0089] In practical applications, the training samples differ for different training tasks at different levels. Therefore, it is necessary to construct sentence positive and negative sample pairs for performing sentence-level contrastive learning tasks based on the first alignment embedding and text sentence embedding obtained from sentence-level semantic alignment processing. Each sentence positive and negative sample pair includes both sentence positive and negative sample pairs. Similarly, based on the second alignment embedding and text segment embedding obtained from segment-level alignment processing, segment positive and negative sample pairs are constructed for performing segment-level contrastive learning tasks. Each segment positive and negative sample pair includes both segment positive and negative sample pairs.

[0090] Furthermore, generating sentence positive and negative sample pairs based on the first alignment embedding and the text sentence embedding includes: determining a first text sentence embedding and a second text sentence embedding in the text sentence embedding based on the first alignment embedding, wherein the first text sentence embedding and the first alignment embedding have a semantic relationship; constructing sentence positive sample pairs based on the first alignment embedding and the first text sentence embedding, and constructing sentence negative sample pairs based on the first alignment embedding and the second text sentence embedding.

[0091] In this context, since the first aligned embedding is an embedding representation of the non-textual modality generated through semantic alignment of text word embeddings and non-textual embeddings, in order to construct positive sample pairs of sentences with similar semantic information, it is necessary to determine the first and second text sentence embeddings from the text sentence embeddings based on the first aligned embedding. The first text sentence embedding and the first aligned embedding have a semantic relationship, meaning that the first text sentence embedding and the first aligned embedding come from the same context or content. This can be understood as follows: when the non-textual modality is image data, the semantically related text modality is the text data describing that image. Therefore, based on the first aligned embedding, the semantically related first text sentence embedding and the unrelated second text sentence embedding can be filtered out from the text sentence embeddings.

[0092] In practical applications, positive sentence pairs are constructed by first-aligned embeddings of non-textual modalities and corresponding semantically related first-text sentence embeddings. These positive sentence pairs represent similar semantic information. Negative sentence pairs are constructed by first-aligned embeddings and unrelated second-text sentence embeddings. The purpose of negative sentence pairs is to distinguish semantically unrelated sample pairs. This allows the model to bring non-textual modal embeddings closer to aligned textual modal embeddings (positive sample pairs) while pushing unrelated textual modal embeddings further away (negative sample pairs). When selecting semantically related first-text sentence embeddings based on first-aligned embeddings, this semantic relationship can be determined through annotation information, contextual association, or other prior knowledge.

[0093] In practice, for the first alignment embedding of each non-text modality Select the corresponding sentence-level text modality embedding As positive sample pairs And the remaining sentence-level text embeddings as negative sample pairs

[0094] In a specific embodiment of this specification, text sentence embeddings are classified according to a first alignment embedding. The classification result includes a first text sentence embedding and a second text sentence embedding. The first alignment embedding is combined with the first text sentence embedding to construct a sentence positive sample pair, and the first alignment embedding is combined with the second text sentence embedding to construct a sentence negative sample pair. The sentence positive sample pair and the sentence negative sample pair are called sentence positive-negative sample pairs, which are used to perform subsequent sentence-level contrastive learning training tasks.

[0095] Based on this, positive sentence sample pairs are constructed by embedding the first text sentence that has a semantic relationship with the first alignment embedding and then embedding it with the first alignment embedding; negative sentence sample pairs are constructed by embedding the second text sentence that is unrelated to the first alignment embedding and then embedding it with the first alignment embedding. This allows the model to shorten the distance between positive sample pairs and widen the distance between negative sample pairs in the semantic space during subsequent contrastive learning training tasks. This not only narrows the semantic gap between data from different modalities but also significantly improves the quality and interpretability of the fused representation, providing strong support for multimodal tasks (such as text retrieval, sentiment analysis, and content understanding).

[0096] Further, generating positive and negative sample pairs of fragments based on the second alignment embedding and the text fragment embedding includes: determining a first text fragment embedding and a second text fragment embedding in the text fragment embedding based on the second alignment embedding, wherein the first text fragment embedding and the second alignment embedding have a semantic relationship; constructing positive sample pairs of fragments based on the second alignment embedding and the second text fragment embedding; and constructing negative sample pairs of fragments based on the second alignment embedding and the second text fragment embedding.

[0097] Similar to the construction of sentence-level positive and negative sample pairs, the construction of segment-level positive and negative sample pairs involves determining the first and second text segment embeddings based on the second alignment embedding within the text segment embedding. The first text segment embedding and the second alignment embedding are semantically related, while the second text segment embedding and the second alignment embedding are unrelated. The second alignment embedding is combined with the first text segment embedding to construct a segment positive sample pair, and the second alignment embedding is combined with the second text segment embedding to construct a segment negative sample pair. These segment positive and negative sample pairs are used to perform subsequent segment-level contrastive learning training tasks.

[0098] In practical applications, for each non-text modality, the second alignment embedding... Select the corresponding fragment-level text modality embedding As positive sample pairs and the remaining fragment-level text embeddings as negative sample pairs

[0099] In a specific embodiment of this specification, text fragment embeddings are classified according to a second alignment embedding. The classification results include a first text fragment embedding and a second text fragment embedding. The second alignment embedding is combined with the second text fragment embedding to construct a positive fragment sample pair, and the second alignment embedding is combined with the second text fragment inclusion to construct a negative fragment sample pair.

[0100] Based on this, positive sample pairs are constructed by embedding first text segments that are semantically related to the second alignment embedding and then embedding them with the second alignment embedding; negative sample pairs are constructed by embedding second text segments that are unrelated to the second alignment embedding and then embedding them with the second alignment embedding. This allows the model to shorten the distance between positive sample pairs and widen the distance between negative sample pairs in the semantic space during subsequent contrastive learning training tasks. This not only narrows the semantic gap between different modalities of data but also significantly improves the quality and interpretability of the fused representation, providing strong support for multimodal tasks (such as text retrieval, sentiment analysis, and content understanding).

[0101] Step 208: Perform a fusion representation training task on the pre-trained model using the sentence positive and negative sample pairs and the segment positive and negative sample pairs to obtain the fusion representation model.

[0102] The fusion representation training task is used to train the pre-trained model, enabling it to possess the capability of multimodal data fusion representation. In practical applications, after obtaining sentence positive-negative sample pairs and segment positive-negative sample pairs, the fusion representation training task can be performed based on these pairs to train the pre-trained model into a fusion representation model. The fusion representation model can be understood as the model obtained after training, possessing the ability to output multimodal data fusion representations.

[0103] In practice, since sentence positive and negative sample pairs are sentence-level positive and negative sample pairs, and segment positive and negative sample pairs are segment-level positive and negative sample pairs, the fusion representation training task for the pre-trained model is performed based on sentence positive and negative sample pairs and segment positive and negative sample pairs. Therefore, sentence-level and segment-level training tasks are performed separately. That is, the fusion representation training task is divided into sentence-level training tasks and segment-level training tasks.

[0104] Specifically, the fusion representation training task includes a first training task and a second training task; the fusion representation training task for the pre-trained model is performed using the sentence positive and negative sample pairs and the segment positive and negative sample pairs, including: performing the first training task for the pre-trained model using the sentence positive and negative sample pairs; and performing the second training task for the pre-trained model using the segment positive and negative sample pairs.

[0105] The first training task can be understood as a sentence-level training task, which is performed using positive and negative sample pairs of sentences; the second training task can be understood as a segment-level training task, which is performed using positive and negative sample pairs of segments.

[0106] In practical applications, the semantic gap between different modalities can lead to poor interpretability of the fused features. This specification addresses this issue by employing a contrastive learning method to train a pre-trained model, thereby achieving semantic alignment between non-textual and textual information and improving the interpretability of the final fused representation. When training the fusion representation on the pre-trained model, the training task is divided into sentence-level and fragment-level tasks based on document multi-level hierarchy, achieving semantic alignment at both the sentence and fragment levels and enhancing the model's ability to understand data at different levels.

[0107] In a specific embodiment of this specification, a first training task is performed using sentence positive sample pairs and sentence negative sample pairs, and a second training task is performed using segment positive sample pairs and segment negative sample pairs. It should be noted that the execution of the first and second training tasks can be sequential. For example, the first training task can be executed first to obtain model A after the first training task, and then the second training task can be performed based on model A to obtain the final model corresponding to the fusion representation training task. In another case, the first and second training tasks can also be executed in parallel. For example, the first and second training tasks can be executed simultaneously to obtain model A after the first training task and model B after the second training task, and then model A and model B can be fused to obtain the final model corresponding to the fusion representation training task.

[0108] Based on this, a multimodal data fusion model was trained using a contrastive learning method. This ensured that the aligned non-textual modal representations were closer to their corresponding textual modal representations and more distinct from mismatched textual modal representations, thereby improving the accuracy of semantic alignment between modalities. Furthermore, during the contrastive learning process, the fusion representation training task was divided into different levels of training tasks, bridging the semantic gap between different modalities and enhancing the interpretability of the final data representation.

[0109] Furthermore, performing the first training task for the pre-trained model using the sentence positive and negative sample pairs includes: calculating the first sample pair similarity of the sentence positive and negative sample pairs by performing the first training task, and calculating the first loss value corresponding to the pre-trained model based on the first sample pair similarity; and tuning the pre-trained model using the first loss value.

[0110] The first training task can be understood as a sentence-level training task. The training objective of the first training task is to maximize the similarity within positive sample pairs of sentences and minimize the similarity within negative sample pairs of sentences, so that the generated non-textual modal embeddings correspond to the textual modal embeddings to the greatest extent. When performing the first training task using positive and negative sample pairs of sentences, the first sample pair similarity of the positive and negative sample pairs of sentences can be calculated. The first sample pair similarity is the highest similarity among the pairwise similarities within the positive and negative sample pairs of sentences. The formula for calculating the first sample pair similarity is given in formula (2).

[0111]

[0112] In practical applications, the sample pair similarity of positive and negative sample pairs of sentences is calculated separately. Then, the pair with the highest similarity is selected as the overall similarity, i.e., the first sample pair similarity. The first loss value is calculated using the first sample pair. Based on the first loss value, the model parameters of the pre-trained model are adjusted to complete the first training task of the pre-trained model. Subsequently, the pre-trained model can continue to be trained according to the above steps until the model training stops. Specifically, by comparing the similarity between positive and negative sample pairs, the non-textual modal embeddings are aligned to the data distribution space of sentence-level textual modal embeddings. The loss function for contrastive learning uses the softmax cross-entropy loss function. The first loss value is calculated using the softmax cross-entropy loss function. The calculation formula of the softmax cross-entropy loss function is shown in formula (3).

[0113]

[0114] Based on this, the first training task at the sentence level is performed on the pre-trained model by constructing positive and negative sample pairs of sentences. This enables the model to distinguish between matching and non-matching modal information when processing multimodal data at the sentence level. This not only narrows the semantic gap between different modal data, but also significantly improves the quality and interpretability of the fused representation.

[0115] Accordingly, the second training task for the pre-trained model is performed using the positive and negative sample pairs of the fragments, including: calculating the second sample pair similarity of the positive and negative sample pairs of the fragments by performing the second training task, and calculating the second loss value corresponding to the pre-trained model based on the second sample pair similarity; and tuning the parameters of the pre-trained model using the second loss value.

[0116] The second training task can be understood as a fragment-level training task. The training objective of the second training task is to maximize the similarity within positive sample pairs of fragments and minimize the similarity within negative sample pairs of fragments, so that the generated non-textual modal embeddings correspond to the textual modal embeddings to the greatest extent possible. When performing the second training task using fragment positive and negative sample pairs, the second sample pair similarity can be calculated. The second sample pair similarity is the highest similarity among the pairwise similarities within the fragment positive and negative sample pairs. The formula for calculating the second sample pair similarity is given in formula (4).

[0117]

[0118] In practical applications, the sample pair similarity of positive and negative sample pairs is calculated separately. Then, the pair with the highest similarity is selected as the overall similarity, i.e., the second sample pair similarity. The second loss value is calculated using the second sample pair. Based on the second loss value, the model parameters of the pre-trained model are adjusted to complete the second training task of the pre-trained model. The pre-trained model can then be trained again following the above steps until the model training stops. Specifically, by comparing the similarity between positive and negative sample pairs, the non-textual modal embeddings are aligned to the data distribution space of the fragment-level textual modal embeddings. The loss function for contrastive learning is the softmax cross-entropy loss function. The second loss value is calculated using the softmax cross-entropy loss function. The calculation formula for the softmax cross-entropy loss function is shown in formula (5).

[0119]

[0120] Based on this, a second training task at the fragment level is performed on the pre-trained model by constructing fragment positive and negative sample pairs. This enables the model to distinguish between matching and non-matching modal information when processing fragment-level multimodal data. This not only narrows the semantic gap between different modal data, but also significantly improves the quality and interpretability of the fused representation.

[0121] In summary, after the first training task at the sentence level and the second training task at the segment level, a fusion representation model for the training sequence can be obtained. This fusion representation model can output a unified representation of the text space that incorporates non-textual modal information. In the fusion representation model, the mask between encoders is switched from the attention layer to... Figure 3 The multimodal causal self-attention masking mechanism allows each text modal embedding to acquire information from all non-text modal embeddings and the already generated text modal embeddings during generation. Based on this mechanism, the encoder in the model can output fragment-level hybrid text embeddings that fuse text and non-text modal information. That is, to generate a unified and interpretable multimodal data representation, where M is the number of fragments.

[0122] This specification provides a training method for a fusion representation model of multimodal data, comprising: determining text word embeddings and non-text embeddings of target data; inputting the text word embeddings and non-text embeddings into a pre-trained model, the pre-trained model including a first encoder pair and a second encoder pair; generating a first aligned embedding and a text sentence embedding based on the text word embeddings and non-text embeddings using the first encoder pair; generating a second aligned embedding and a text fragment embedding based on the first aligned embedding and the text sentence embedding using the second encoder pair; wherein the alignment dimensions of the first aligned embedding and the text sentence embedding are different from those of the second aligned embedding and the text fragment embedding; generating sentence positive and negative sample pairs based on the first aligned embedding and the text sentence embedding; generating fragment positive and negative sample pairs based on the second aligned embedding and the text fragment embedding; and performing a fusion representation training task on the pre-trained model using the sentence positive and negative sample pairs and the fragment positive and negative sample pairs to obtain a fusion representation model. This method achieves feature extraction of both text modal data and non-text modal data from the target data, and determines the text word embeddings and non-text embeddings of the target data. The pre-trained model uses a first encoder to generate first-aligned embeddings and text-sentence embeddings based on text word embeddings and non-text embeddings in the text sentence alignment dimension. A second encoder generates second-aligned embeddings and text-sentence embeddings based on the first-aligned embeddings and text-sentence embeddings in the text fragment alignment dimension. This allows for the fusion of non-text modal information through two different alignment dimensions, enabling its combination with corresponding text sentences and fragments, thus enhancing the semantic association between different modalities. Subsequently, positive and negative sample pairs are generated based on the first-aligned embeddings and text-sentence embeddings, and positive and negative sample pairs are generated based on the second-aligned embeddings and text-fragment embeddings. These positive and negative sample pairs are then used for training the fusion representation task. This allows the fusion representation model trained through contrastive learning to output interpretable multimodal data fusion representations, facilitating the execution of downstream tasks based on the model's output fusion representations, and improving the reliability of the model output and the accuracy of downstream tasks.

[0123] The following is in conjunction with the appendix Figure 4 Taking the application of the training method for the fusion representation model of multimodal data provided in this specification in the field of government affairs as an example, the training method for the fusion representation model of multimodal data will be further explained. Figure 4 The flowchart of a training method for a fusion representation model of multimodal data provided in one embodiment of this specification is shown, which specifically includes the following steps.

[0124] Step 402: Extract text data and non-text data from the target data, encode the text data and non-text data to obtain the text word embeddings of the text data and the non-text embeddings corresponding to the non-text data.

[0125] In one feasible approach, the target data is government document data, which includes various modalities such as text, images, and videos. The modal data within the document data is differentiated to obtain text data (text modality) and non-text data (non-text modality). The text and non-text data are then encoded separately to obtain text word embeddings and non-text embeddings.

[0126] Step 404: Input the text word embeddings and non-text embeddings into the pre-trained model, which includes a first encoder pair and a second encoder pair.

[0127] In one feasible approach, the extracted text word embeddings and non-text embeddings are input into a pre-trained model, and semantic alignment and mapping at different levels are performed through the first encoder pair and the second encoder pair in the pre-trained model.

[0128] Step 406: Align the text word embeddings and non-text embeddings using a first alignment encoder to obtain a first aligned embedding, and encode the text word embeddings using a sentence encoder to obtain a text sentence embedding.

[0129] In one feasible approach, a first alignment encoder performs sentence-level semantic alignment processing on the text word embeddings and non-text embeddings to obtain the first aligned embeddings of the non-text modality. A sentence encoder then encodes the text word embeddings to obtain sentence-level text sentence embeddings.

[0130] Step 408: Align the first alignment embedding and the text sentence embedding using the second alignment encoder to generate the second alignment embedding, and encode the text sentence embedding using the fragment encoder to obtain the text fragment embedding.

[0131] In one feasible approach, a second alignment encoder performs fragment-level semantic alignment processing on the first alignment embedding and the text sentence embedding to obtain a non-text modality second alignment embedding. A fragment encoder then encodes the text sentence embedding to obtain fragment-level text fragment embeddings.

[0132] Step 410: Generate positive and negative sample pairs of sentences based on the first alignment embedding and the text sentence embedding.

[0133] In one feasible approach, a first text sentence embedding and a second text sentence embedding are determined in the text sentence embedding based on a first alignment embedding, wherein the first text sentence embedding and the first alignment embedding have a semantic relationship, and a sentence positive sample pair is constructed based on the first alignment embedding and the first text sentence embedding, and a sentence negative sample pair is constructed based on the first alignment embedding and the second text sentence embedding.

[0134] Step 412: Generate positive and negative sample pairs of fragments based on the second alignment embedding and the text fragment embedding.

[0135] In one feasible approach, a first text fragment embedding and a second text fragment embedding are determined in the text fragment embedding based on a second alignment embedding, wherein the first text fragment embedding and the second alignment embedding have a semantic relationship, and a positive fragment sample pair is constructed based on the second alignment embedding and the second text fragment embedding, and a negative fragment sample pair is constructed based on the second alignment embedding and the second text fragment embedding.

[0136] Step 414: By performing the first training task, calculate the first sample pair similarity of positive and negative sample pairs of sentences, and calculate the first loss value corresponding to the pre-trained model based on the first sample pair similarity, and use the first loss value to tune the parameters of the pre-trained model.

[0137] In one feasible approach, a sentence-level first training task is performed based on sentence positive and negative sample pairs. The sample pair similarity of sentence positive and negative sample pairs is calculated separately. The highest similarity among the calculated sample pair similarities is selected as the first sample similarity. The cross-entropy loss value is calculated based on the first sample similarity as the first loss value. The pre-trained model is then tuned using the first loss value.

[0138] Step 416: By performing the second training task, calculate the second sample pair similarity of positive and negative sample pairs of the fragment, and calculate the second loss value corresponding to the pre-trained model based on the second sample pair similarity. Use the second loss value to tune the parameters of the pre-trained model.

[0139] In one feasible approach, a second training task at the fragment level is performed based on positive and negative sample pairs. The similarity between positive and negative sample pairs is calculated separately. The highest similarity among the calculated similarity scores is selected as the second sample similarity. A cross-entropy loss value is then calculated based on this second sample similarity and used as the second loss value to tune the pre-trained model. Through the sentence-level first training task and the fragment-level second training task described above, a fusion representation model is trained. This fusion representation model can effectively understand government information documents and outputs a multimodal fusion representation based on these documents, facilitating subsequent processing of downstream tasks based on this multimodal fusion representation.

[0140] This specification provides a training method for a fusion representation model of multimodal data. It extracts features from both textual and non-textual modal data of the target data, determining the text word embeddings and non-textual embeddings. A first encoder of the pre-trained model generates first-aligned embeddings and text sentence embeddings based on text word embeddings and non-textual embeddings along the text sentence alignment dimension. A second encoder generates second-aligned embeddings and text segment embeddings based on the first-aligned embeddings and text sentence embeddings along the text segment alignment dimension. This allows for the fusion of non-textual modal information using two different alignment dimensions, combining it with corresponding text sentences and segments to enhance the semantic association between different modalities. Subsequently, positive and negative sample pairs are generated based on the first-aligned embeddings and text sentence embeddings, and positive and negative sample pairs are generated based on the second-aligned embeddings and text segment embeddings. These positive and negative sample pairs are then used for fusion representation training. This enables the fusion representation model, trained through contrastive learning, to output interpretable multimodal data fusion representations, facilitating the execution of downstream tasks based on the model's output fusion representations and improving the reliability of the model output and the accuracy of downstream tasks.

[0141] Corresponding to the above method embodiments, this specification also provides embodiments of a training apparatus for a fusion representation model of multimodal data. Figure 5 This diagram illustrates a structural schematic of a training apparatus for a fusion representation model of multimodal data, provided in one embodiment of this specification. Figure 5 As shown, the device includes:

[0142] Input module 502 is configured to determine the text word embeddings and non-text embeddings of target data, and input the text word embeddings and the non-text embeddings into a pre-trained model, the pre-trained model including a first encoder pair and a second encoder pair;

[0143] The first generation module 504 is configured to generate a first alignment embedding and a text sentence embedding based on the text word embedding and non-text embedding through the first encoder, and to generate a second alignment embedding and a text fragment embedding based on the first alignment embedding and the text sentence embedding through the second encoder, wherein the alignment dimensions of the first alignment embedding and the text sentence embedding are different from those of the second alignment embedding and the text fragment embedding.

[0144] The second generation module 506 is configured to generate sentence positive and negative sample pairs based on the first alignment embedding and the text sentence embedding, and to generate segment positive and negative sample pairs based on the second alignment embedding and the text segment embedding;

[0145] The execution module 508 is configured to perform a fusion representation training task on the pre-trained model using the sentence positive and negative sample pairs and the segment positive and negative sample pairs, to obtain the fusion representation model.

[0146] Optionally, the first generation module 504 is further configured to perform alignment processing on the text word embedding and the non-text embedding through the first alignment encoder to obtain a first aligned embedding; and to perform encoding processing on the text word embedding through the sentence encoder to obtain a text sentence embedding.

[0147] Optionally, the first generation module 504 is further configured to perform alignment processing on the first alignment embedding and the text sentence embedding through the second alignment encoder to generate a second alignment embedding; and to perform encoding processing on the text sentence embedding through the fragment encoder to obtain a text fragment embedding.

[0148] Optionally, the second generation module 506 is further configured to determine a first text sentence embedding and a second text sentence embedding in the text sentence embedding based on the first alignment embedding, wherein the first text sentence embedding and the first alignment embedding have a semantic relationship; construct positive sentence sample pairs based on the first alignment embedding and the first text sentence embedding, and construct negative sentence sample pairs based on the first alignment embedding and the second text sentence embedding.

[0149] Optionally, the second generation module 506 is further configured to determine a first text fragment embedding and a second text fragment embedding in the text fragment embedding based on the second alignment embedding, wherein the first text fragment embedding and the second alignment embedding have a semantic association relationship; construct positive fragment sample pairs based on the second alignment embedding and the second text fragment embedding, and construct negative fragment sample pairs based on the second alignment embedding and the second text fragment embedding.

[0150] Optionally, the execution module 508 is further configured to perform the first training task for the pre-trained model using the sentence positive and negative sample pairs; and to perform the second training task for the pre-trained model using the segment positive and negative sample pairs.

[0151] Optionally, the execution module 508 is further configured to calculate the first sample pair similarity of the positive and negative sample pairs of the sentence by executing the first training task, and calculate the first loss value corresponding to the pre-trained model based on the first sample pair similarity; and tune the pre-trained model using the first loss value.

[0152] Optionally, the execution module 508 is further configured to calculate the second sample pair similarity of the positive and negative sample pairs of the fragment by executing the second training task, and calculate the second loss value corresponding to the pre-trained model based on the second sample pair similarity; and use the second loss value to tune the parameters of the pre-trained model.

[0153] Optionally, the input module 502 is further configured to extract text data and non-text data from the target data; and to encode the text data and non-text data to obtain the text word embeddings of the text data and the non-text embeddings corresponding to the non-text data.

[0154] The above is a schematic scheme of a training device for a fusion representation model of multimodal data according to this embodiment. It should be noted that the technical solution of this training device for a fusion representation model of multimodal data belongs to the same concept as the technical solution of the training method for a fusion representation model of multimodal data described above. For details not described in detail in the technical solution of the training device for a fusion representation model of multimodal data, please refer to the description of the technical solution of the training method for a fusion representation model of multimodal data described above.

[0155] See Figure 6 , Figure 6 A flowchart of a fusion representation method for multimodal data provided according to an embodiment of this specification is shown, which specifically includes the following steps.

[0156] Step 602: Determine the multimodal data and input the multimodal data into the fusion representation model, wherein the fusion representation model is trained using the above-described training method for the fusion representation model of multimodal data.

[0157] Step 604: Obtain the target fusion representation corresponding to the multimodal data output by the fusion representation model, wherein the target fusion representation is used to perform data processing tasks on the multimodal data.

[0158] In a specific embodiment of this specification, multimodal data can be understood as industry data containing multiple different modalities. For example, multimodal data in the government affairs field can be government documents containing image data and text data. In the field of image and text retrieval, multimodal data can be web articles containing image data, video data, and text data. After inputting multimodal data into the fusion representation model, the model can output the target fusion representation corresponding to the multimodal data. The target fusion representation can be understood as a unified representation of the text space that integrates non-textual modal information from the multimodal data. Subsequently, downstream tasks such as image and text retrieval, image and text generation, sentiment analysis, and other multimodal tasks can be performed based on the target fusion representation, providing reliable data support for downstream tasks and improving the execution accuracy of downstream tasks.

[0159] This specification provides a fusion representation method for multimodal data, which realizes the output of a fusion representation that incorporates non-textual modal information for multimodal data through a fusion representation model, bridging the semantic gap between different modal data and thus enhancing the interpretability of the final data representation.

[0160] Corresponding to the above method embodiments, this specification also provides embodiments of a fusion representation apparatus for multimodal data. Figure 7 A schematic diagram of a fusion representation apparatus for multimodal data provided in one embodiment of this specification is shown.

[0161] like Figure 7 As shown, the device includes:

[0162] The input module 702 is configured to determine multimodal data and input the multimodal data into the fusion representation model, wherein the fusion representation model is trained using the above-described training method for the fusion representation model of multimodal data.

[0163] The acquisition module 704 is configured to acquire a target fusion representation corresponding to the multimodal data output by the fusion representation model, wherein the target fusion representation is used to perform a data processing task for the multimodal data.

[0164] This specification provides a fusion representation device for multimodal data, which realizes the output of a fusion representation that incorporates non-textual modal information for multimodal data through a fusion representation model, bridging the semantic gap between different modal data and thus enhancing the interpretability of the final data representation.

[0165] The above is a schematic scheme of a fusion representation device for multimodal data according to this embodiment. It should be noted that the technical solution of this fusion representation device for multimodal data belongs to the same concept as the technical solution of the fusion representation method for multimodal data described above. For details not described in detail in the technical solution of the fusion representation device for multimodal data, please refer to the description of the technical solution of the fusion representation method for multimodal data described above.

[0166] See Figure 8 , Figure 8 A flowchart of a government data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0167] Step 802: Determine the target government data for the data analysis task, wherein the target government data includes government data in at least two modalities.

[0168] Step 804: Input the target government data into the fusion representation model to obtain the target fusion representation of the target government data output by the fusion representation model, wherein the fusion representation model is trained by the above-mentioned training method for the fusion representation model of multimodal data.

[0169] Step 806: Execute the data analysis task according to the target fusion representation to obtain the data analysis results corresponding to the target government data.

[0170] In a specific embodiment of this specification, a data analysis task can be understood as a multimodal data analysis task targeting government data. This task can include tasks such as summary generation and document translation. Determining the target government data for the data analysis task means identifying the target government data to be analyzed. This target government data can be documents such as service guides or policy notices. The target government data includes at least two modalities, such as text data in the text modality and image data in the image modality. After inputting the target government data into the fusion representation model, the model can output a target fusion representation corresponding to the multimodal data. This target fusion representation can be understood as a unified representation of the text space that integrates non-text modality information from the multimodal data. Data analysis is performed based on the target fusion representation to obtain the corresponding data analysis result. For example, if the data analysis task is summary generation, the data analysis result is the summary generated based on the target government data.

[0171] This specification provides a method for processing government data. It achieves this by using a fusion representation model to output a fused representation of target government data that incorporates non-textual modal information. This bridges the semantic gap between different modalities, thereby enhancing the interpretability of the final data representation. The target fused representation output by the model is then used to perform data analysis tasks, improving the accuracy of data analysis and enhancing the user experience.

[0172] Corresponding to the above method embodiments, this specification also provides embodiments of a government data processing device. Figure 9 A schematic diagram of the structure of a government data processing device according to one embodiment of this specification is shown. Figure 9 As shown, the device includes:

[0173] The determination module 902 is configured to determine the target government data for the data analysis task, wherein the target government data includes government data in at least two modalities.

[0174] The input module 904 is configured to input the target government data into the fusion representation model to obtain the target fusion representation of the target government data output by the fusion representation model, wherein the fusion representation model is trained by the above-mentioned training method for the fusion representation model of multimodal data.

[0175] The execution module 906 is configured to perform the data analysis task based on the target fusion representation and obtain the data analysis results corresponding to the target government data.

[0176] This specification provides a government data processing device that, through a fusion representation model, outputs a fused representation of target government data that incorporates non-textual modal information. This bridges the semantic gap between different modalities of data, thereby enhancing the interpretability of the final data representation. The target fused representation output by the model is used to perform data analysis tasks, improving the accuracy of data analysis and enhancing the user experience.

[0177] The above is an illustrative scheme of a government data processing device according to this embodiment. It should be noted that the technical solution of this government data processing device and the technical solution of the above-described government data processing method belong to the same concept. For details not described in detail in the technical solution of the government data processing device, please refer to the description of the technical solution of the above-described government data processing method.

[0178] Figure 10 A structural block diagram of a computing device 1000 according to one embodiment of this specification is shown. The components of the computing device 1000 include, but are not limited to, a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 via a bus 1030, and a database 1050 is used to store data.

[0179] The computing device 1000 also includes an access device 1040, which enables the computing device 1000 to communicate via one or more networks 1060. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1040 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0180] In one embodiment of this specification, the above-described components of the computing device 1000 and Figure 10 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 10 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0181] The computing device 1000 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1000 can also be a mobile or stationary server.

[0182] The processor 1020 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned training method for the fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method.

[0183] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solutions of the training method for the fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the training method for the fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method described above.

[0184] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described training method for a fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method.

[0185] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the training method for the fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method described above. Details not described in detail in the technical solution of the storage medium can be found in the descriptions of the technical solutions of the training method for the fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method described above.

[0186] An embodiment of this specification also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described training method for a fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method.

[0187] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the training method for the fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method described above. Details not described in detail in the technical solution of the computer program product can be found in the descriptions of the technical solutions of the training method for the fusion representation model of multimodal data, the fusion representation method for multimodal data, and the government data processing method described above.

[0188] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0189] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0190] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0191] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0192] The preferred embodiments disclosed above are merely illustrative of this specification. Optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described in this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification.

Claims

1. A training method for a fusion representation model of multimodal data, characterized in that, include: Determine the text word embeddings and non-text embeddings of the target data, and input the text word embeddings and non-text embeddings into a pre-trained model, the pre-trained model including a first encoder pair and a second encoder pair; The first encoder generates a first alignment embedding and a text sentence embedding based on the text word embedding and non-text embedding. The second encoder generates a second alignment embedding and a text fragment embedding based on the first alignment embedding and the text sentence embedding. The alignment dimensions of the first alignment embedding and the text sentence embedding are different from those of the second alignment embedding and the text fragment embedding. Sentence positive and negative sample pairs are generated based on the first alignment embedding and the text sentence embedding, and fragment positive and negative sample pairs are generated based on the second alignment embedding and the text fragment embedding; The step of generating positive and negative sentence sample pairs based on the first alignment embedding and the text sentence embedding includes: determining a first text sentence embedding and a second text sentence embedding in the text sentence embedding based on the first alignment embedding, wherein the first text sentence embedding and the first alignment embedding have a semantic relationship; constructing positive sentence sample pairs based on the first alignment embedding and the first text sentence embedding, and constructing negative sentence sample pairs based on the first alignment embedding and the second text sentence embedding; The step of generating positive and negative sample pairs of fragments based on the second alignment embedding and the text fragment embedding includes: determining a first text fragment embedding and a second text fragment embedding in the text fragment embedding based on the second alignment embedding, wherein the first text fragment embedding and the second alignment embedding have a semantic relationship; constructing positive sample pairs of fragments based on the second alignment embedding and the second text fragment embedding; and constructing negative sample pairs of fragments based on the second alignment embedding and the second text fragment embedding. By performing a fusion representation training task on the pre-trained model using the sentence positive and negative sample pairs and the segment positive and negative sample pairs, a fusion representation model is obtained.

2. The method according to claim 1, characterized in that, The first encoder pair includes a first alignment encoder and a sentence encoder; The first encoder generates a first aligned embedding and a text sentence embedding based on the text word embedding and non-text embedding, including: The first alignment encoder aligns the text word embeddings and the non-text embeddings to obtain the first aligned embedding; The sentence encoder encodes the text word embeddings to obtain the text sentence embeddings.

3. The method according to claim 1, characterized in that, The second encoder pair includes a second alignment encoder and a segment encoder; The second encoder generates a second alignment embedding and a text fragment embedding based on the first alignment embedding and the text sentence embedding, including: The second alignment encoder aligns the first alignment embedding and the text sentence embedding to generate the second alignment embedding. The text sentence embedding is encoded by the fragment encoder to obtain the text fragment embedding.

4. The method according to claim 1, characterized in that, The fusion representation training task includes a first training task and a second training task; The task of training a fusion representation for the pre-trained model is performed using the sentence positive and negative sample pairs and the segment positive and negative sample pairs, including: The first training task for the pre-trained model is performed using the positive and negative sample pairs of the sentences; The second training task for the pre-trained model is performed using the positive and negative sample pairs of the fragment.

5. The method according to claim 4, characterized in that, Performing the first training task for the pre-trained model using the positive and negative sample pairs of the sentences includes: By performing the first training task, the first sample pair similarity of the positive and negative sample pairs of the sentence is calculated, and the first loss value corresponding to the pre-trained model is calculated based on the first sample pair similarity. The pre-trained model is tuned using the first loss value.

6. The method according to claim 4, characterized in that, Performing the second training task for the pre-trained model using the positive and negative sample pairs of the fragments includes: By performing the second training task, the second sample pair similarity of the positive and negative sample pairs of the segment is calculated, and the second loss value corresponding to the pre-trained model is calculated based on the second sample pair similarity. The pre-trained model is then tuned using the second loss value.

7. The method according to any one of claims 1-6, characterized in that, Determine the text word embeddings and non-text embeddings of the target data, including: Extract the text data and non-text data from the target data; The text data and non-text data are encoded to obtain the text word embeddings of the text data and the non-text embeddings corresponding to the non-text data.

8. A fusion representation method for multimodal data, characterized in that, include: Determine multimodal data and input the multimodal data into a fusion representation model, wherein the fusion representation model is obtained by training using the method described in any one of claims 1-7; Obtain the target fusion representation corresponding to the multimodal data output by the fusion representation model, wherein the target fusion representation is used to perform data processing tasks on the multimodal data.

9. A method for processing government data, characterized in that, include: The target government data for the data analysis task is determined, wherein the target government data includes government data in at least two modalities; The target government data is input into the fusion representation model to obtain the target fusion representation corresponding to the target government data output by the fusion representation model, wherein the fusion representation model is trained by the method described in any one of claims 1-7; The data analysis task is performed based on the target fusion representation to obtain the data analysis results corresponding to the target government data.

10. A training device for a fusion representation model of multimodal data, characterized in that, include: An input module is configured to determine the text word embeddings and non-text embeddings of the target data, and input the text word embeddings and the non-text embeddings into a pre-trained model, the pre-trained model including a first encoder pair and a second encoder pair; The first generation module is configured to generate a first alignment embedding and a text sentence embedding based on the text word embedding and non-text embedding using the first encoder, and to generate a second alignment embedding and a text fragment embedding based on the first alignment embedding and the text sentence embedding using the second encoder, wherein the alignment dimensions of the first alignment embedding and the text sentence embedding are different from those of the second alignment embedding and the text fragment embedding. The second generation module is configured to generate sentence positive and negative sample pairs based on the first alignment embedding and the text sentence embedding, and to generate segment positive and negative sample pairs based on the second alignment embedding and the text segment embedding; The step of generating positive and negative sentence sample pairs based on the first alignment embedding and the text sentence embedding includes: determining a first text sentence embedding and a second text sentence embedding in the text sentence embedding based on the first alignment embedding, wherein the first text sentence embedding and the first alignment embedding have a semantic relationship; constructing positive sentence sample pairs based on the first alignment embedding and the first text sentence embedding, and constructing negative sentence sample pairs based on the first alignment embedding and the second text sentence embedding; The step of generating positive and negative sample pairs of fragments based on the second alignment embedding and the text fragment embedding includes: determining a first text fragment embedding and a second text fragment embedding in the text fragment embedding based on the second alignment embedding, wherein the first text fragment embedding and the second alignment embedding have a semantic relationship; constructing positive sample pairs of fragments based on the second alignment embedding and the second text fragment embedding; and constructing negative sample pairs of fragments based on the second alignment embedding and the second text fragment embedding. The execution module is configured to perform a fusion representation training task on the pre-trained model using the sentence positive and negative sample pairs and the segment positive and negative sample pairs, to obtain the fusion representation model.

11. A fusion representation device for multimodal data, characterized in that, include: An input module is configured to determine multimodal data and input the multimodal data into a fusion representation model, wherein the fusion representation model is obtained by training using the method described in any one of claims 1-7; The acquisition module is configured to acquire a target fusion representation corresponding to the multimodal data output by the fusion representation model, wherein the target fusion representation is used to perform data processing tasks on the multimodal data.

12. A government data processing device, characterized in that, include: The determination module is configured to determine the target government data for the data analysis task, wherein the target government data includes government data in at least two modalities; The input module is configured to input the target government data into the fusion representation model to obtain the target fusion representation of the target government data output by the fusion representation model, wherein the fusion representation model is trained by the method described in any one of claims 1-7; The execution module is configured to perform the data analysis task based on the target fusion representation and obtain the data analysis results corresponding to the target government data.

13. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9.

14. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.

15. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-modal pre-training model training method and device and multi-modal data processing method and device

    CN116861995A

  • Multi-language speech neural machine translation method based on multi-modal comparative learning

    CN117494730A