Task execution method and system based on multi-modal information
By constructing a model collaborative transformer and a modal perception layer, the problems of inaccurate and complex multimodal information processing in the prior art are solved, and more accurate and efficient multimodal information fusion and task execution are achieved.
Patent Information
- Application Number
- CN202411937268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to capture the differences in different mode data when processing multimodal information, resulting in inaccurate output results, and complex processing flow and inefficient efficiency.
By constructing a model collaborative transformer, multiple modal embedding layers are used to extract multimodal embedding vectors, and the dynamic weight vectors for the modal perception layer to obtain multimodal information are added, and the output of the multi-head self-attention layer in the encoder is adjusted to realize the fusion of multimodal information and task execution.
It improves the understanding and description of complex scenarios, improves the accuracy of task execution results, and realizes accurate and efficient multimodal information extraction to adapt to the existence or absence of different modal information.
Smart Images

Figure CN120030485A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a task execution method and system based on multimodal information. Background Art
[0002] With the iteration and update of large language models, their functions are becoming more and more general and intelligent, with stronger question-answering and text generation capabilities, and they are widely used in many fields. How to effectively process and integrate multimodal field data so that large language models can perform better in this field and improve the accuracy of intelligent question-answering is an important issue currently faced.
[0003] The existing multimodal information processing methods mainly focus on text and speech, but rarely involve images. At the same time, due to the differences between different types of modal data, the existing models that use multi-class modal data for analysis are difficult to capture the differences between different modal data, resulting in inaccurate final output results.
[0004] In addition, different models are usually used to process different modal data, which makes the processing flow complicated and inefficient, and ignores the relationship between multimodal information, resulting in inaccurate task execution results. Summary of the invention
[0005] In view of the above analysis, an embodiment of the present invention aims to provide a task execution method and system based on multimodal information, so as to solve the problem that the existing multimodal information mining is insufficient and the processing is complex, resulting in inaccurate task execution results.
[0006] On the one hand, an embodiment of the present invention provides a task execution method based on multimodal information, comprising the following steps:
[0007] Collect and preprocess historical multimodal information to construct a multimodal sample set; each sample includes one modal information or multiple different types of but content-related modal information;
[0008] Construct a modal co-transformer. The modal co-transformer constructs multiple modal embedding layers in the Transformer model to extract multimodal embedding vectors and pass them to the encoder. It also adds a modal perception layer to obtain dynamic weight vectors of multimodal information to adjust the output of the multi-head self-attention layer in the encoder.
[0009] Using multimodal sample sets, the modal cooperative transformer is trained through multi-task self-supervised learning.
[0010] The trained modal cooperative transformer is used to fuse the multimodal information input in the actual task and output the task execution result.
[0011] Based on the further improvement of the above method, multiple modal embedding layers are constructed to extract multimodal embedding vectors, including:
[0012] Each type of modal information is input into its respective modal embedding layer. After the embedding vector is extracted, it is added to the modal identification vector corresponding to a modality type and then the positional encoding is added. The positional encoding is generated by adding a trainable offset to the position index and then using sine and cosine functions.
[0013] Based on the further improvement of the above method, the modal perception layer includes a modal quality assessment module, a splicing module and a softmax output module which are connected in sequence; the modal quality assessment module is used to perform quality assessment on the input multimodal information and then calculate a reliability score; the splicing module is used to obtain a modal perception vector according to the type and reliability score of the input multimodal information; the softmax output module uses the softmax function to convert the modal perception vector into a dynamic weight vector.
[0014] Based on the further improvement of the above method, the modal perception vector is obtained by splicing the perception values of each modality in a preset order, and the perception value of each modality is obtained by multiplying the existence flag of each modality by its reliability score, wherein the modal existence flag is set to 1 or 0 according to the presence or absence of modal information of this type.
[0015] Based on the further improvement of the above method, the dynamic weight vector adjusts the output of the multi-head self-attention layer in the encoder using the following formula:
[0016]
[0017] Among them, A j represents the output of the self-attention layer of the jth head, W modality represents the dynamic weight vector, Q, K, and V represent the query matrix, key matrix, and value matrix in the self-attention layer, d k represents the dimension of each row of key vector in the key matrix; ⊙ represents the Hadamard product; T represents the transpose operation.
[0018] Based on the further improvement of the above method, the multiple tasks include: mask collaborative prediction task, modal contrast learning task and modal collaborative transformer learning task; the mask collaborative prediction task is to mask the content of the multimodal information in each sample and then predict the masked content using the modal collaborative transformer; the modal contrast learning task is to construct positive and negative data pairs according to the type and content of the modal information, and use the modal collaborative transformer to maximize the similarity of the feature representation of the positive data pairs and minimize the similarity of the feature representation of the negative data pairs; the modal collaborative transformer learning task is an actual learning task set according to the application scenario.
[0019] Based on the further improvement of the above method, the positive data pairs in the modal contrastive learning task are data pairs consisting of different types of modal information in the same sample; the negative data pairs are data pairs consisting of different types of modal information in different samples; the feature representation of the positive data pairs and the features of the negative data pairs are both features extracted by the encoder.
[0020] Based on the further improvement of the above method, the loss function of the modal cooperative transformer is obtained by calculating the loss function of each task and weighted summing them.
[0021] Based on the further improvement of the above method, the multimodal information includes: text, speech and image; the reliability score of the text is obtained by calculating the scores of text length, grammatical correctness and professional terminology matching respectively and then weighting them; the reliability score of the speech is obtained by calculating the signal-to-noise ratio and clarity of the speech respectively and then weighting them; the reliability score of the image is obtained by calculating the edge clarity, contrast and noise level of the image respectively and then weighting them.
[0022] On the other hand, an embodiment of the present invention provides a task execution system based on multimodal information, including:
[0023] The sample construction module is used to collect and preprocess historical multimodal information and construct a multimodal sample set; each sample includes one modal information or multiple different types of but content-related modal information;
[0024] The model building module is used to build a modal co-transformer. The modal co-transformer constructs multiple modal embedding layers in the Transformer model to extract multi-modal embedding vectors and pass them to the encoder. It also adds a modal perception layer to obtain dynamic weight vectors of multi-modal information to adjust the output of the multi-head self-attention layer in the encoder.
[0025] The multi-task training module uses a multi-modal sample set to train the modal cooperative transformer through multi-task self-supervised learning;
[0026] The task execution module is used to use the trained modal cooperative transformer to fuse the multimodal information input in the actual task and output the task execution result.
[0027] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0028] 1. Construct a unified modal collaborative transformer, which converts different modal information into a unified embedding vector through multiple modal embedding layers, making it easier to process and understand multiple types of information at the same time. By integrating multimodal information, it is helpful to capture richer features and contextual relationships, thereby improving its ability to understand and describe complex scenes and improve the accuracy of task execution results.
[0029] 2. Generate dynamic weight vectors through the modal perception layer to flexibly adjust the degree of attention to different modal information according to the quality of multimodal information, better capture the intrinsic connection and interaction between multimodal information when fusing multimodal information, and achieve accurate and efficient multimodal information extraction; moreover, when faced with the presence or absence of different modal information, the modal perception layer helps the modal cooperative transformer make reasonable adjustments, so that it can maintain good performance in various situations.
[0030] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. In the entire drawings, the same reference symbols represent the same components;
[0032] Figure 1 This is a flow chart of a task execution method based on multimodal information in Example 1 of the present invention. DETAILED DESCRIPTION
[0033] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0034] Example 1
[0035] A specific embodiment of the present invention discloses a task execution method based on multimodal information, such as Figure 1 As shown, the following steps are included:
[0036] S1. Collect and preprocess historical multimodal information to construct a multimodal sample set; each sample includes one modal information or multiple different types of but content-related modal information.
[0037] It should be noted that the multimodal information in this embodiment includes: text, voice and image. For example, in the field of medical self-service diagnosis, text information includes: electronic medical records, symptoms described by patients, medical literature, etc.; voice information includes: patient voice messages, doctor's voice records; image information includes: photos of patients' own symptoms uploaded by patients, medical images, etc.
[0038] Furthermore, preprocess the historical multimodal information. Among them, the preprocessing of text includes: removing noise and irrelevant information in the text, such as HTML tags, special characters, numbers, spaces, line breaks, etc.; removing stop words, such as "de", "shi", etc.; using the jieba word segmentation library to split the text into words or phrases.
[0039] The preprocessing of speech includes: calculating the signal-to-noise ratio of the speech. If the signal-to-noise ratio is within the preset signal-to-noise ratio threshold range, spectral subtraction is used for denoising. If the signal-to-noise ratio is less than the minimum signal-to-noise ratio threshold, after preliminary denoising using spectral subtraction, adaptive filtering is used for fine denoising; identifying the start and end positions of the speech and removing non-speech parts; performing framing, pre-emphasis, and windowing on the remaining speech information to extract the Mel spectrogram.
[0040] The preprocessing of images includes: scaling the images to a unified size; using filters to remove noise in the images and smooth the image details.
[0041] Preferably, to increase the robustness and generalization of data processing, data augmentation is performed on the multimodal information, including: identifying entities in the text information, obtaining relevant entities and their attributes from the knowledge spectrogram through similarity to enhance the semantics of the text information; randomly selecting a certain proportion of words and replacing them with the special token [MASK] to generate new text information; rotating, translating, and flipping the images to generate new image information; dividing the images into several blocks, randomly covering some image blocks, and generating new image information by filling with gray or adding noise; using the Mixup method for the speech information to generate new speech signals, and performing frequency or time masking on the Mel spectrogram of the speech information.
[0042] Considering that historical data does not simultaneously contain multimodal information related to content, therefore, when constructing the multimodal sample set, each sample includes one type of modal information or multiple different types of but content-related modal information.
[0043] S2. Construct a modal collaborative transformer. The modal collaborative transformer constructs multiple modal embedding layers in the Transformer model to extract multimodal embedding vectors and pass them to the encoder, and adds a modal perception layer to obtain the dynamic weight vector of the multimodal information to adjust the output of the multi-head self-attention layer in the encoder.
[0044] It should be noted that the existing large language model is usually composed of multiple Transformer models superimposed, and the modal cooperative transformer of this embodiment is obtained by improving the Transformer model in the large language model. The existing Transformer model usually includes four parts: an embedding module, multiple encoders, multiple decoders, and an output module. This embodiment constructs multiple modal embedding layers in the embedding module to extract multimodal embedding vectors, and adds a modal perception layer to obtain the dynamic weight vector of multimodal information to adjust the output of the multi-head self-attention layer in the encoder.
[0045] Specifically, multiple modal embedding layers are constructed to extract multimodal embedding vectors, including:
[0046] Each type of modal information is input into its respective modal embedding layer. After the embedding vector is extracted, it is added to the modal identification vector corresponding to a modality type and then the positional encoding is added. The positional encoding is generated by adding a trainable offset to the position index and then using sine and cosine functions.
[0047] It should be noted that the modal embedding layer includes a text embedding layer, a speech embedding layer and an image embedding layer, and each modal embedding layer includes: linear transformation, adding a modal identification vector and adding a position code.
[0048] Specifically, the linear transformation of the text embedding layer is to map each word in the text into a vector space of a fixed size, including: converting each word in the input text into a corresponding vocabulary index, where the vocabulary index is an integer representing the position of the word in the vocabulary table; then searching the corresponding embedding vector from the word embedding matrix according to the vocabulary index to achieve linear transformation, thereby converting the vocabulary index in the text into a word vector of fixed dimension to obtain a text embedding vector; wherein the word embedding matrix is initially randomly initialized and is continuously updated and optimized during the training process; the added modality identification vector of the text embedding layer is to add the text embedding vector output by the linear transformation to the modality identification vector corresponding to the text, so that the model can perceive the modality type to which the data belongs, and then add the position encoding, which is generated by adding a trainable offset to the position index and then using sine and cosine functions. This method can better handle the position information in the sequence by learning the relative position.
[0049] It should be noted that the linear transformation of the speech embedding layer is to map the MFCC features of the speech signal into the vector space, including: extracting the MFCC features from the Mel spectrum of the speech information, normalizing it and then linearly transforming it to obtain the speech embedding vector; adding the modal identification vector of the speech embedding layer is to add the speech embedding vector output by the linear transformation to the modal identification vector corresponding to the speech, and then adding the speech position coding using the same method as the text position coding.
[0050] Preferably, before the linear transformation, a convolutional neural network is first used to extract a speech feature tensor from the MFCC features, and then the speech feature tensor is linearly transformed to facilitate capturing local patterns and time dependencies in the MFCC features.
[0051] It should be noted that the linear transformation of the image embedding layer is to map the pixel data of the image into the vector space, including: dividing the image into multiple non-overlapping blocks of fixed size, flattening the pixel value of each image block into a one-dimensional vector, and obtaining the image embedding vector through linear transformation; the added modality identification vector of the image embedding layer is to add the image embedding vector output by the linear transformation to the modality identification vector corresponding to the image, and then add the image position code using the same method as the text position code.
[0052] It should be noted that the modal perception layer added in this embodiment is used to perform quality assessment on the input multimodal information setting to obtain a reliability score, and then obtain a dynamic weight vector of the multimodal information.
[0053] Specifically, the modal perception layer includes a modal quality assessment module, a splicing module and a softmax output module which are connected in sequence; the modal quality assessment module is used to perform quality assessment on the input multimodal information and then calculate a reliability score in the range of (0,1); the splicing module is used to splice the input multimodal information into a modal perception vector according to the type and reliability score of the input multimodal information; the softmax output module uses the softmax function to convert the modal perception vector into a dynamic weight vector.
[0054] Among them, the reliability score R of the text in the modality quality assessment module is text It is obtained by calculating the scores of text length, grammatical correctness and professional term matching and then weighting them. The formula is as follows:
[0055] R text =w L ·exp(-α|LL 0 |)+w g ·(1-Eg)+w t ·T m
[0056] Among them, w L ,w g ,w t L and L respectively represent the weight coefficients of text length, grammatical correctness, and professional term matching. 0 They represent the actual length of the text and the length threshold, respectively. α represents the preset scaling factor used to adjust the impact of the text length difference. E g represents the grammatical error rate, T m Indicates the degree of professional term matching.
[0057] The reliability score of the speech speech It is obtained by calculating the signal-to-noise ratio and clarity score of the speech separately and then weighting them. The formula is as follows:
[0058]
[0059] Among them, w snr ,w c Represent the weight coefficients of signal-to-noise ratio and clarity respectively, represents the speech signal-to-noise ratio normalized to the interval [0, 1], C represents the clarity value, and in this embodiment, the speech clarity is represented by calculating the STOI (Short-Time ObjectiveIntelligibility) score.
[0060] The reliability score R of the image image It is obtained by calculating the edge clarity, contrast and noise level of the image and then weighting them. The formula is as follows:
[0061] R image =w cl Clarity+w co Contrast+w nl (1-NoiseLevel)
[0062] Among them, w cl ,w co ,w nl They represent the weight coefficients of edge clarity, contrast and noise level of the image, respectively. Clarity, Contrast and NoiseLevel represent the normalized edge clarity, contrast and noise level, respectively. In this example, edge clarity is obtained by using the Sobel operator to calculate the gradient amplitude of all edge points in the image, contrast is obtained by calculating the histogram statistics of the image, and noise level is measured by filtering or noise estimation algorithm.
[0063] Furthermore, in the splicing module, the modal perception vector is obtained by splicing the perception values of each modality in a preset order, and the perception value of each modality is obtained by multiplying the existence flag of each modality by its reliability score, wherein the modal existence flag is set to 1 or 0 according to the presence or absence of modal information of this type.
[0064] Use M i Indicates the modality existence flag of the i-th modality information, M i =1 indicates the existence of the i-th modal information, M i = 0 means that there is no i-th modal information; R irepresents the reliability score of the i-th modality, then the modality perception vector is expressed as V modality =[M 1 R 1 ,M 2 R 2 ,...,M i R i ,...,M n R n ], n represents the number of modal types.
[0065] Furthermore, in the softmax output module, the softmax function is used to obtain the dynamic weight vector W modality =softmax(V modality ), ensuring that the sum of the modal weights is 1.
[0066] The multimodal embedding vectors output by multiple modal embedding layers are concatenated and passed to the encoder. The length of the concatenated feature vector when all modal information exists is fixed. When a certain modal information is missing, it is padded with 0 or a specific padding value. When calculating self-attention in the decoder, the padded part is masked so that these padded parts do not participate in the calculation. The attention score of the padded part is infinitely close to 0 after softmax.
[0067] It should be noted that the encoder consists of a two-layer connection structure: the first layer is a multi-head self-attention layer, and the second layer is a feedforward neural network layer; each layer is followed by a residual connection and a layer normalization.
[0068] There are three weight matrices in the multi-head self-attention layer: query weight matrix, key weight matrix and value weight matrix. These weight matrices are the parameters learned during model training. The multi-head self-attention layer multiplies the input multimodal embedding vector with the three weight matrices respectively to obtain the query matrix Q, key matrix K and value matrix V corresponding to each head.
[0069] Furthermore, the modality perception layer obtains the dynamic weight vector of multimodal information to adjust the output of the multi-head self-attention layer in the encoder, including:
[0070] For each self-attention layer, the query matrix is multiplied by the transpose of the key matrix to obtain an attention score matrix, which is used to represent the attention intensity of each element in the input to other elements. In order to obtain a stable gradient, the score is scaled by dividing by the square root of the key vector dimension. Finally, the modal perception layer is used to obtain the dynamic weight vector of multimodal information to weight the scaled score, and then the softmax function is applied to convert the weighted score into a probability distribution to obtain the adjusted attention weight; the value matrix is multiplied by the adjusted attention weight, and the weighted sum result is the adjusted output of each self-attention layer. The formula is as follows:
[0071]
[0072] Among them, A j represents the output of the self-attention layer of the jth head, W modality represents the dynamic weight vector, Q, K, and V represent the query matrix, key matrix, and value matrix in the self-attention layer, d k represents the dimension of each row of key vector in the key matrix; ⊙ represents the Hadamard product, and T represents the transpose operation.
[0073] The outputs of the multi-head self-attention layer are concatenated and mapped through a linear layer to integrate the information learned by different heads to obtain the final output of the multi-head self-attention layer. The output of the multi-head self-attention layer is residually connected to the input and then normalized. The result of the layer normalization is input into the feedforward neural network layer for nonlinear transformation. The output of the feedforward neural network layer is residually connected and layer normalized with the input before the multi-head self-attention layer to obtain the output of the encoder.
[0074] It should be noted that the decoder consists of a three-layer connection structure: the first layer is a masked multi-head self-attention sublayer, the second layer is the encoder-decoder attention layer, and the third layer is a feedforward neural network layer; each layer is followed by a residual connection and a layer normalization.
[0075] The output module after the decoder usually sets different structures according to the needs of the actual task to meet the different types of outputs required by different tasks, such as classification tasks, text generation tasks, and named entity recognition tasks. For classification tasks, the output module maps the output of the decoder to the dimension of the number of categories through a linear layer, and then converts the output to a probability distribution through a softmax function for classification; for text generation tasks, the output module maps to the size of the vocabulary through a linear layer, and then generates the probability distribution of the next word through a softmax function, and selects the next word based on the probability distribution; for named entity recognition tasks, the output module maps to the dimension of the number of labels through a linear layer, and then performs entity annotation through a softmax function or other classification functions (such as a CRF layer).
[0076] S3. Use multimodal sample sets to train the modal collaborative transformer through multi-task self-supervised learning.
[0077] It should be noted that the modality co-transformer is trained through a multi-task self-supervised learning method, including: mask co-prediction task, modality contrast learning task and modality co-transformer learning task.
[0078] Specifically, the masked co-prediction task is to mask the multimodal information in each sample and use the modal co-transformer to predict the masked content, which helps to improve the tolerance of the modal co-transformer to noise and incomplete data, promote the ability to understand the context, and is conducive to learning richer and more general feature representations, so that it can better generalize when facing new and unseen tasks. The loss function L of the masked co-prediction task is mask Using a multi-category cross entropy loss function, the modal collaborative transformer generates a probability distribution for each masked position, and uses the multi-category cross entropy loss function to measure the difference between the probability distribution of the generated token and the probability distribution of the real token, so that the generated mask content is close to the real content.
[0079] The modality contrast learning task is to construct positive and negative data pairs according to the type and content of modal information, and use the modality co-transformer to maximize the similarity of the feature representation of the positive data pair and minimize the similarity of the feature representation of the negative data pair, so as to capture the differences between different modalities and improve the robustness of the modality co-transformer to different changes. Among them, the positive data pairs in the modality contrast learning task are data pairs composed of different types of modal information in the same sample, that is, different modal information with related content constitutes positive data pairs; negative data pairs are data pairs composed of different types of modal information in different samples, that is, different modal information with different content constitutes negative data pairs; the feature representation of the positive data pair and the feature representation of the negative data pair are both features extracted by the encoder.
[0080] The loss function L for the modality contrastive learning task modal Calculated by the following formula:
[0081]
[0082] Where N represents the number of modal information, represents the feature representation of positive data pairs, h p Represents the feature representation of the p-th modal information; Represents the feature representation of the modal information that forms a positive data pair with the p-th modal information; represents the feature representation of negative data pairs, represents the feature representation of the rth modal information that forms a negative data pair with the pth modal information, Nr represents the number of modal information that forms a negative data pair with the pth modal information; sim(·,·) represents the similarity function, and this embodiment adopts cosine similarity; τ represents a conditional parameter, which is used to control the scale of similarity to avoid gradient vanishing or excessive gradient.
[0083] The modal cooperative transformer learning task is an actual learning task set according to the application scenario, and the corresponding loss function L is used according to the learning goal. task .
[0084] For example, in the scenario of medical self-diagnosis, the modal co-transformer is used to extract multimodal information and then output the response text. The corresponding modal co-transformer learning task is the text generation task, and the corresponding loss function adopts the negative log-likelihood loss function to maximize the likelihood of the entire sequence, so that the model can generate coherent and reasonable text.
[0085] Finally, the loss function L of the modal cooperative transformer is total It is obtained by calculating the loss function of each task and weighted summing them up. The formula is as follows:
[0086] L total =λ 1 L mask +λ 2 L modal +λ 3 L task
[0087] Among them, λ 1 ,λ 2 ,λ 3 is the weight coefficient of the loss function of the corresponding task, which is adjusted according to the task requirements.
[0088] During the training process of the modal cooperative transformer, as the number of training times increases, the model parameters are continuously updated according to the loss function, and a trained modal cooperative transformer is obtained at the end of the training.
[0089] S4. Use the trained modal cooperative transformer to fuse the multimodal information input in the actual task and output the task execution result.
[0090] During implementation, multimodal information is received, pre-processed and passed into the trained modal cooperative transformer, which extracts information from the multimodal information and outputs the task execution results.
[0091] For example, in the scenario of medical self-diagnosis, the patient inputs a text-type symptom description through the online medical diagnosis assistance system, leaves more symptom details through voice, uploads a tongue coating photo, pre-processes the input multimodal information, and searches for relevant information in combination with the medical knowledge graph to enhance the text semantics. The multimodal information is then passed into the modal collaborative transformer, and the modal perception layer performs a quality assessment on the multimodal information to generate a dynamic weight vector. The dynamic weight vector is used in the encoder to adjust the multi-head self-attention, extract and fuse the deep features of text, voice and image, and generate a preliminary diagnosis report on the patient's symptoms through the decoder and output module. The method of this embodiment can also be applied in intelligent customer interaction scenarios to achieve more intelligent customer service by understanding the user's voice, text and expression; it can be applied in security monitoring scenarios to improve the intelligence level of security monitoring by combining video, audio and sensor data.
[0092] Compared with the prior art, the present embodiment provides a task execution method based on multimodal information, constructs a unified modal cooperative transformer, and converts different modal information into a unified embedding vector through multiple modal embedding layers, which is convenient for simultaneously processing and understanding multiple types of information; by integrating multimodal information, it is beneficial to capture richer features and contextual relationships, thereby improving its understanding and description capabilities of complex scenes and improving the accuracy of task execution results; a dynamic weight vector is generated through the modal perception layer, so that the degree of attention to different modal information can be flexibly adjusted according to the quality of multimodal information, and the intrinsic connection and interaction between multimodal information can be better captured when fusing multimodal information, thereby achieving accurate and efficient multimodal information extraction; moreover, when faced with the presence or absence of different modal information, the modal perception layer is used to help the modal cooperative transformer make reasonable adjustments, so that it can maintain good performance in various situations.
[0093] Example 2
[0094] Another embodiment of the present invention discloses a task execution system based on multimodal information, thereby implementing a task execution method based on multimodal information in embodiment 1. The specific implementation of each module refers to the corresponding description in embodiment 1. The system includes:
[0095] The sample construction module is used to collect and preprocess historical multimodal information and construct a multimodal sample set; each sample includes one modal information or multiple different types of but content-related modal information;
[0096] The model building module is used to build a modal co-transformer. The modal co-transformer constructs multiple modal embedding layers in the Transformer model to extract multi-modal embedding vectors and pass them to the encoder. It also adds a modal perception layer to obtain dynamic weight vectors of multi-modal information to adjust the output of the multi-head self-attention layer in the encoder.
[0097] The multi-task training module uses a multi-modal sample set to train the modal cooperative transformer through multi-task self-supervised learning;
[0098] The task execution module is used to use the trained modal cooperative transformer to fuse the multimodal information input in the actual task and output the task execution result.
[0099] Since the present embodiment and the aforementioned method for executing a task based on multimodal information can be mutually referenced, the description here is repeated, so it will not be repeated here. Since the present system embodiment and the aforementioned method embodiment have the same principle, the present system embodiment also has the corresponding technical effects of the aforementioned method embodiment.
[0100] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0101] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A task execution method based on multimodal information, characterized in that: The following steps are involved: Collect and preprocess historical multimodal information to construct a multimodal sample set; each sample includes one modal information or multiple different types of but content-related modal information; Construct a modal cooperative transformer, which constructs multiple modal embedding layers in the Transformer model to extract multimodal embedding vectors and pass them to the encoder, and adds a modal perception layer to obtain a dynamic weight vector of multimodal information to adjust the output of the multi-head self-attention layer in the encoder; Using the multimodal sample set, the modal cooperative transformer is trained by a multi-task self-supervised learning method; The trained modal cooperative transformer is used to fuse the multimodal information input in the actual task and output the task execution result.
2. The task execution method based on multimodal information according to claim 1, characterized in that: The step of constructing multiple modal embedding layers to extract multimodal embedding vectors includes: Each type of modal information is input into its respective modal embedding layer. After the embedding vector is extracted, it is added to the modal identification vector corresponding to a modality type and then the position code is added; the position code is generated by adding a trainable offset to the position index and then using sine and cosine functions.
3. The task execution method based on multimodal information according to claim 1, characterized in that: The modal perception layer includes a modal quality assessment module, a splicing module and a softmax output module connected in sequence; the modal quality assessment module is used to perform quality assessment on the input multimodal information respectively and then calculate a reliability score; the splicing module is used to obtain a modal perception vector according to the type and reliability score of the input multimodal information; the softmax output module uses a softmax function to convert the modal perception vector into a dynamic weight vector.
4. The task execution method based on multimodal information according to claim 3 is characterized in that: The modal perception vector is obtained by splicing the perception values of each modality in a preset order, and the perception value of each modality is obtained by multiplying the existence flag of each modality by its reliability score, wherein the modal existence flag is set to 1 or 0 according to the presence or absence of modal information of this type.
5. The task execution method based on multimodal information according to claim 1, characterized in that: The dynamic weight vector adjusts the output of the multi-head self-attention layer in the encoder using the following formula: Among them, A j represents the output of the self-attention layer of the jth head, W modality represents the dynamic weight vector, Q, K, and V represent the query matrix, key matrix, and value matrix in the self-attention layer, d k represents the dimension of each row of key vector in the key matrix; ⊙ represents the Hadamard product; T represents the transpose operation.
6. The task execution method based on multimodal information according to claim 1, characterized in that: The multiple tasks include: a mask collaborative prediction task, a modal contrast learning task and a modal collaborative transformer learning task; the mask collaborative prediction task is to mask the content of the multimodal information in each sample and then use the modal collaborative transformer to predict the masked content; the modal contrast learning task is to construct positive and negative data pairs according to the type and content of the modal information, and use the modal collaborative transformer to maximize the similarity of the feature representation of the positive data pair and minimize the similarity of the feature representation of the negative data pair; the modal collaborative transformer learning task is an actual learning task set according to the application scenario.
7. The task execution method based on multimodal information according to claim 6, characterized in that: The positive data pairs in the modality contrast learning task are data pairs consisting of different types of modal information in the same sample; the negative data pairs are data pairs consisting of different types of modal information in different samples; the feature representation of the positive data pairs and the features of the negative data pairs are both features extracted using an encoder.
8. The task execution method based on multimodal information according to claim 6, characterized in that: The loss function of the modal cooperative transformer is obtained by calculating the loss function of each task and weighted summing them.
9. The task execution method based on multimodal information according to claim 3, characterized in that: The multimodal information includes: text, voice and image; the reliability score of the text is obtained by respectively calculating the scores of text length, grammatical correctness and professional terminology matching and then weighting them; the reliability score of the voice is obtained by respectively calculating the scores of the signal-to-noise ratio and clarity of the voice and then weighting them; the reliability score of the image is obtained by respectively calculating the edge clarity, contrast and noise level of the image and then weighting them.
10. A task execution system based on multimodal information, characterized in that: include: The sample construction module is used to collect and preprocess historical multimodal information and construct a multimodal sample set; each sample includes one modal information or multiple different types of but content-related modal information; A model building module is used to build a modal cooperative transformer, which constructs multiple modal embedding layers in the Transformer model to extract multimodal embedding vectors and pass them to the encoder, and adds a modal perception layer to obtain a dynamic weight vector of multimodal information to adjust the output of the multi-head self-attention layer in the encoder; A multi-task training module, using the multi-modal sample set to train the modal cooperative transformer through a multi-task self-supervised learning method; The task execution module is used to use the trained modal cooperative transformer to fuse the multimodal information input in the actual task and output the task execution result.