Multi-modal feature fusion method and device, electronic equipment and storage medium

By judging the modal type and using flexible computing paths to generate multimodal feature fusion, the problem of ignoring single-modal personalized features in the existing technology is solved, the efficiency and accuracy of multimodal feature fusion is improved, the interference between modes is reduced, and information interaction is optimized.

CN120354338APending Publication Date: 2025-07-22SHENZHEN YIWANKE DATA EQUIP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510365777.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art tends to ignore the personalized feature information in the single mode in multimodal feature fusion, resulting in the lack of sufficient expression ability of multimodal fusion features.

Method used

By judging whether the modal types of the reference modal data and the associated modal data in the modal data group are the same, different calculation paths are used to generate the first fusion feature, flexibly adjust the interaction intensity between the modal data groups, retain personalized information in the reference modal data, and avoid unnecessary calculations and information flow.

Benefits of technology

The efficiency and accuracy of multimodal feature fusion are improved, interference between modes is reduced, information interaction is optimized, and information fusion effect is enhanced between different modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354338A_ABST
    Figure CN120354338A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of feature extraction, and discloses a multi-modal feature fusion method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a plurality of modal data sets of target information, carrying out the feature extraction of each modal data set, obtaining a modal fusion feature, and storing the modal fusion feature in a database; whether the modal types of the reference modal data and the associated modal data in each modal data set are the same is judged, if yes, the modal fusion feature and the reference modal data are fused to generate a first fusion feature, and if not, the modal fusion feature is determined as the first fusion feature to generate a second fusion feature; and splicing the first fusion features corresponding to the modal data groups with the same reference modal data to generate second fusion features corresponding to the reference modal data, and determining multi-modal fusion features corresponding to the target information according to the second fusion features. According to the mode, not only can the correlation information among the multiple modal data be extracted, but also the personalized characteristics of the reference modal data can be reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of feature extraction, and in particular, to a multi-modal feature fusion method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] With the rapid development of big data and storage technologies, the data stored in storage servers usually contains rich user information. The data it stores is formed on the basis of text and develops towards the combination of multiple modalities such as text, images, and audio. Comprehensive analysis of multi-modal information such as text, audio, and images can analyze and understand the changing needs of users more comprehensively and accurately.

[0003] Currently, cross-modal features between multiple modality information are mainly extracted through pre-trained feature extraction models (for example, Transformer models, attention mechanisms, etc.). However, when extracting cross-modal features, it is easy to ignore the individual feature information within a single modality, resulting in the multi-modal fusion features obtained lacking sufficient expressive power. Summary of the Invention

[0004] In view of the above problems, the embodiments of the present application provide a multi-modal feature fusion method, a multi-modal feature fusion apparatus, an electronic device, and a computer-readable storage medium, which are used to solve the problem in the prior art that it is easy to ignore the personalized feature information within a single modality, resulting in the multi-modal fusion features obtained lacking sufficient expressive power.

[0005] According to one aspect of the embodiments of the present application, a multi-modal feature fusion method is provided. The method includes: obtaining multiple groups of modality data groups of target information, where the modality data group includes reference modality data and associated modality data; respectively performing feature extraction on each group of modality data groups to obtain the modality fusion features corresponding to each modality data group; respectively determining whether the modality types of the reference modality data and the associated modality data in each modality data group are the same; if they are the same, fusing the modality fusion feature corresponding to the modality data group and the reference modality data to generate the first fusion feature corresponding to the modality data group; otherwise, determining the modality fusion feature corresponding to the modality data group as the first fusion feature corresponding to the modality data group; splicing the first fusion features corresponding to the modality data groups with the same reference modality data to generate the second fusion feature corresponding to the reference modality data; and determining the multi-modal fusion feature corresponding to the target information according to the second fusion features of each reference modality data.

[0006] In an alternative approach, feature extraction is performed on each group of modal data groups respectively to obtain the modal fusion features corresponding to each modal data group, which specifically includes: for each modal data group, the following operations are performed: extracting the dependency information between the reference modal data and the associated modal data to obtain the global feature corresponding to the modal data group; extracting the local feature and the deep feature of the modal data group to obtain the modal feature corresponding to the modal data group; splicing the global feature and the modal feature corresponding to the modal data group to generate the modal fusion feature corresponding to the modal data group.

[0007] In an alternative approach, extracting the dependency information between the reference modal data and the associated modal data to obtain the global feature corresponding to the modal data group specifically includes: determining the reference modal data as the query, and respectively determining the associated modal data as the key and the value; processing the query, the key, and the value through an attention unit to obtain the global feature corresponding to the modal data group.

[0008] In an alternative approach, extracting the local feature and the deep feature of the modal data group to obtain the modal feature corresponding to the modal data group specifically includes: performing regularization processing on the reference modal data and the associated modal data; processing the regularized reference modal data and the associated modal data through a multi-layer perceptron to obtain the modal feature corresponding to the modal data group.

[0009] In an alternative approach, fusing the modal fusion feature corresponding to the modal data group and the reference modal data to generate the first fusion feature corresponding to the modal data group specifically includes: obtaining the first weight corresponding to the modal fusion feature and the second weight corresponding to the reference modal data; splicing the modal fusion feature and the reference modal data according to the first weight and the second weight to generate the first fusion feature corresponding to the modal data group.

[0010] In an alternative approach, determining the multi-modal fusion feature corresponding to the target information according to the second fusion feature of each reference modal data specifically includes: determining the feature weights corresponding to each reference modal data through a pre-trained weight determination model; determining the third fusion feature corresponding to each reference modal data according to the second fusion feature corresponding to each reference modal data and the feature weights corresponding to each reference modal data; splicing the third fusion features corresponding to each reference modal data to generate the multi-modal fusion feature corresponding to the target information.

[0011] In an alternative approach, obtaining multiple groups of modal data groups of the target information specifically includes: obtaining the target information, where the target information includes modal information of multiple modal types; using a pre-trained feature extraction model to perform feature extraction on each modal information respectively to obtain multiple modal data of the target information; respectively using each modal data as the reference modal data, and using each modal data as the associated modal data to generate multiple groups of modal data groups.

[0012] According to another aspect of the embodiments of the present application, a multi-modal feature fusion device is provided, including: an acquisition module, configured to acquire multiple groups of modal data groups of target information, where the modal data group includes reference modal data and associated modal data; a feature extraction module, configured to respectively perform feature extraction on each group of modal data groups to obtain modal fusion features corresponding to each modal data group; a judgment module, configured to respectively judge whether the modal types of the reference modal data and the associated modal data in each modal data group are the same; a fusion module, configured to, when the modal types of the reference modal data and the associated modal data in the modal data group are the same, fuse the modal fusion feature corresponding to the modal data group and the reference modal data to generate a first fusion feature corresponding to the modal data group; a first determination module, configured to, when the modal types of the reference modal data and the associated modal data in the modal data group are different, determine the modal fusion feature corresponding to the modal data group as the first fusion feature corresponding to the modal data group; a splicing module, configured to splice the first fusion features of each modal data group with the same reference modal data to generate a second fusion feature corresponding to the reference modal data; a second determination module, configured to determine the multi-modal fusion feature corresponding to the target information according to the second fusion features of each reference modal data.

[0013] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the multi-modal feature fusion method described in any one of the above.

[0014] According to yet another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the multi-modal feature fusion method described in any one of the above is implemented.

[0015] In the embodiments of the present application, by determining whether the modal types of the reference modal data and the associated modal data in the modal data group are the same, the multi-modal feature fusion method can use different calculation paths to generate the first fusion feature according to different judgment results, flexibly adjust the interaction intensity between the reference modal data and the associated modal data in different modal data groups, not only avoid data redundancy, but also effectively capture the dependency information within a single modality. On the one hand, when the modal types are the same, the reference modal data is fused with the modal fusion feature, so that the second fusion feature corresponding to the reference modal data not only has the association information between the reference modal data and other associated modal data, but also can retain the personalized information in the reference modal data; on the other hand, when the modal types are different, the modal fusion feature is directly used as the first fusion feature, which can not only avoid unnecessary calculations and information flows, but also reduce the interference of personalized information between modalities, improve the fusion effect of information between different modalities, and optimize information interaction.

[0016] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings are only used to illustrate the embodiments and are not considered as a limitation to the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0018] Figure 1 shows a schematic flowchart of the multi-modal feature fusion method provided by the embodiments of the present application;

[0019] Figure 2 shows a schematic flowchart of the modal data group generation process provided by the embodiments of the present application;

[0020] Figure 3 shows a partial flowchart of the multi-modal feature fusion method provided by the embodiments of the present application;

[0021] Figure 4 shows a flowchart of the multi-modal fusion feature generation process provided by the embodiments of the present application;

[0022] Figure 5 shows a flowchart of the sentiment classification task provided by the embodiments of the present application;

[0023] Figure 6 shows a schematic structural diagram of the dataset statistical information table involved in the embodiments of the present application;

[0024] Figure 7The structural schematic diagram of the ablation experiment result information table involved in the embodiments of the present application is shown;

[0025] Figure 8 The flowchart of the multi-modal feature fusion method provided by another embodiment of the present application is shown;

[0026] Figure 9 The structural schematic diagram of the multi-modal feature fusion device provided by the embodiments of the present application is shown;

[0027] Figure 10 The structural schematic diagram of the electronic device provided by the embodiments of the present application is shown. Detailed implementation manners

[0028] Hereinafter, the exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein.

[0029] A feature refers to an identifiable attribute or dimension extracted from the original data, representing certain important information of the data, which helps the model identify patterns and make predictions. Through research, it is found that after the initial feature extraction of modal information such as text, audio, and images, the initial features usually contain the personalized features of each modal information. After extracting the cross-modal association information from the initial features of multiple modal information to generate cross-modal features, the initial features and cross-modal features can be fused to generate new fusion features. In this way, not only can the mutual association information between modalities be fused to enhance the cross-modal features, but also the personalized features in the initial features can be retained.

[0030] For example, when obtaining the fusion feature of the text modal information, the cross-modal feature between the text modal information and the audio modal information can be extracted from the initial features of the audio modal information and the text modal information, and then the cross-modal feature, the initial feature of the text modal information, and the initial feature of the audio modal information are fused to generate the fusion feature. Such a fusion feature not only has the mutual association information between the text modality and the audio modality, but also retains the personalized features in the initial features of the text modal information.

[0031] However, when obtaining the fusion feature corresponding to the reference modal information, if the initial features of each modal information are directly fused with the cross-modal features, the initial features of the other modal information except the reference modal information will not only cause too many unnecessary personalized features in the fusion feature, but also increase the interference features in the fusion feature, thereby affecting the accuracy of the subsequent tasks. Moreover, when the initial features of the reference modal information are combined with multiple modal information respectively, it will also cause redundancy of the initial features of the reference modal information.

[0032] For example, the fusion features of the above-generated text modality information not only retain the personalized features in the initial features of the text modality information, but also retain the personalized features in the initial features of the audio modality information. For the fusion features of the text modality information, the personalized features in the initial features of the audio modality information are unnecessary features, which will not only increase unnecessary calculations and information flow, but also interfere with the fusion features and affect the accuracy of subsequent tasks. In addition, since the text modality information needs to be associated with multiple modality information respectively and generate multiple fusion features corresponding to the text modality information, subsequent tasks based on the text modality information need to analyze all these fusion features. However, the initial features of the text modality information are fused in all these fusion features, resulting in redundancy of the initial features of the text modality information and the need to analyze the initial features of the text modality information repeatedly for many times.

[0033] Therefore, in order to improve the efficiency and performance of feature fusion, the present application provides a multi-modal feature fusion method. After extracting cross-modal features from the reference modality information and the associated modality information, first determine whether the modality types of the reference modality information and the associated modality information are the same. Only when the modality types of the reference modality information and the associated modality information are the same, fuse the reference modality information with the cross-modal features to generate fusion features. If the modality types are different, directly determine the cross-modal features as the fusion features. Such a method can not only avoid unnecessary calculations and information flow and improve the efficiency of feature fusion, but also the fusion features corresponding to each reference modality information only include the personalized features in its own initial features, reducing the interference in the fusion features, optimizing the information interaction, and improving the fusion effect of information between different modalities.

[0034] The multi-modal feature fusion method provided by the embodiments of the present application can be applied to various tasks, such as, for example, sentiment analysis, recommendation systems, content generation, medical diagnosis, etc. For the convenience of description, the embodiments of the present application are only described by taking sentiment analysis as an example.

[0035] As an example, in the process of sentiment analysis, first obtain a multi-modal data set, where the multi-modal data set may include modality information such as text modality information, audio modality information, and image modality information, extract multi-modal fusion features from the text modality information, audio modality information, and image modality information by using the multi-modal feature fusion method provided by the embodiments of the present application, and finally analyze the multi-modal fusion features to determine the sentiment type corresponding to the multi-modal data set.

[0036] According to one aspect of the embodiments of the present application, a multi-modal feature fusion method is provided, as Figure 1 shown Figure 1The flowchart of the multi-modal feature fusion method provided by the embodiment of the present application is shown. This method is executed by an electronic device. The electronic device can be a terminal device such as a mobile phone or a tablet, or a local server, a cloud server, etc. As Figure 1 shown, the method includes the following steps:

[0037] Step S100: Obtain multiple groups of modal data groups of the target information, where the modal data group includes reference modal data and associated modal data.

[0038] Among them, the target information includes data in various forms of representation, such as text data, image data, audio data, etc. The modal data group is a combination of multiple modal data, mainly including reference modal data and associated modal data. For example, in the modal data group based on text data, the reference modal data is text data, and the associated modal data can be text data, image data, audio data, etc. The data included in the target information is the original data, that is, the data represented in the forms of text, image, audio, etc. Therefore, the modal data in the modal data group can be the original data.

[0039] In addition, since the original data usually contains a large amount of information, in order to improve the efficiency of the multi-modal feature fusion method, the modal data in the modal data group can also be the feature data obtained by extracting features from the original data, that is, the recognizable attributes or dimensions extracted from the original data, which can represent some important information of the original data. Specifically, as Figure 2 shown, Figure 2 The flowchart of the modal data group generation process provided by the embodiment of the present application is shown. Step S100 may include the following steps (step S110 to step S130):

[0040] Step S110: Obtain the target information, where the target information includes modal information of multiple modal types.

[0041] Step S120: Use the pre-trained feature extraction model to extract features from each modal information respectively to obtain multiple modal data of the target information.

[0042] Step S130: Respectively use each modal data as the reference modal data, and use each modal data as the associated modal data to generate multiple groups of modal data groups.

[0043] Among them, the modal information is the original data, that is, the data stored in the form of text, images, or audio, etc. The modal type refers to the different forms of the original data, such as text, images, audio, etc. The feature extraction model is used to extract key features from the modal information. Specifically, the feature extraction model can be a deep learning model such as a convolutional neural network, a recurrent neural network, a Transformer, etc., or a model using feature extraction algorithms such as scale-invariant feature transform, term frequency-inverse document frequency, Word2Vec, Mel-frequency cepstral coefficients, principal component analysis, etc.

[0044] After obtaining the modal data corresponding to each modal information, the modal data is combined to generate multiple groups of modal data groups. Specifically, each type of modal data among the multiple modal data is used as a reference modal data once. And after determining the reference modal data, each type of modal data is used as an associated modal data once. After combining the reference modal data and the associated modal data, a modal data group is generated.

[0045] As an example, assume that the target information includes text data and image data. Then the multiple modal data include text feature data and image feature data. First, using the text feature data as the reference modal data, and using the text feature data and the image feature data as the associated modal data respectively, two groups of modal data groups, namely text feature data → text feature data (t→t) and image feature data → text feature data (v→t), are combined and generated. Then, using the image feature data as the reference modal data, and using the text feature data and the image feature data as the associated modal data respectively, two groups of modal data groups, namely image feature data → image feature data (v→v) and text feature data → image feature data (t→v), are combined and generated. Finally, the multiple groups of modal data groups of the target information altogether include four groups of modal data groups: (t→t), (v→t), (v→v), and (t→v).

[0046] Through steps S110 to S130, feature extraction is performed on the multiple modal information in the target information, so that the modal data only includes the key features of the modal information, filtering out irrelevant or redundant information, and effectively reducing the dimension of the modal data. In addition, in the process of generating the modal data group, after taking a specific modal data as the reference, all modal data are combined with this specific data. Assume that the modal data includes text modal data, audio modal data, and image modal data. When using the text modal data as the reference modal data, the reference modal data needs to be combined with the text modal data, the audio modal data, and the image modal data respectively to generate the modal data group, fully considering the dependency information between different modal data and within a single modal data, and improving the performance of the multi-modal feature fusion method.

[0047] After step S100, step S200 is executed: Feature extraction is respectively performed on each group of modal data groups to obtain modal fusion features corresponding to each modal data group.

[0048] Among them, as Figure 3 shown, Figure 3 FIG. shows a schematic diagram of the model structure corresponding to the multi-modal feature fusion method provided by the embodiment of the present application. This feature extraction process is executed by the Figure 3 feature extraction module in, which is mainly used to extract deeper semantic information in the modal data and the mutual correlation information between the fused reference modal data and associated modal data, so as to enrich the information of the reference modal data through the associated modal data. The feature extraction module can be a pre-trained model (for example, deep learning models such as convolutional neural networks, recurrent neural networks, Transformers, etc.), and directly extract deep features and mutual correlation information from the reference modal data and the associated modal data through this model.

[0049] Of course, the feature extraction module can also include two extraction units, and parallel processing is performed on the modal data group through these two extraction units. Specifically, cross-modal long-distance dependence information and global alignment information between the reference modal data and the associated modal data can be extracted through one of the extraction units (this extraction can be an attention unit, a neural network unit, etc.). At the same time, through another extraction unit (this extraction unit can be a multi-layer perceptron unit, a convolutional neural network unit, etc.), local features of the reference modal data and the associated modal data and deep features for modal non-linear transformation are extracted. For the audio modality, non-linear representations of audio features can be extracted. For the text modality, high-level semantic features of word embeddings can be extracted. For the video modality, spatial features or dynamic features of frame-level features can be extracted.

[0050] Step S300: Determine whether the modal types of the reference modal data and the associated modal data in each modal data group are the same respectively.

[0051] Among them, if the modal types of the reference modal data and the associated modal data in the modal data group are the same, then the reference modal data and the associated modal features need to be fused with the modal fusion features, and step S400 is executed to fuse the unique information within the single modality.

[0052] If the modal types of the reference modal data and the associated modal data in the modal data group are different, it means that the personalized features in the associated modal data may interfere with the personalized features of the reference modal data, and step S500 is executed to avoid unnecessary calculations and information flows, reduce interference, and optimize information interaction.

[0053] Specifically, as Figure 3As shown, this step can be executed by the judgment module in the figure. The judgment module can be a gating unit. Through this gating unit, it is judged whether the reference modal data and the associated modal data belong to the same modal type. If the reference modal data and the associated modal data belong to the same modal type, 1 is output to activate the gating unit. If the reference modal data and the associated modal data belong to different modal types, 0 is output, indicating that the gating unit is disabled.

[0054] In addition, in order to improve the accuracy of the gating unit, when training the feature extraction module, the gating unit is trained synchronously, so that the gating unit can learn how to judge whether the reference modal data and the associated modal data belong to the same modal type, and gradually optimize the judgment mechanism as the training progresses. Specifically, through cross-validation and the input of data for various tasks, the judgment logic of the gating unit can be gradually adjusted, enabling it to flexibly adapt to different data modal combinations, enhancing the effect of multi-modal learning, and maintaining high stability under multi-modal data combinations.

[0055] Step S400: Fuse the modal fusion feature corresponding to the modal data group and the reference modal data to generate the first fusion feature corresponding to the modal data group.

[0056] Among them, when the output signal of the gating unit is 1 (i.e., the modal types are the same), the gating unit is activated, and the multi-modal feature fusion method will normally execute Figure 3 the feature extraction module in it to generate a modal fusion feature, and connect the gating unit to fuse the modal fusion feature and the reference modal feature, and generate the first fusion feature. This can reduce the attenuation of modal data of the same modal type in the feature extraction module, maintain the transmission of signals, and help retain the unique information of the modal data of the same modal type in the first fusion feature, thereby improving the accuracy of the multi-modal feature fusion method.

[0057] Furthermore, since different tasks have different emphases on associated information and personalized information, when generating the first fusion feature, fusion can be completed through weighted averaging or using attention weights. Specifically, step S400 can include the following steps (step S410 to step S420):

[0058] Step S410: Obtain the first weight corresponding to the modal fusion feature and the second weight corresponding to the reference modal data.

[0059] Step S420: Stitch the modal fusion feature and the reference modal data according to the first weight and the second weight to generate the first fusion feature corresponding to the modal data group.

[0060] Among them, the first weight and the second weight can be set to fixed values according to actual task requirements. That is, in an actual task, if the importance attached to associated information is higher than that of personalized information, the first weight can be set higher than the second weight; if the importance attached to personalized information is higher than that of associated information, the first weight can be set lower than the second weight. Of course, a weight determination model can also be pre-trained according to the actual task, and the first weight and the second weight can be dynamically assigned to the modality fusion feature and the reference modality data through this weight determination model.

[0061] Through steps S410 to S420, the fusion of the modality fusion feature and the reference modality data is completed through the first weight and the second weight, enabling the adjustment of the influence degrees of the associated information and the personalized information on subsequent tasks according to the actual task, and improving the fusion effect of the modality fusion feature and the reference modality data.

[0062] Step S500: Determine the modality fusion feature corresponding to the modality data group as the first fusion feature corresponding to the modality data group.

[0063] Among them, when the output signal of the gating unit is 0 (i.e., the modality types are different), the gating unit is disabled, and the multi-modal feature fusion method only normally executes Figure 3 the feature extraction module in, fuse the modality data of different modality types in the same representation space to generate a modality fusion feature, and determine the modality fusion feature as the first fusion feature. In this way, when the modality types of the reference modality data and the associated modality data are different, unnecessary calculation paths are closed, avoiding ineffective information fusion and computational burden, and reducing interference between modalities.

[0064] Specifically, the judgment process of the judgment module can be expressed by the following formula:

[0065]

[0066] Among them, m is the reference modality data, n is the associated modality data, t is the text modality data, a is the audio modality data, and v is the image modality data. Then, the first fusion feature T is generated through the following formula ′ m,n :

[0067] T′ m,n = T m,n + αG m,n

[0068] Among them, T m,nIt represents the modal fusion features obtained by the feature extraction module from the modal data group with m as the baseline modal data and n as the associated modal data. α is the threshold of the gating unit, which is a pre-set hyperparameter used to adjust the fusion ratio between the baseline modal data and the modal fusion features after activating the gating unit.

[0069] After step S400 or step S500, step S600 is executed: first fusion features corresponding to each modality data group having the same reference modality data are concatenated to generate second fusion features corresponding to the reference modality data.

[0070] Among them, after generating the first fusion features corresponding to each modal data group, it is necessary to fuse the first fusion features corresponding to the reference modal data to obtain the cross-modal features based on the reference modal data. As an example, assuming that the modal data includes text modal data t, audio modal data a, and image modal data v, and the modal data group with audio modal data a as the reference modal data includes three groups, namely (a→a), (t→a), and (v→a), and the first fusion features corresponding to these modal data groups are T a ′ ,a , T a ′ ,t and T a ′ ,v , therefore, the second fusion feature T of audio modality data a a It can be generated by the following formula:

[0071] T a =T a ′ ,a +T a ′ ,t +T a ′ ,v

[0072] Similarly, the second fusion feature T of text modal data t t =T t ′ ,a +T t ′ ,t +T t ′ ,v , the second fusion feature T of the image modality data v v =T v ′ ,a +T v ′,t +T v ′ ,v 。

[0073] Step S700: Determine the multi-modal fusion feature corresponding to the target information according to the second fusion feature of each reference modal data.

[0074] Specifically, the second fusion features of each reference modal data can be directly concatenated, or concatenated in a way of constant value weighting to generate the multi-modal fusion feature.

[0075] Furthermore, in view of the different degrees of influence of features of different modalities on downstream tasks, in order to distinguish the contribution degrees of features of different modalities to downstream tasks, step S700 may include the following steps (step S710 to step S730):

[0076] Step S710: Determine the feature weights corresponding to each reference modal data through a pre-trained weight determination model.

[0077] Step S720: Determine the third fusion feature corresponding to each reference modal data according to the second fusion feature corresponding to each reference modal data and the feature weight corresponding to each reference modal data.

[0078] Step S730: Concatenate the third fusion features corresponding to each reference modal data to generate the multi-modal fusion feature corresponding to the target information.

[0079] Specifically, as Figure 4 shown, Figure 4 shows a flowchart of the multi-modal fusion feature generation process provided by an embodiment of the present application. Assuming that the modal data includes text modal data t, audio modal data a, and image modal data v, first, input the second fusion feature T p corresponding to each reference modal data into the weight determination model, and calculate the weight coefficient W p corresponding to each reference modal data, and then multiply W p by the corresponding T p to obtain the weighted third fusion feature T p ′ . The specific formula is as follows:

[0080]

[0081] where, represents the multiplication operation.

[0082] Finally, concatenate the third fusion features T p ′ of all reference modal data to form the final multi-modal fusion feature R. The specific formula is as follows:

[0083]

[0084] Among them, represents a splicing operation.

[0085] In addition, the weight determination model can be trained separately according to different tasks. For example, when performing sentiment analysis, first obtain a modal data set related to sentiment analysis, generate multi-modal fusion features corresponding to each modal data, then determine the sentiment type according to the multi-modal fusion features, and finally determine the loss function according to the accuracy of the sentiment type, and optimize the weight determination model through the loss function. In this way, the weights corresponding to each modal data can be dynamically allocated for different tasks, enhancing the robustness of the multi-modal feature fusion method. The weight determination model can be composed of a Softmax non-linear activation function. Specifically, the probabilities corresponding to each benchmark modal data are obtained through this activation function to obtain a probability distribution, that is, the feature weights.

[0086] To further verify the effectiveness of the multi-modal feature fusion method provided in steps S710 to S730, the above multi-modal feature fusion method will be applied to the sentiment classification task below. As Figure 5 shown, Figure 5 shows the flow block diagram of the sentiment classification task provided by the embodiment of the present application, and tests are carried out on two publicly available multi-modal data sets: IEMOCAP and MELD. The statistical information of the data sets is as Figure 6 shown in Table 1 in Figure 6 shows the structural schematic diagram of the data set statistical information table involved in the embodiment of the present application.

[0087] All models use the weighted average F1 value (weighted average F1-score, abbreviated as w-F1), the seven-class classification accuracy (Accuracy of the seven classifications, abbreviated as Acc-7), and the six-class classification accuracy (Accuracy of the six classifications, abbreviated as Acc-6) as the evaluation indicators of the multi-modal feature fusion method.

[0088] To verify the effectiveness of steps S710 to S730, a set of ablation experiments was designed. The experimental results are as Figure 7 shown in Table 2 in Figure 7The structural schematic diagram of the ablation experiment result information table involved in the embodiments of the present application is shown. Specifically, S_Fusion is a model that performs feature fusion only by simple splicing method, and Origin is a model that performs feature fusion using steps S710 to S730. By comparing Experiment 3 and Experiment 4, it can be seen that after using steps S710 to S730 for feature fusion, the model improves by 1.21% and 0.83% respectively in terms of Acc-6 and w-F1 metrics on the IEMOCAP dataset, and improves by 1.20% and 0.36% respectively in terms of Acc-7 and w-F1 metrics on the MELD dataset. This shows that steps S710 to S730 can dynamically adjust the weights according to the contribution degree of each modality data to the final task by using the weights determined by the pre-trained model, which can effectively improve the accuracy of the model.

[0089] Through steps S710 to S730, it is possible to dynamically adjust the proportion weights of each modality data in the multi-modal fusion features according to the contribution degree to the final task, thereby optimizing the performance of the multi-modal feature fusion method, not only improving the flexibility of multi-modal feature fusion, but also enhancing the adaptability of the multi-modal feature fusion method to complex data structures.

[0090] Furthermore, in order to verify the effectiveness of the above embodiments, a set of ablation experiments were designed, and the experimental results are as Figure 7 shown in Table 2 below. Specifically, Gate is a model that directly determines the features extracted in step S200 as the first fusion feature, and Origin is a model that generates the first fusion feature using steps S300 to S500. By comparing Experiment 2 and Experiment 4, it can be seen that after using steps S300 to S500 to generate the first fusion feature, the model improves by 1.55% and 1.46% respectively in terms of Acc-6 and w-F1 metrics on the IEMOCAP dataset, and improves by 2.39% and 2.29% respectively in terms of Acc-7 and w-F1 metrics on the MELD dataset. This shows that generating the first fusion feature through steps S300 to S500 can effectively capture the dependency information within a single modality, which is beneficial to improving the accuracy of the model.

[0091] In the above embodiments, by determining whether the modal types of the reference modal data and the associated modal data in the modal data group are the same, the multi-modal feature fusion method can use different calculation paths to generate the first fusion feature according to different judgment results, flexibly adjust the interaction intensity between the reference modal data and the associated modal data in different modal data groups, not only avoid data redundancy, but also effectively capture the dependency information within a single modality. On the one hand, when the modal types are the same, the reference modal data is fused with the modal fusion feature, so that the second fusion feature corresponding to the reference modal data not only has the association information between the reference modal data and other associated modal data, but also can retain the personalized information in the reference modal data; on the other hand, when the modal types are different, the modal fusion feature is directly used as the first fusion feature, which can not only avoid unnecessary calculations and information flows, but also reduce the interference of personalized information between modalities, improve the fusion effect of information between different modalities, and optimize information interaction.

[0092] Furthermore, currently, most traditional serial attention mechanism structures are used to extract modal fusion features, but a simple serial attention mechanism structure will ignore the common feature information between modalities and is difficult to extract deeper multi-modal feature semantic information, resulting in insufficient expressive ability of modal features. For example Figure 8 as shown Figure 8 FIG. shows a flowchart of a multi-modal feature fusion method provided by another embodiment of the present application, and the method includes the following steps:

[0093] For each modal data group, the following operations are performed:

[0094] Step S210: Extract the dependency information between the reference modal data and the associated modal data to obtain the global feature corresponding to the modal data group.

[0095] Among them, the global feature is used to represent the association information between the reference modal data and the associated modal data, and the dependency information mainly includes cross-modal long-distance dependency information and global alignment information between modal data. The dependency information can be specifically extracted by algorithms such as attention mechanism and neural network.

[0096] Furthermore, in order to accurately extract the dependency information, as Figure 3 shown, see Figure 3 the structure in the small dotted box within the large dotted box on the right in, the dependency information can be extracted by an attention module. Step S210 specifically includes the following steps (steps S211 to S212):

[0097] Step S211: Determine the reference modal data as the query Q, and determine the associated modal data as the key K and the value V respectively.

[0098] Step S212: Process the query Q, key K, and value V through the attention unit to obtain the global features corresponding to the modal data group.

[0099] Specifically, process the reference modal data and associated modal data through multiple network blocks within the attention unit. For example, Figure 3 the structure within the small dashed box in the large right dashed box in left is the structure of a network block, and the global feature Tran

[0100] block i = Focus(QW i Q , KW i k , VW i V )

[0101] Tran left = Concat(block i , ···, block i )W o

[0102] where block i represents the feature extracted by the i-th network block in the attention unit, and W i Q , W i k , W i V and W o all represent the weight parameters in the attention unit, and h represents the number of network blocks in the attention unit.

[0103] In steps S211 to S212, the attention mechanism effectively captures the long-range dependence information between the reference modal data and the associated modal data, can effectively fuse the information between different modal data, and effectively improves the performance of the multi-modal feature fusion method.

[0104] Step S220: Extract the local features and deep features of the modal data group to obtain the modal features corresponding to the modal data group.

[0105] Among them, the modal features are used to represent the deeper semantic information in the modal data, mainly including the local features of the modal data and the deep features obtained by performing non-linear transformation on the modal data. For audio modal data, the modal features are the non-linear representation of spectral features; for text modal data, the modal features are the high-level semantic features of word embeddings; for video features, the modal features are the spatial features or dynamic features of frame-level features. Specifically, feature extraction can be performed through algorithms such as multi-layer perceptrons and neural networks.

[0106] Furthermore, in order to accurately obtain local features and deep features, such as Figure 3 shown, see in detail Figure 3 For the branch on the right side within the large black solid line box, the local features and deep features can be extracted through a multi-layer perceptron. Step S220 can specifically include the following steps (Step S221 to Step S222):

[0107] Step S221: Regularize the reference modal data and the associated modal data.

[0108] Step S222: Process the regularized reference modal data and the associated modal data through a multi-layer perceptron to obtain the modal features corresponding to the modal data group.

[0109] Specifically, the modal feature Tran right can be extracted through the following formula:

[0110] Tran right = Relu(MLP(Norm(MLP m,n )))

[0111] where H m,n represents the modal data group, Norm represents the regularization operation, MLP represents the multi-layer perceptron, and Relu represents the activation function.

[0112] In Steps S221 to S222, the multi-layer perceptron can effectively capture the complex non-linear relationship between the reference modal data and the associated modal data, and then accurately extract the deeper multi-modal feature semantic information in the modal data.

[0113] It should be particularly noted that Step S210 and Step S220 are independent steps, and there is no requirement for the execution order between them. That is, Step S210 can be executed first, and then Step S220, or Step S220 can be executed first, and then Step S210. In addition, Steps S210 and S220 can also be executed in parallel as Figure 3 shown, which can effectively accelerate the running speed of the multi-modal feature fusion method and improve the efficiency of feature fusion.

[0114] Step S230: Concatenate the global features and the modal features corresponding to the modal data group to generate the modal fusion features corresponding to the modal data group.

[0115] Specifically, the global feature Tran left and the modal feature Tran right can be concatenated through the following formula to generate the modal fusion feature T m,n :

[0116] Tm,n =Tran left +Tran right

[0117] Furthermore, in order to verify the effectiveness of steps S210 to S230, a set of ablation experiments was designed, and the experimental results are shown in Table 2 in Figure 7 . Specifically, P_Trans is a model that uses the traditional serial attention mechanism structure for feature extraction, and Origin is a model that uses steps S210 to S230 for feature extraction. By comparing Experiment 1 and Experiment 4, it can be seen that after using steps S210 to S230 for feature extraction, the model improves by 1.85% and 1.73% respectively in terms of the Acc-6 and w-F1 metrics on the IEMOCAP dataset, and improves by 2.68% and 2.86% respectively in terms of the Acc-7 and w-F1 metrics on the MELD dataset. This shows that steps S210 to S230 can fully extract deeper multi-modal semantic information and fusion of inter-modal correlation information in the conversation through parallel feature extraction, and can effectively improve the accuracy of the model.

[0118] In the above embodiments, by processing the modal data groups respectively and directly extracting the global features and modal features from the modal data groups, the gradients in the feature extraction process can be propagated through different paths, which helps to alleviate the problems of gradient disappearance or gradient explosion. Moreover, the global features and modal features can learn multiple features respectively, thereby enhancing the expression ability of the model.

[0119] According to another embodiment of the present application, a multi-modal feature fusion device is provided, as shown in Figure 9 . Figure 9 FIG. shows a schematic structural diagram of the multi-modal feature fusion device provided by the embodiment of the present application. The multi-modal feature fusion device 1 includes: an acquisition module 11, a feature extraction module 12, a judgment module 13, a fusion module 14, a first determination module 15, a splicing module 16, and a second determination module 17.

[0120] The acquisition module 11 is used to acquire multiple groups of modal data groups of target information, where the modal data group includes reference modal data and associated modal data. The feature extraction module 12 is used to perform feature extraction on each group of modal data groups respectively to obtain the modal fusion features corresponding to each modal data group. The judgment module 13 is used to judge whether the modal types of the reference modal data and the associated modal data in each modal data group are the same respectively. The fusion module 14 is used to fuse the modal fusion feature corresponding to the modal data group and the reference modal data when the modal types of the reference modal data and the associated modal data in the modal data group are the same, and generate the first fusion feature corresponding to the modal data group. The first determination module 15 is used to determine the modal fusion feature corresponding to the modal data group as the first fusion feature corresponding to the modal data group when the modal types of the reference modal data and the associated modal data in the modal data group are different. The splicing module 16 is used to splice the first fusion features of each modal data group with the same reference modal data to generate the second fusion feature corresponding to the reference modal data. The second determination module 17 is used to determine the multi-modal fusion feature corresponding to the target information according to the second fusion features of each reference modal data.

[0121] In the above embodiment, by judging whether the modal types of the reference modal data and the associated modal data in the modal data group are the same, the multi-modal feature fusion method can use different calculation paths to generate the first fusion feature according to different judgment results, and flexibly adjust the interaction intensity between the reference modal data and the associated modal data in different modal data groups. This can not only avoid data redundancy, but also effectively capture the dependency information within a single modality. On the one hand, when the modal types are the same, the reference modal data is fused with the modal fusion feature, so that the second fusion feature corresponding to the reference modal data not only has the association information between the reference modal data and other associated modal data, but also can retain the personalized information in the reference modal data. On the other hand, when the modal types are different, the modal fusion feature is directly used as the first fusion feature, which can not only avoid unnecessary calculations and information flows, but also reduce the interference of personalized information between modalities, improve the fusion effect of information between different modalities, and optimize information interaction.

[0122] According to another aspect of the embodiments of the present application, an electronic device is provided. Figure 10 The structural schematic diagram of the electronic device provided by the embodiments of the present application is shown. The specific implementation of the electronic device is not limited in the specific embodiments of the present application.

[0123] As Figure 10 shown, the electronic device 2 may include: a processor 21 and a memory 22.

[0124] Among them, the memory 22 is used to store the computer program 23. The memory 22 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory. The computer program 23 may include computer-executable instructions.

[0125] The processor 21 is used to execute the computer program 23 to implement the above-mentioned multi-modal feature fusion method embodiments.

[0126] The processor 21 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the electronic device 2 may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0127] The embodiments of the present application provide a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned multi-modal feature fusion method embodiments.

[0128] The embodiments of the present application provide a computer program. The computer program can be executed by a processor to implement the above-mentioned multi-modal feature fusion method embodiments.

[0129] The embodiments of the present application provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned multi-modal feature fusion method embodiments.

[0130] In several embodiments provided by the present application, if any function is implemented in the form of a software functional module / unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, part or all of the technical solutions of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be an electronic device such as a personal computer, a server, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., which can store computer program codes.

[0131] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. A variety of general-purpose systems may also be used in conjunction with the teachings based herein. The structure required to construct such systems will be apparent from the above description. Additionally, the embodiments of the present application are not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present application.

[0132] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a claim listing several devices, several units or modules of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

[0133] The above-described embodiments merely represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but should not be construed as limiting the patent scope of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A multimodal feature fusion method, characterized in that, The method includes: Obtaining multiple groups of modal data groups of target information, where each modal data group includes reference modal data and associated modal data; Respectively performing feature extraction on each group of the modal data groups to obtain modal fusion features corresponding to each modal data group; Respectively determining whether the modal types of the reference modal data and the associated modal data in each modal data group are the same; If they are the same, fusing the modal fusion feature corresponding to the modal data group and the reference modal data to generate a first fusion feature corresponding to the modal data group; Otherwise, determining the modal fusion feature corresponding to the modal data group as the first fusion feature corresponding to the modal data group; Concatenating the first fusion features corresponding to the modal data groups with the same reference modal data to generate a second fusion feature corresponding to the reference modal data; Determining multi-modal fusion features corresponding to the target information according to the second fusion features of each reference modal data.

2. The multimodal feature fusion method according to claim 1, wherein The step of respectively performing feature extraction on each group of the modal data groups to obtain modal fusion features corresponding to each modal data group specifically includes: For each modal data group, the following operations are performed: Extracting the dependency information between the reference modal data and the associated modal data to obtain the global feature corresponding to the modal data group; Extracting the local feature and the deep feature of the modal data group to obtain the modal feature corresponding to the modal data group; Concatenating the global feature corresponding to the modal data group and the modal feature to generate the modal fusion feature corresponding to the modal data group.

3. The multimodal feature fusion method according to claim 2, wherein The step of extracting the dependency information between the reference modal data and the associated modal data to obtain the global feature corresponding to the modal data group specifically includes: Determining the reference modal data as the query, and respectively determining the associated modal data as the key and the value; Processing the query, the key, and the value through an attention unit to obtain the global feature corresponding to the modal data group.

4. The multimodal feature fusion method according to claim 2, wherein The step of extracting the local feature and the deep feature of the modal data group to obtain the modal feature corresponding to the modal data group specifically includes: Performing regularization processing on the reference modal data and the associated modal data; Processing the reference modal data and the associated modal data after regularization processing through a multi-layer perceptron to obtain the modal feature corresponding to the modal data group.

5. The multimodal feature fusion method according to claim 1, wherein The step of fusing the modal fusion feature corresponding to the modal data group and the reference modal data to generate a first fusion feature corresponding to the modal data group specifically includes: Obtaining a first weight corresponding to the modal fusion feature and a second weight corresponding to the reference modal data; According to the first weight and the second weight, concatenating the modal fusion feature and the reference modal data to generate a first fusion feature corresponding to the modal data group.

6. The multimodal feature fusion method according to claim 1, wherein The step of determining multi-modal fusion features corresponding to the target information according to the second fusion features of each reference modal data specifically includes: Determining the feature weights corresponding to each reference modal data through a pre-trained weight determination model; Determine the third fusion feature corresponding to each of the reference modal data according to the second fusion feature corresponding to each of the reference modal data and the feature weight corresponding to each of the reference modal data; Concatenate the third fusion features corresponding to each of the reference modal data to generate a multi-modal fusion feature corresponding to the target information.

7. The multimodal feature fusion method according to claim 1, wherein The obtaining of multiple groups of modal data groups of the target information specifically includes: Obtain target information, where the target information includes modal information of multiple modal types; Use a pre-trained feature extraction model to respectively extract features from each of the modal information to obtain multiple modal data of the target information; Respectively use each of the modal data as reference modal data, and use each of the modal data as associated modal data to generate multiple groups of modal data groups.

8. A multimodal feature fusion device, characterized in that, The device includes: An obtaining module, configured to obtain multiple groups of modal data groups of target information, where the modal data group includes reference modal data and associated modal data; A feature extraction module, configured to respectively extract features from each of the groups of modal data groups to obtain a modal fusion feature corresponding to each of the groups of modal data groups; A judgment module, configured to respectively judge whether the modal type of the reference modal data and the modal type of the associated modal data in each of the modal data groups are the same; A fusion module, configured to, when the modal type of the reference modal data and the modal type of the associated modal data in the modal data group are the same, fuse the modal fusion feature corresponding to the modal data group and the reference modal data to generate a first fusion feature corresponding to the modal data group; A first determination module, configured to, when the modal type of the reference modal data and the modal type of the associated modal data in the modal data group are different, determine the modal fusion feature corresponding to the modal data group as the first fusion feature corresponding to the modal data group; A concatenation module, configured to concatenate the first fusion features of the modal data groups with the same reference modal data to generate a second fusion feature corresponding to the reference modal data; A second determination module, configured to determine a multi-modal fusion feature corresponding to the target information according to the second fusion feature of each of the reference modal data.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the multi-modal feature fusion method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the multi-modal feature fusion method according to any one of claims 1 to 7.