Multi-modal Feature Fusion Method, Apparatus, Computing Device, and Medium
By combining multi-channel features in modal dimensions and using convolution processing, the problem of insufficient information complementarity in modal feature fusion in multimedia resources is solved, and more accurate multimedia content detection is achieved.
Patent Information
- Application Number
- CN202110291490.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-03-18
AI Technical Summary
The multimodal feature fusion method of multimedia resources in the prior art fails to effectively consider the relationship between each modal, resulting in insufficient information complementarity.
By combining multi-channel features in a modal dimension and fusing them using convolution processing methods, fusion features of multi-channel features are generated.
The information complementarity between different modal features is achieved, the expression ability of fusion features is improved, the impact of a large single feature value is reduced, and the accuracy of multimedia content detection is improved.
Smart Images

Figure CN113033647B_ABST
Abstract
Description
Background Art
[0002] This section aims to provide background or context for the embodiments of the present disclosure described in the claims. The descriptions herein are not admitted to be prior art merely by virtue of their inclusion in this section.
[0003] With the development of Internet technology, more and more multimedia resources have emerged on the network. For multimedia resources, such as videos composed of various modal data such as images and audio, how to extract the features of multimedia resources has become the focus of attention.
[0004] In related technical solutions, the features of each modality of multimedia resources are extracted, and the features of each modality are directly numerically concatenated in a certain dimension. For example, assume that the features of two modalities of multimedia resources include: one-dimensional feature a and one-dimensional feature b, where a = [1, 2, 3] and b = [7, 6, 5, 4]. The features a and b are concatenated in the current single dimension to obtain one-dimensional feature c = [1, 2, 3, 7, 6, 5, 4]. Summary of the Invention
[0005] However, in the above technical solution, since the concatenation operation is performed in a single dimension, the mutual relationship between the features of each modality is not considered, and information complementarity between modalities cannot be achieved.
[0006] Therefore, there is a great need for an improved multi-modal feature fusion method to enable information complementarity between different modal features and effectively fuse the features of different modalities.
[0007] In this context, embodiments of the present disclosure are expected to provide a multi-modal feature fusion method, a multi-modal feature fusion device, a computing device, and a medium.
[0008] In a first aspect of the embodiments of the present disclosure, a multi-modal feature fusion method is provided, including: extracting the features of each modality in multiple modalities of multimedia resources; combining the features of each modality in the modality dimension to generate multi-channel features, where each channel of the multi-channel features corresponds to one of the modalities; performing convolution processing on the multi-channel features to generate fusion features corresponding to the multi-channel features.
[0009] According to the first aspect of the present disclosure, in some exemplary embodiments, the multi-channel features are feature vectors of L×D×N dimensions, where N is the number of modalities of the multiple modalities, and the performing convolution processing on the multi-channel features includes: performing convolution processing on the multi-channel features in the direction of dimension D through C1 first convolution kernels, where the size of the first convolution kernel is determined according to the value of dimension D, and the dimension of the first fusion feature is L×D×C1.
[0010] According to a first aspect of the present disclosure, in some example embodiments, the convolutional processing of the multi-channel features further includes: performing convolutional processing on the first fused feature in the direction of dimension D through C2 second convolutional kernels to generate a second fused feature, where the dimension of the second fused feature is L×D×C2, and the size of the second convolutional kernel is smaller than the size of the first convolutional kernel.
[0011] According to a first aspect of the present disclosure, in some example embodiments, the method further includes: after extracting the features of each modality, performing feature aggregation on the features of each modality based on a predetermined grouping; and generating feature vectors of each modality with the same dimension based on the result of the feature aggregation.
[0012] According to a first aspect of the present disclosure, in some example embodiments, the performing feature aggregation on the features of each modality based on a predetermined grouping includes: using a Nextvlad model to perform feature aggregation on the features of different dimensions of each modality based on a predetermined grouping.
[0013] According to a first aspect of the present disclosure, in some example embodiments, the method further includes: performing stretching processing on the fused feature in the modality dimension to generate a corresponding third fused feature.
[0014] According to a first aspect of the present disclosure, in some example embodiments, the multimedia resource includes image frame data, audio data, and text data of a video, and the extracting of features of multiple modalities of the multimedia resource includes: extracting image features corresponding to the video from the image frame data; extracting audio features corresponding to the video from the audio data; and extracting text features corresponding to the video from the text data.
[0015] According to a first aspect of the present disclosure, in some example embodiments, the multiple modalities include at least two of an image modality, an audio modality, and a text modality.
[0016] In a second aspect of the embodiments of the present disclosure, there is provided a multi-modal feature fusion device, including: a feature extraction module configured to extract features of each modality among multiple modalities of a multimedia resource; a modality combination module configured to combine the features of each modality in the modality dimension to generate multi-channel features, where each channel of the multi-channel features corresponds to one of the modalities; and a convolutional processing module configured to perform convolutional processing on the multi-channel features to generate a fused feature corresponding to the multi-channel features.
[0017] According to a second aspect of the present disclosure, in some exemplary embodiments, the multi-channel feature is a feature vector of L×D×N dimensions, where N is the number of modalities of the multiple modalities, and the convolution processing module is further configured to: perform convolution processing on the multi-channel feature in the direction of dimension D through C1 first convolution kernels to generate a first fusion feature, the size of the first convolution kernel being determined according to the value of dimension D, and the dimension of the first fusion feature being L×D×C1.
[0018] According to a second aspect of the present disclosure, in some exemplary embodiments, the convolution processing module is further configured to: perform convolution processing on the first fusion feature in the direction of dimension D through C2 second convolution kernels to generate a second fusion feature, the dimension of the second fusion feature being L×D×C2, and the size of the second convolution kernel being smaller than the size of the first convolution kernel.
[0019] According to a second aspect of the present disclosure, in some exemplary embodiments, the apparatus further includes: a feature aggregation module, configured to perform feature aggregation on the features of each modality based on a predetermined grouping after extracting the features of each modality; and a feature generation module, configured to generate feature vectors of each modality with the same dimension based on the result of the feature aggregation.
[0020] According to a second aspect of the present disclosure, in some exemplary embodiments, the feature aggregation module is further configured to: adopt a Nextvlad model to perform feature aggregation on the features of different dimensions of each modality based on a predetermined grouping.
[0021] According to a second aspect of the present disclosure, in some exemplary embodiments, the apparatus further includes: a stretching processing module, configured to perform stretching processing on the fusion feature in the modality dimension to generate a corresponding third fusion feature.
[0022] According to a second aspect of the present disclosure, in some exemplary embodiments, the multimedia resource includes image frame data, audio data, and title data of a video, and the feature extraction module is further configured to: extract image features corresponding to the video from the image frame data; extract audio features corresponding to the video from the audio data; and extract title features corresponding to the video from the title data.
[0023] According to a second aspect of the present disclosure, in some exemplary embodiments, the multiple modalities include at least two of image, audio, and text.
[0024] In a third aspect of the embodiments of the present disclosure, there is provided a computing device, including: a processor and a memory, the memory storing executable instructions, and the processor being configured to call the executable instructions stored in the memory to execute the method according to any one of the above first aspects.
[0025] In a fourth aspect of the embodiments of the present disclosure, a medium is provided, on which a program is stored, and when the program is executed by a processor, the method described in any one of the above first aspects is implemented.
[0026] According to the technical solution of the embodiments of the present disclosure, on the one hand, by adopting the convolution processing method to perform inter-channel fusion processing on the multi-channel features composed of different modality features, it is possible to achieve information complementarity between different modality features and effectively fuse the features of different modalities; on the other hand, since the fusion result does not depend on the eigenvalue size of the feature vector itself, the influence of a relatively large single eigenvalue is reduced; on the further hand, since the internal connection between each modality feature is mined, the expression ability of the fusion feature is improved, so that more accurate multimedia content detection can be achieved by using the fusion feature. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] By referring to the drawings and reading the detailed description below, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become easily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, where:
[0028] Figure 1 A schematic diagram showing an application scenario of the multi-modal feature fusion method according to the embodiments of the present disclosure is shown;
[0029] Figure 2 A flowchart of the multi-modal feature fusion method according to some embodiments of the present disclosure is schematically shown;
[0030] Figure 3 A flowchart of the multi-modal feature fusion method according to some other embodiments of the present disclosure is schematically shown;
[0031] Figure 4 A flowchart of the multi-modal feature fusion method according to some embodiments of the present disclosure is schematically shown;
[0032] Figure 5 A schematic diagram of a computer-readable storage medium according to some embodiments of the present disclosure is schematically shown;
[0033] Figure 6 A schematic diagram of the structure of a multi-modal feature fusion device according to some embodiments of the present disclosure is schematically shown;
[0034] Figure 7 A schematic diagram of the structure of a computing device according to some embodiments of the present disclosure is schematically shown.
[0035] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0037] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0038] According to the embodiments of the present disclosure, a multi-modal feature fusion method, device, computing device, and medium are provided.
[0039] In this article, it should be understood that the terms involved:
[0040] Modal: Represents the source or form of information. Each source or form of information can be called a modality. For example, the information of multimedia resources has multiple modalities such as voice, image, and text.
[0041] Feature: Refers to the high-dimensional numerical representation of data or information after being processed by a computer model, and can also be called a feature vector. Video features refer to the high-dimensional numerical representation of video data after being processed by a computer model, and can also be called video vectors. Since features are the numerical output of the model for videos, different models have different preference effects and output features for videos. For example, a visual model is used to extract the visual information of a video, then the output is the visual feature of the video; an audio model is used to extract the sound information of a video, and the output is the audio feature; a text model is used to extract the text or title information content of a video, and the output is the text feature, etc.
[0042] Multi-modal features: The feature representations of each modality of a multimedia resource together constitute the numerical representation of the multimedia resource. The features of multiple modalities of a multimedia resource can be called the multi-modal features of the multimedia resource. Suppose the multimedia resource is a video, then the multi-modal features include audio modality features, image modality features, and text modality features. The feature representations of each modality together constitute the numerical representation of the video. The feature representations of multiple modalities of the video are called video multi-modal features.
[0043] Multi-modal feature fusion: Integrate the features of each modality of a multimedia resource to achieve information complementarity between the features of each modality.
[0044] Convolution processing: Use multiple convolutional kernels to aggregate the features of each modality of the multimedia resource. Assume that the multimedia resource has features of n modalities, and the features of the n modalities of the multimedia resource are aggregated by m convolutional kernels.
[0045] In addition, the number of any elements in the accompanying drawings is for illustration rather than limitation, and any naming is only for distinction without any limiting meaning.
[0046] Next, with reference to several representative embodiments of the present disclosure, the principles and spirit of the present disclosure will be elaborated in detail. Summary of the Invention
[0048] The inventor of the present invention found that in order to solve the problem of how to fuse the features of each modality, in the related technical solutions, the features of each modality of the multimedia resource are extracted, and the features of each modality are fused in a bitwise fusion manner, for example, by element-wise sum or element-wise average or element-wise product. Taking element-wise sum as an example, assume that the features of multiple modalities of the multimedia resource include: one-dimensional feature a and one-dimensional feature b, where a = [1, 2, 3], and b = [3, 6, 5], and the feature a and the feature b are added bitwise to obtain the feature c, and the feature c = [4, 8, 8]. In the above technical solution, on the one hand, since the feature values between different modalities are fused by mathematical calculation methods, if the feature value corresponding to a certain modality's feature is large, it will mask the role of other modality features; on the other hand, fusing different modalities by mathematical calculation methods does not fully consider the mutual relationship between the features of each modality and cannot achieve information complementarity between each modality.
[0049] Based on the above content, the basic idea of the embodiments of the present disclosure is: combine or stack the features of different modalities of the multimedia resource in the modality dimension direction, and perform convolution processing on the combined or stacked multi-channel features to generate fused features. According to the technical solution of the embodiments of the present disclosure, on the one hand, using convolution processing to perform inter-channel fusion processing on the multi-channel features composed of different modality features can achieve information complementarity between different modality features and effectively fuse the features of different modalities; on the other hand, since the fusion result does not depend on the feature value size of the feature vector itself, the influence of a single large feature value is reduced; on the other hand, since the internal connection between the features of each modality is mined, the expression ability of the fused features is improved, so that more accurate multimedia content detection can be achieved using the fused features.
[0050] After introducing the basic principles of the present disclosure, the various non-limiting embodiments of the present disclosure will be specifically introduced below.
[0051] Overview of Application Scenarios
[0052] It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0053] Figure 1 A block diagram schematically showing an application scenario of a multi-modal feature fusion method according to an embodiment of the present disclosure is shown.
[0054] Referring to Figure 1 As shown, the application scenario may include: at least one client 110 and a server 120, wherein a multimedia application is installed on the client 110. Communication between the client 110 and the server 120 is carried out through a network 130. Taking the multimedia as a video as an example, the video uploaded by the user on the client 110 will be sent to the server 120 through the network 130. The server 120 can apply the multi-modal feature fusion method of the embodiment of the present disclosure to extract features and fuse features from the received video data, and obtain fusion features of multiple modalities of the video. The fusion features can be used for further processing such as video classification processing or video recommendation processing.
[0055] It should be noted that the client 110 can be a mobile phone, a tablet computer, a desktop computer, a portable notebook computer or a vehicle-mounted terminal, etc. The server 120 can be a physical server including an independent host, or a virtual server hosted by a host cluster, or a cloud server. The network 130 can be a wired network or a wireless network. For example, the network 130 can be a PSTN (Public Switched Telephone Network) or the Internet.
[0056] It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0057] Exemplary Method
[0058] Next, in combination with the above application scenario, referring to Figure 2 to describe a multi-modal feature fusion method according to an exemplary embodiment of the present disclosure. The multi-modal feature fusion method can be applied to Figure 1 the server 120.
[0059] Referring to Figure 2 As shown, in step S210, features of each modality among multiple modalities of the multimedia resource are extracted.
[0060] In an exemplary embodiment, taking the multimedia resource as a video as an example, video data is obtained. The video data may include image frame data, audio data, and text data. The text data may include one or more of subtitle data, title data, and bullet screen data. The modalities of the video include: image modality, audio modality, and text modality.
[0061] Furthermore, features of each modality of the video are extracted from the video data. For example, feature extraction models of different modality forms that are pre-trained can be used to extract features of each modality of the video from the video data. For example, an image feature extraction model such as the inceptionv3 model is used to extract the image features of the video from the image frame data; an audio feature extraction model such as the vggish model is used to extract the audio features of the video from the audio data; a text feature extraction model such as the word2vec model is used to extract the text features of the video from the text data.
[0062] It should be noted that although the multimedia resource is taken as a video as an example for description, the multimedia resource can also be other appropriate content such as audio resources or animation resources, etc., which is also within the protection scope of the present disclosure.
[0063] In step S220, the features of each modality are combined in the modality dimension to generate multi-channel features, where each channel of the multi-channel features corresponds to a modality.
[0064] In an exemplary embodiment, the features of each obtained modality are combined or stacked in the modality dimension to generate multi-channel features. Taking the multimedia resource as a video as an example, assuming that the obtained image features, audio features, and text features of the video are all feature vectors of L×D dimensions, then the image features, audio features, and text features are combined into multi-channel features of L×D×3 dimensions. Each channel feature is a modality feature, where 3 represents the modality dimension or channel dimension. Since the video has three modalities: image, audio, and text, the number of modalities or channels is 3. Among them, each channel of the multi-channel features corresponds to a modality, that is, the image channel feature corresponds to the image modality feature, the audio channel feature corresponds to the audio modality feature, and the text channel feature corresponds to the text modality feature.
[0065] It should be noted that although the modality dimension or the number of channels is taken as 3 as an example for description, the embodiments of the present disclosure are not limited thereto. For example, the modality dimension can also be other appropriate numbers such as 2 or 4, etc., which is also within the protection scope of the present disclosure.
[0066] In step S230, the multi-channel features are subjected to convolution processing to generate fusion features corresponding to the multi-channel features.
[0067] Convolution processing can utilize multiple convolutional kernels to aggregate features of different channels. For example, it can aggregate the features of three channels, namely images, audio, and text, using multiple convolutional kernels. The convolutional kernels are obtained by training based on a machine learning model. Taking the video classification scenario as an example, a video classification model is adopted. The video classification model includes a convolutional network, and the convolutional network includes multiple convolutional kernels. The parameters of each convolutional kernel can be obtained by training the video classification model with labeled training samples.
[0068] In an exemplary embodiment, a predetermined number of convolutional kernels are used to perform convolution processing on the multi-channel features corresponding to the multimedia resource, generating the fusion features of the corresponding multimedia resource. The size of the convolutional kernel can be adjusted according to the vector length of each modality.
[0069] For example, in the video classification scenario, assume that the multi-channel features of the video include image channel features, audio channel features, and text channel features. Among them, the image channel features, audio channel features, or text channel features are all one-dimensional features, and the length of the feature is 256. Assume that the multi-channel features are 1×256×3, the number of convolutional kernels is 128, and the size of the convolutional kernel is 64. Then the obtained fusion features are 1×256×128-dimensional features.
[0070] According to Figure 2 the technical solution in the exemplary embodiment, on the one hand, by using convolution processing to perform inter-channel fusion processing on the multi-channel features composed of different modality features, it can achieve information complementarity between different modality features and effectively fuse the features of different modalities; on the other hand, since the fusion result does not depend on the eigenvalue size of the feature vector itself, the influence of a single large eigenvalue is reduced; on the further hand, since the internal connection between the features of each modality is mined, the expression ability of the fusion features is improved, so that more accurate multimedia content detection can be achieved using the fusion features.
[0071] In addition, to facilitate the combination of the features of each modality and reduce the subsequent data processing volume, before combining the features of each modality in the modality dimension, feature aggregation processing can be performed on the features of each modality.
[0072] In an exemplary embodiment, after extracting the features of each modality, feature aggregation is performed on the features of each modality based on a predetermined grouping; feature vectors of each modality with the same dimension are generated based on the result of feature aggregation. For example, the nextvlad model can be used to perform feature aggregation on the features of each modality, that is, the nextvald model is used to divide the features of each modality into a predetermined number of groups, perform feature aggregation or clustering on the features of each group, and generate modality features with the same dimension based on the aggregation result, such as image modality features, audio modality features, and text modality features with the same dimension.
[0073] Figure 3 A flowchart of a multimodal feature fusion method according to some other embodiments of the present disclosure is schematically shown. The multimodal feature fusion method can be applied to Figure 1 server 120.
[0074] Referring to Figure 3 As shown, in step S310, feature aggregation is performed on the features of each modality to generate feature vectors of the same dimension.
[0075] In an exemplary embodiment, different feature extraction models are used to extract corresponding modality features, and feature aggregation is performed on the extracted features of each modality. For example, an image feature extraction model such as the inceptionv3 model is used to extract the image features of a video from image frame data; an audio feature extraction model such as the vggish model is used to extract the audio features of a video from audio data; a text feature extraction model such as the word2vec model is used to extract the text features of a video from text data.
[0076] Furthermore, the nextvlad model can be used to perform feature aggregation on the features of each modality, that is, the nextvald model is used to divide the features of each modality into a predetermined number of groups, perform feature aggregation or clustering on the features of each group, and generate modality features of the same dimension based on the aggregation result, such as image modality features, audio modality features, and text modality features of the same dimension. For example, the image modality features, audio modality features, and text modality features can be processed into one-dimensional vectors, and the vectors have the same length D. At this time, the sizes of the image modality features, audio modality features, and text modality features can be expressed as 1×D.
[0077] In step S320, the features of each modality are combined in the modality dimension to generate multi-channel features, where each channel of the multi-channel features corresponds to one modality.
[0078] The image modality features, audio modality features, and text modality features respectively represent information of a video in different dimensions, and the features of each modality are relatively important for understanding the video content. The key to multimodal feature fusion is to comprehensively integrate the content information of each modality so that the information between the features of each modality complements each other. A main feature of convolution operation is that multiple convolution kernels can be used to aggregate the features between different channels and the local features of each channel, so as to obtain deeper deep features, and to a certain extent, consider the adjacent values of the features and the relationship between the features of different channels. Borrowing the idea of convolution operation, considering the mutual relationship between different modality features, the modality is extended to the "channel" of convolution operation, that is, the features of each modality are superimposed in the modality dimension or channel dimension, and the number of modalities can be 0 to N.
[0079] Therefore, in the exemplary embodiment, the features of different modalities are superimposed in the modality dimension or the channel dimension direction. Assuming there are N modalities and the feature size of each modality is 1×D, the multi-channel feature after superposition can be represented as 1×D×N. Taking a video as an example, a video has three modalities: image, audio, and text, and the multi-channel feature after superposition is 1×D×3. Figure 4 a shows the 1×D×3-dimensional multi-channel feature of the video. Referring to Figure 4 as shown in a, the multi-channel feature includes 3 1×D-dimensional features, i.e., squares. Among them, the square filled with vertical lines is the image channel feature; the square filled with grid lines is the audio channel feature; the square filled with dots is the text channel feature.
[0080] In step S330, C1 first convolutional kernels are used to perform convolutional processing on the L×D×N-dimensional multi-channel feature in the direction of dimension D to generate a first fusion feature, and the size of the first convolutional kernel is determined according to the value of dimension D.
[0081] In the exemplary embodiment, C1 convolutional kernels of size L×K are used to perform a convolutional operation on the L×D×N vector along the direction of D. Among them, the stride is set to 1, and the padding method is same. The first fusion feature after convolutional processing is an L×D×C1-dimensional feature, that is, the number of channels of the first fusion feature is the same as the number of the first convolutional kernels.
[0082] Taking a video as an example, the multi-channel feature is 1×D×3-dimensional. Convolutional calculation is performed on the superimposed multi-channel feature, and the input is a feature map of 1×D×3. Suppose the number of convolutional kernels is C1, the size of the convolutional kernel is 1×K, the stride is set to 1, and the padding method is same. The feature vector of the first fusion feature output after one convolutional operation is a 1×D×C1-dimensional feature.
[0083] It should be noted that the size K of the convolutional kernel can be adjusted according to the vector length of a single modality, and the number C1 of the convolutional kernels can be set according to the machine learning task, for example, set according to the classification number or the label number of the video classification or multi-label task. The purpose of this convolutional operation is to fuse the features of each modality in a larger receptive field. That is to say, the size of the convolutional kernel in step S330 can be set relatively large. In video classification or multi-label processing, assuming that the feature lengths of the image modality feature, the audio modality feature, and the text modality feature after aggregation are all 256, the value of the first-layer convolutional kernel K is 64, and the number of convolutional kernels is 128, then the output first fusion feature is a 1×256×128-dimensional feature. Figure 4 b shows a schematic diagram of performing convolutional processing on the multi-channel feature using a convolutional kernel of size 1×64.
[0084] In step S340, the first fused feature is convolved in the direction of dimension D by C2 second convolutional kernels to generate a second fused feature. The dimension of the second fused feature is L×D×C2, and the size of the second convolutional kernel is smaller than that of the first convolutional kernel.
[0085] In an exemplary embodiment, convolving the first fused feature with C2 convolutional kernels having a smaller size than the first convolutional kernel can fuse the multi-channel features again at a fine-grained level, and the processed feature vector is a feature of dimension L×D×C2.
[0086] By extracting more fine-grained local relationships of the first fused feature, it is possible to effectively fuse the features in the local ranges of multiple modalities, achieving the effect of information complementarity among the modalities. Taking a video as an example, because at the same moment or within a short time interval, the video, audio, or subtitles often have a closer connection. By using a smaller convolutional kernel, the video, audio, or subtitles at the same moment or within a short time interval can be more effectively fused, achieving the effect of information complementarity among the modalities. In video classification or multi-label processing, the size K of the second convolutional kernel can be set to 4, and the number of convolutional kernels is 64. Therefore, the output second fused feature is a feature of dimension 1×256×64. Figure 4 c shows a schematic diagram of convolving the first fused feature with a convolutional kernel of size 1×4.
[0087] Since the parameters of the convolutional kernel for convolution processing can be learned through a machine learning model, according to the technical solution of the embodiments of the present disclosure, it is possible to extract representative feature contents of different modalities from the perspective of model learning, and perform feature fusion based on these feature contents to achieve better information expression.
[0088] In step S350, the second fused feature is stretched in the modality dimension or the channel dimension to generate a corresponding third fused feature.
[0089] In an exemplary embodiment, the second fused feature after two convolution operations is stretched in the modality dimension or the channel dimension to generate a corresponding third fused feature, that is, a feature of dimension L×(D×C2).
[0090] In the present exemplary embodiment, considering that feature fusion does not change the original vector form, the fused feature can be changed to a feature having the same dimension as the input features of each modality, that is, the third fused feature is stretched in the channel dimension direction. Taking video classification or multi-label tasks as an example, the second fused feature of the video, that is, the feature of dimension 1×256×64, is stretched to a feature of dimension 1×16384, which has the same dimension as the original image, audio, or text features. Figure 4 d shows the feature after stretching the second fused feature.
[0091] According to Figure 3 In the technical solution of the exemplary embodiment, on the one hand, the features of different modalities are fused by means of convolution processing, effectively extracting the key components in the features of each modality and fusing them, realizing the complementary effect of the information of each modality. Since the convolution operation is a non-linear calculation method, it can fit the multi-modal features of the video to a great extent. And because there is a correlation and locality between different modalities of the video and between the modality features at adjacent times, for example, for a video content, there will be a certain corresponding relationship between the video image content and the audio content in a short period of time. Therefore, the convolution processing can effectively fuse the features in the local range to achieve the effect of information complementation.
[0092] On the other hand, the different modality features are processed to the same length. Without restricting the specific numerical size of the features, the fusion result does not depend on the numerical value of the vector itself. Since the convolution result is the combined action of multiple convolution kernels, different convolution kernels have different extraction effects on the multi-modal superposition features, that is, multi-channel features, and the fusion result is the comprehensive result of multi-faceted features, greatly weakening the influence of a single feature value.
[0093] On yet another hand, since the convolution parameters can be specifically set and learned according to the specific problems of processing tasks such as video multi-label tasks or video recommendation tasks, the number of parameters of the model can be flexibly controlled to avoid the model being too complex.
[0094] Exemplary Medium
[0095] After introducing the method of the exemplary embodiment of the present disclosure, next, the medium of the exemplary embodiment of the present disclosure will be described.
[0096] In some possible embodiments, each aspect of the present disclosure can also be implemented as a medium having program code stored thereon, which is used to implement the steps in the multi-modal feature fusion method according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification when the program code is executed by a processor of a device.
[0097] In some possible embodiments, when the processor of the device executes the program code, it is used to implement the following steps: Step S210, extracting the features of each modality in multiple modalities of a multimedia resource; Step S220, combining the features of each modality in the modality dimension to generate multi-channel features, where each channel of the multi-channel features corresponds to a modality; Step S230, performing convolution processing on the multi-channel features to generate fusion features corresponding to the multi-channel features.
[0098] Refer to Figure 5As shown, a program product 500 for implementing the above multi-modal feature fusion method according to an embodiment of the present disclosure is described. It can be a portable compact disc read-only memory and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto.
[0099] It should be noted that the above medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disc, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disc read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0100] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to: an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program used by or in conjunction with an instruction execution system, apparatus, or device.
[0101] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, optical cable, radio frequency signal, etc., or any suitable combination of the above.
[0102] The program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network - including a local area network or a wide area network - or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0103] Exemplary Apparatus
[0104] After introducing the medium of the exemplary embodiment of the present disclosure, next, referring to Figure 6A multimodal feature fusion device according to an exemplary embodiment of the present disclosure will be described.
[0105] Figure 6 A structural diagram of a multimodal feature fusion device according to some embodiments of the present disclosure is schematically shown.
[0106] Referring to Figure 6 As shown, the multimodal feature fusion device 600 includes: a feature extraction module 610 for extracting features of each modality in multiple modalities of a multimedia resource; a modality combination module 620 for combining the features of each modality in the modality dimension to generate multi-channel features, wherein each channel of the multi-channel features corresponds to one of the modalities; and a convolution processing module 630 for performing convolution processing on the multi-channel features to generate fusion features corresponding to the multi-channel features.
[0107] According to Figure 6 the technical solution of the exemplary embodiment, on the one hand, by using the convolution processing method to perform inter-channel fusion processing on the multi-channel features composed of different modality features, it is possible to achieve information complementarity between different modality features and enable effective fusion of different modality features; on the other hand, since the fusion result does not depend on the eigenvalue size of the feature vector itself, the influence of a single large eigenvalue is reduced; on the further hand, since the internal connection between the features of each modality is mined, the expression ability of the fusion features is improved, so that more accurate multimedia content detection can be achieved by using the fusion features.
[0108] In some exemplary embodiments, the multi-channel features are feature vectors of L×D×N dimensions, where N is the number of modalities of the multiple modalities, and the convolution processing module 630 is further configured to: perform convolution processing on the multi-channel features in the direction of dimension D through C1 first convolution kernels to generate first fusion features, the size of the first convolution kernels is determined according to the value of the dimension D, and the dimension of the first fusion features is L×D×C1.
[0109] In some exemplary embodiments, the convolution processing module 630 is further configured to: perform convolution processing on the first fusion features in the direction of dimension D through C2 second convolution kernels to generate second fusion features, the dimension of the second fusion features is L×D×C2, and the size of the second convolution kernels is smaller than the size of the first convolution kernels.
[0110] In some exemplary embodiments, the device 600 further includes: a feature aggregation module for performing feature aggregation on the features of each modality based on a predetermined grouping after extracting the features of each modality; and a feature generation module for generating feature vectors of each modality with the same dimension based on the result of the feature aggregation.
[0111] In some exemplary embodiments, the feature aggregation module is further configured to: use the Nextvlad model to perform feature aggregation on features of different dimensions of each modality based on a predetermined grouping.
[0112] In some exemplary embodiments, the apparatus 600 further includes: a stretching processing module, configured to perform stretching processing on the fused features in the modality dimension to generate corresponding third fused features.
[0113] In some exemplary embodiments, the multimedia resource includes image frame data, audio data, and title data of a video, and the feature extraction module 610 is further configured to: extract image features corresponding to the video from the image frame data; extract audio features corresponding to the video from the audio data; and extract title features corresponding to the video from the title data.
[0114] In some exemplary embodiments, the multiple modalities include at least two of image, audio, and text.
[0115] Since Figure 6 Each functional module of the multi-modal feature fusion device in the exemplary embodiments corresponds to the steps of the exemplary embodiments of the above multi-modal feature fusion method. Therefore, for details not disclosed in the device embodiments of the present disclosure, please refer to the embodiments of the above multi-modal feature fusion method of the present disclosure.
[0116] Exemplary Computing Device
[0117] After introducing the methods, media, and devices of the exemplary embodiments of the present disclosure, next, a computing device according to another exemplary embodiment of the present disclosure is introduced.
[0118] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0119] In some possible embodiments, a computing device according to an embodiment of the present disclosure may at least include at least one processor and at least one memory. Wherein, the memory stores program code, and when the program code is executed by the processor, the processor executes the steps in the multi-modal feature fusion method according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processor may execute as Figure 2The steps shown in: Step S210, extract the features of each modality among multiple modalities of the multimedia resource; Step S220, combine the features of each modality in the modality dimension to generate multi-channel features, where each channel of the multi-channel features corresponds to a modality; Step S230, perform convolution processing on the multi-channel features to generate fused features corresponding to the multi-channel features. For another example, the processor may also execute as Figure 3 the steps shown in.
[0120] Next, refer to Figure 7 to describe the electronic device 700 according to an exemplary embodiment of the present disclosure. Figure 7 The electronic device 700 shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0121] As Figure 7 shown, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include but are not limited to: the at least one processing unit 710 described above, the at least one storage unit 720 described above, and a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710).
[0122] The bus 730 includes a data bus, an address bus, and a control bus.
[0123] The storage unit 720 may include a readable medium in the form of volatile memory, such as RAM (Random Access Memory) 721 and / or cache memory 722, and may further include ROM (Read-Only Memory) 723.
[0124] The storage unit 720 may further include a program / utility 725 having a set (at least one) of program modules 724. Such program modules 724 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.
[0125] The electronic device 700 can also communicate with one or more external devices 740 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and such communication can be carried out through the input / output (I / O) interface 750. Moreover, the electronic device 700 can also communicate with one or more networks (such as a local area network, a wide area network, and / or a public network, such as the Internet) through the network adapter 760. As shown in the figure, the network adapter 760 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID (Redundant Arrays of Independent Disks) systems, tape drives, and data backup storage systems, etc.
[0126] It should be noted that, although several units or subunits of the multimodal feature fusion device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0127] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0128] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for the convenience of expression. The present disclosure aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A multimodal feature fusion method, characterized in that, including: extracting features of each modality among multiple modalities of the multimedia resource; combining the features of each modality in the modality dimension to generate multi-channel features, where each channel of the multi-channel features corresponds to one of the modalities; the multi-channel features are feature vectors of dimension L×D×N, and N is the number of modalities of the multiple modalities; performing convolution processing on the multi-channel features in the direction of dimension D through C1 first convolutional kernels to generate first fusion features corresponding to the multi-channel features, the size of the first convolutional kernels being determined according to the value of dimension D, and the dimension of the first fusion features being L×D×C1; performing convolution processing on the first fusion features in the direction of dimension D through C2 second convolutional kernels to generate second fusion features corresponding to the multi-channel features, the dimension of the second fusion features being L×D×C2, and the size of the second convolutional kernels being smaller than the size of the first convolutional kernels; performing stretching processing on the second fusion features in the modality dimension to generate corresponding third fusion features, and the third fusion features having the same modality dimension as the features of each modality among the multiple modalities.
2. The method according to claim 1, wherein The method further includes: after extracting the features of each modality, performing feature aggregation on the features of each modality based on a predetermined grouping; generating feature vectors of each modality having the same dimension based on the result of the feature aggregation.
3. The method according to claim 2, characterized in that, The performing feature aggregation on the features of each modality based on a predetermined grouping includes: adopting a Nextvlad model to perform feature aggregation on the features of different dimensions of each modality based on a predetermined grouping.
4. The method according to claim 1, wherein The multimedia resource includes image frame data, audio data, and text data of a video, and the extracting features of multiple modalities of the multimedia resource includes: extracting image features corresponding to the video from the image frame data; extracting audio features corresponding to the video from the audio data; and extracting text features corresponding to the video from the text data.
5. The method according to any one of claims 1 to 3, characterized in that, The multiple modalities include at least two of an image modality, an audio modality, and a text modality.
6. A multi-modal feature fusion device, characterized in that, including: a feature extraction module for extracting features of each modality among multiple modalities of the multimedia resource; a modality combination module for combining the features of each modality in the modality dimension to generate multi-channel features, where each channel of the multi-channel features corresponds to one of the modalities; the multi-channel features are feature vectors of dimension L×D×N, and N is the number of modalities of the multiple modalities; a convolution processing module for performing convolution processing on the multi-channel features in the direction of dimension D through C1 first convolutional kernels to generate first fusion features corresponding to the multi-channel features, the size of the first convolutional kernels being determined according to the value of dimension D, and the dimension of the first fusion features being L×D×C1; the convolution processing module is further configured to perform convolution processing on the first fusion features in the direction of dimension D through C2 second convolutional kernels to generate second fusion features corresponding to the multi-channel features, the dimension of the second fusion features being L×D×C2, and the size of the second convolutional kernels being smaller than the size of the first convolutional kernels; A stretching processing module for performing stretching processing on the second fusion feature in the modality dimension to generate a corresponding third fusion feature, where the third fusion feature has the same modality dimension as the features of each modality in the multiple modalities.
7. The device according to claim 6, wherein The device further includes: A feature aggregation module for, after extracting the features of each modality, performing feature aggregation on the features of each modality based on a predetermined grouping; A feature generation module for generating feature vectors of each modality with the same dimension based on the result of feature aggregation.
8. The device according to claim 7, characterized in that, The feature aggregation module is further configured to: Adopt a Nextvlad model to perform feature aggregation on the features of different dimensions of each modality based on a predetermined grouping.
9. The device according to claim 6, characterized in that, The multimedia resource includes image frame data, audio data, and caption data of a video, and the feature extraction module is further configured to: Extract the image features corresponding to the video from the image frame data; Extract the audio features corresponding to the video from the audio data; and Extract the caption features corresponding to the video from the caption data.
10. The device according to any one of claims 6 to 8, characterized in that The multiple modalities include at least two modalities of image, audio, and text.
11. A computing device, characterized in that, Includes: A processor and a memory, where the memory stores executable instructions, and the processor is configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 5.
12. A medium on which a program is stored, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
A socio-emotional classification method based on multimodal fusion
CN109508375A
Image classification method based on deep interleaved fusion packet convolutional network
CN110197217A