Media content processing method and device, equipment and medium
By using the automatic fusion module and gating module in the multimodal model, the correlation degree of features at different levels is automatically determined for cross-modal fusion, which solves the problems of high cost and complexity in existing technologies and realizes refined fusion and efficient feature extraction of media content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, the analysis and fusion of different modal data of media content requires a lot of experimentation and adjustment, which is costly and makes it difficult to achieve effective cross-modal data fusion.
By employing the automatic fusion module and gating module in the multimodal model, the correlation between different levels of features in each modality is automatically determined, and cross-modal fusion processing is performed to adaptively select more matching level features for fusion.
It enables refined cross-modal data fusion of media content, improves the extraction effect and efficiency of multimodal features, and reduces the complexity of design and experimentation.
Smart Images

Figure CN121744172A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data technology, and in particular to a media content processing method, apparatus, device, and medium. Background Technology
[0002] To better understand the complex and rich content of media platforms, such as videos and text / image content, analysis can be conducted to examine the relationships and semantics between different modalities of the media content. However, due to the varying levels of abstraction in different modalities of media content, analyzing and fusing this data requires extensive experimental comparisons and adjustments, necessitating the manual design of fusion strategies to achieve satisfactory results. This process is costly and requires improvement. Summary of the Invention
[0003] To address the aforementioned technical problems, this disclosure provides a media content processing method, apparatus, device, and medium.
[0004] This disclosure provides a media content processing method, the method comprising:
[0005] Acquire multimodal data of the target media content;
[0006] The multimodal data of the target media content is input into the multimodal model, and multiple hierarchical features corresponding to each modality are extracted.
[0007] The automatic fusion module in the multimodal model is used to determine the correlation degree between different level features of each modality when fusion is performed based on multiple level features and level information of each modality, and the level fusion features of each modality at each level are determined based on the correlation degree.
[0008] The gating module in the multimodal model is used to perform fusion processing on the hierarchical fusion features of each submodality at each level to determine the submodal fusion features at each level.
[0009] The primary modal features and secondary modal features of each level are fused to obtain multimodal features of each level. The target features of the target media content are determined based on the highest level multimodal features and the highest level features of each modality.
[0010] This disclosure also provides a media content processing apparatus, the apparatus comprising:
[0011] The acquisition module is used to acquire multimodal data of the target media content;
[0012] The feature module is used to input the multimodal data of the target media content into the multimodal model and extract multiple hierarchical features corresponding to each modality;
[0013] The hierarchical module is used to utilize the automatic fusion module in the multimodal model to determine the correlation degree between different hierarchical features of each modality when fusion is performed based on multiple hierarchical features and hierarchical information of each modality, and to determine the hierarchical fusion features of each modality at each level based on the correlation degree.
[0014] The first fusion module is used to perform fusion processing on the hierarchical fusion features of each submodality at each level using the gating module in the multimodal model, and to determine the submodal fusion features at each level.
[0015] The second fusion module is used to fuse the main modal features and submodal fusion features of each level to obtain multimodal features of each level, and to determine the target features of the target media content based on the highest level multimodal features and the highest level features of each modality.
[0016] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the media content processing method provided in this disclosure.
[0017] This disclosure also provides a computer-readable storage medium storing a computer program for performing the media content processing method provided in this disclosure.
[0018] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: The media content processing solution provided in this disclosure acquires multimodal data of the target media content; inputs the multimodal data of the target media content into a multimodal model to extract multiple hierarchical features corresponding to each modality; utilizes the automatic fusion module in the multimodal model to determine the correlation degree between different hierarchical features of each modality when fusion, based on the multiple hierarchical features and hierarchical information of each modality, and determines the hierarchical fusion features of each modality at each level based on the correlation degree; utilizes the gating module in the multimodal model to perform fusion processing on the hierarchical fusion features of each modality at each level to determine the submodal fusion features of each level; fuses the primary modality features and submodal fusion features of each level to obtain the multimodal features of each level, and determines the target features of the target media content based on the highest level multimodal features and the highest level features of each modality. By employing the above technical solution, the automatic fusion module of the multimodal model can automatically determine the correlation between different levels of features in each modality when fusing multimodal data of the target media content. Based on this correlation, the hierarchical fusion features of each modality at each level are determined. Subsequently, the gating module is used to perform cross-modal fusion processing of the hierarchical fusion features of different modalities, which are then fused with the main modality features to obtain multimodal features at different levels, thereby obtaining the target features of the target media content. This not only enables refined fusion of cross-modal data of media content, but also considers the establishment of connections between hierarchical information and features at different levels during the fusion process. It achieves adaptive selection of more matching hierarchical features for cross-modal fusion, avoiding complicated design and experiments, and effectively improving the extraction effect and efficiency of multimodal features of media content. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 A schematic flowchart illustrating a media content processing method provided in an embodiment of this disclosure;
[0021] Figure 2 A flowchart illustrating another media content processing method provided in this embodiment of the disclosure;
[0022] Figure 3 This is a schematic diagram of the structure of a multimodal model provided in an embodiment of the present disclosure;
[0023] Figure 4 A schematic diagram illustrating the execution process of an automatic fusion module provided in an embodiment of this disclosure;
[0024] Figure 5 This is a schematic diagram of the structure of a media content processing device provided in an embodiment of the present disclosure;
[0025] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0026] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0027] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0028] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0029] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0030] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0031] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0032] In related technologies, the analysis and processing of multimodal data are mainly divided into the following three types: 1. Early-fusion: Early-fusion refers to the interaction of raw data from different modalities before the processing stage of different modal data. However, the data types of different modalities are quite different, the fusion method is complex, and there is overlap and redundancy between information, which negatively affects the model's performance and generalization ability. 2. Late-fusion: Late-fusion refers to the merging of feature vectors from different modalities after feature extraction and processing. However, the disadvantage is that fusion is only performed after feature extraction, which cannot fully utilize the correlation information between different modalities, and some important features between modalities are ignored. 3. Intermediate-fusion: Intermediate-fusion refers to the fusion of features from different modalities during the feature extraction process. The features are fused at the intermediate layer or intermediate representation stage. However, due to the different levels of abstraction of different modal data in media content, intermediate-fusion requires the selection of appropriate fusion strategies and the trade-off between modalities. It requires a lot of experimentation and adjustment, which increases the complexity of design and optimization and is costly.
[0033] To address the aforementioned problems, this disclosure provides a media content processing method, which will be described below with reference to specific embodiments.
[0034] Figure 1 This is a flowchart illustrating a media content processing method provided in an embodiment of the present disclosure. The method can be executed by a media content processing device, which can be implemented using software and / or hardware and is generally integrated into an electronic device.
[0035] like Figure 1 As shown, the method includes:
[0036] Step 101: Obtain multimodal data of the target media content.
[0037] The target media content may include videos and / or images / text that require analysis and processing. There may be one or more target media content items; for example, the target media content may include multiple short videos from a video application that need to be analyzed and processed. Analyzing these short videos can lead to application scenarios such as video understanding, video recommendation, title generation, and subtitle generation. Multimodal data may include data with multiple modalities, which refers to various types of data modalities, including images, audio, and text. In this embodiment, multimodal data may include image data, text data, and audio data.
[0038] Specifically, the media content processing device can acquire multimodal data of the target media content. For example, when the target media content is a target video, the device can perform frame extraction processing on the target video to obtain image data, and can extract text data such as the title, subtitles, and description of the target video, as well as obtain the corresponding audio data of the target video. The image data, text data, and audio data of the target video can be combined to obtain multimodal data.
[0039] Step 102: Input the multimodal data of the target media content into the multimodal model and extract multiple hierarchical features corresponding to each modality.
[0040] A multimodal model can be a model that analyzes and processes media content to extract its multimodal features. This multimodal model can handle data from multiple modalities and better understand the media content by analyzing the correlations and semantics between different modalities. Hierarchical features can be features at a specific level, where level refers to the specific hierarchy when extracting features from data step by step. There can be multiple levels, ranging from low to high. The higher the level, the higher the dimensionality. Low-level features include more superficial semantic features of the current data, while high-level features include deeper and more abstract features.
[0041] After acquiring multimodal data of the target media content, the media content processing device can input the multimodal data into a multimodal model and use the feature extraction model in the multimodal model to extract features at multiple levels for each modality, thereby obtaining corresponding multi-level features. The input of each level is the output of the previous level, and multiple levels of features are extracted for each modality.
[0042] Step 103: Using the automatic fusion module in the multimodal model, determine the correlation degree between different level features of each modality when fusion is performed based on the multiple level features and level information of each modality, and determine the level fusion features of each modality at each level based on the correlation degree.
[0043] The automatic fusion module can be a module that learns the relationship between features at different levels during cross-modal fusion based on the self-attention mechanism, and is one module in a multimodal model. There are multiple automatic fusion modules, each corresponding to a level setting of a submodality. The submodality is a modality other than the primary modality among the multiple modalities mentioned above. In this embodiment of the disclosure, since the analysis is performed on media content such as video and / or text and images, the primary modality can be set as image, and the submodality includes text and / or audio.
[0044] Hierarchical fusion features can be features representing each level of each modality, generated by weighting and fusing features from other levels based on their correlation with each other. Each level's hierarchical fusion feature can take other level features into account. Hierarchical information can be identifiers representing which specific level it is; for example, layer n can represent the hierarchical information of the nth layer.
[0045] For example, Figure 2 A flowchart illustrating another media content processing method provided in this disclosure embodiment is shown below. Figure 2 As shown, in one feasible implementation, step 103 above may include the following steps:
[0046] Step 201: For each automatic fusion module, input the multiple hierarchical features and hierarchical information of the corresponding submodal of the automatic fusion module as input data.
[0047] A hierarchical feature includes multiple unit features. A unit feature can be a feature at a single position that makes up a hierarchical feature. A hierarchical feature includes multiple unit features with a sequential relationship. For example, for image data, a hierarchical feature can include features of multiple image frames corresponding to different time points, and the feature of an image frame at a time point is a unit feature. Similarly, for text data, a hierarchical feature can include features of multiple characters with a sequential arrangement, and the feature of a character at a single position is a unit feature.
[0048] After the media content processing device inputs the multimodal data of the target media content into the multimodal model to obtain the multi-layer features corresponding to each modality, it uses each automatic fusion module to perform fusion calculations on each level of the submodality. Specifically, the multi-level features of the submodality corresponding to each automatic fusion module and the level information of each level can be input into the automatic fusion module as input data.
[0049] Step 202: Using the hierarchical position encoder of the automatic fusion module, based on the original position encoding information, feature position, and hierarchical information of each unit feature in this sub-modality, determine the hierarchical position encoding information corresponding to each unit feature.
[0050] The layer position embedding unit is a functional module that determines the layer position encoding information of each unit feature. Layer position encoding information provides unique positional and hierarchical encoding for each unit feature, enabling it to be distinguished from unit features of other layers and positions during the training of the multimodal model. Layer position encoding information is obtained by adding layer information to the original position encoding information of a unit feature, and the original position encoding information refers to the basis information of the original relative position encoding of the unit feature. The feature position can be the specific location of the unit feature within its respective layer of features. For example, if the unit feature is an image frame, the feature position is the time point of that image frame, such as the second.
[0051] Specifically, for each automatic fusion module's hierarchical position encoder, the media content processing device can input the original position encoding information, feature position, and hierarchical information of each unit feature of the corresponding submodality of the automatic fusion module, encode the hierarchical information to obtain new hierarchical encoding information, and calculate the new hierarchical encoding information and the original position encoding information through hyperparameters learned during model training to determine the hierarchical position encoding information of each unit feature.
[0052] For example, the hierarchical position encoding information of each unit feature can be represented as PE layer (pos, l) = α*PE new (l)+(1-α)*PE old (pos), where PE layer (pos, l) represents the hierarchical position encoding information, α represents the hyperparameter, characterizing the weight of the hierarchical information, which has been determined through training, pos represents the feature position of a unit feature, l represents the hierarchical information of the unit feature, PE new (l) represents hierarchical coding information based on hierarchical information, which is only related to hierarchical information. The specific coding parameters have been obtained through training. PE old (pos) represents the original positional encoding information, which is only related to the feature position. The above formula for hierarchical positional encoding information can be used to combine the new hierarchical encoding with the old positional encoding to generate hierarchical positional encoding information. This can effectively utilize the pre-trained weights of the original model and is compatible with the original encoding, thus accelerating the training convergence speed. Initially, the encoding parameters need to remain unchanged. Once the model training reaches a stable state, learning updates are allowed.
[0053] Step 203: Using the self-attention mechanism module of the automatic fusion module, based on the correlation parameter, the multiple unit features included in each level of the submodal features, and the corresponding multiple level position encoding information, determine the correlation degree between different level features of the submodal during fusion, and determine the level fusion feature of the target level of the submodal corresponding to the automatic fusion module according to the correlation degree.
[0054] The self-attention mechanism module can be a neural network module that automatically learns the dependencies between internal data when processing sequence data. In this embodiment, the self-attention mechanism module automatically learns the correlation between different levels of features in a submodal by introducing additional hierarchical information into the parameter learning. The correlation parameter can be a parameter in the trained attention mechanism used to calculate the correlation between different levels of features, and has been learned and determined. The hierarchical fusion feature can be a feature representing a level that is generated by weighting and fusing other hierarchical features based on the correlation between that hierarchical feature and other hierarchical features for each level of each modality.
[0055] The target level can be the specific level of the submodality corresponding to the current automatic fusion module. The correlation degree between features at different levels of the submodality during fusion can include the sub-correlation degree between each unit feature of the submodality at the target level and the corresponding unit features at other levels. Other levels refer to the levels other than the target level among multiple levels. The sub-correlation degree indicates the importance of the unit features at other levels when the unit features of the target level are fused with the features of the corresponding level of the main modality. The larger the sub-correlation degree, the greater the importance of the representation. For a unit feature of a target level, the sum of the multiple sub-correlation degrees between this unit feature and the units at other corresponding levels is 1.
[0056] The media content processing device can take the multiple unit features of each level of the submodal features of the automatic fusion module and the hierarchical position encoding information corresponding to each unit feature as input data for the self-attention mechanism module of each automatic fusion module. For each unit feature of the target level of the submodal, the sub-association degree between it and the corresponding unit features of other levels is calculated based on the correlation degree parameter. The unit features of other levels are weighted and fused according to the sub-association degree to generate the representation information of the unit feature. The representation information of each unit feature of the target level is combined to obtain the corresponding hierarchical fusion feature.
[0057] Optionally, determining the hierarchical fusion features of the target level of the submodality corresponding to the automatic fusion module based on the correlation degree may include: for the target level of the submodality corresponding to the automatic fusion module, performing a weighted summation based on multiple unit features included in each level feature, as well as the hierarchical position encoding information and sub-correlation degree corresponding to each unit feature, to obtain the corresponding hierarchical fusion features.
[0058] For example, the hierarchical fusion feature for the target level of the corresponding submodality of each automatic fusion module can be represented as:
[0059]
[0060] Where F represents the hierarchical fusion feature of the target level corresponding to the submodal of the current automatic fusion module, and TransformerEncoder m This represents a cross-level transform fusion encoder used to encode the hierarchical fusion features of each level. This encoder converts each unit feature into three distinct vectors: a query vector, a key vector, and a value vector. Furthermore, it learns the correlation between features from different levels within the submodal when fused, including the sub-correlation degree of each unit feature relative to the target level. PE represents the unit feature at the nth feature position of the mth level. layer (m, n) represents The corresponding hierarchical position encoding, through the weighted summation calculation of the above formula, can yield the hierarchical fusion feature of the target level.
[0061] The above scheme employs a self-attention mechanism and introduces hierarchical information during learning, enabling each unit feature at a position to simultaneously consider features at that position in other levels. It adaptively selects cross-modal features from more matching levels for fusion and suppresses cross-modal features from mismatched levels, thereby better capturing the dependencies between contexts and maximizing the understanding of cross-modal data.
[0062] Step 104: Use the gating module in the multimodal model to perform fusion processing on the hierarchical fusion features of each submodality at each level, and determine the submodal fusion features at each level.
[0063] The gating module can be a module that considers the attention of different modalities and features of the previous level when performing cross-modal feature fusion for multiple modal features of submodalities at various levels. In this embodiment, there are multiple gating modules, each corresponding to a level setting, that is, one gating module is set for each level.
[0064] Submodal fusion features can be obtained by fusing the hierarchical fusion features of all submodals at various levels in multiple modalities. These features are only related to the features of the submodals and are not related to the features of the main modality.
[0065] Specifically, after determining the hierarchical fusion features of each level of the submodality, the media content processing device can use each gating module to perform fusion calculation processing on all hierarchical fusion features of the submodality at the corresponding level and the submodal fusion features of the previous level. For example, it can perform fusion calculation processing according to the attention of different submodalities and the attention of the previous level as preset, to obtain the submodal fusion features of each level.
[0066] In some embodiments, the gating module in the multimodal model is used to fuse the hierarchical fusion features of each submodality at each level to determine the submodal fusion features at each level. This can include: using each gating module, inputting the hierarchical fusion features of the current level of the submodality and the previous level of the previous level of the current level as input data; using the importance formula of the gating module, based on the hierarchical fusion features of the current level of the submodality, the weight matrix of the current level of the submodality, the weight matrix of the current level, the previous level of the fusion features, and the bias vector, determining the current importance of the submodality at the current level and the previous level of the previous level of the current level; and performing a weighted calculation based on the weighted fusion features of the previous level of the fusion features, the previous level of the fusion features, the weighted fusion features of the current level of the submodality, and the current importance to determine the current submodal fusion feature of the current level corresponding to the gating module.
[0067] The importance formula can be a formula that considers the attention given to data from different modalities. In this embodiment, the importance formula uses the Sigmoid function as an example. The Sigmoid function is a mathematical function of type curve function, usually used to map input values to output values between 0 and 1. The weight matrix of the current layer in the submodality can represent the importance of a feature of a submodality when performing feature fusion at the current layer. When there are multiple submodities, there are multiple weight matrices. The weight matrix of the current layer can represent the importance of the fused features of the submodality at the current layer when performing feature fusion at the current layer. The bias vector can be understood as a constant. The weight matrix and bias vector mentioned above have been learned during the training of the multimodal model. The current importance can be the calculated importance of the submodality at the current layer, and the previous importance can be the calculated importance of the layer above the current layer. The sum of the current importance and the previous importance is 1.
[0068] Specifically, for each gating module, the media content processing device can input the hierarchical fusion features of the current level in each submodality and the previous modal fusion features of the previous level. If the current level is the first level, the previous modal fusion features can be the hierarchical fusion features of each submodality of the current level itself. The importance of the gating module is calculated using its importance formula. For example, if the submodality includes text and audio, the importance formula can be expressed as z(t) = sigmoid(W). a *x a (t)+W b *x b (t)+U h *h(t-1)+b z ), where z(t) represents the current importance of the submodality at the current level, W a x represents the weight matrix of the current level t in the text modality. a (t) represents the hierarchical fusion feature of the current level t in the text modality, W b x represents the weight matrix of the current level t in the audio modality. b (t) represents the hierarchical fusion feature of the current level t in the audio modality, U h Let h(t-1) represent the weight matrix of the current level t, and let b represent the submodal fusion features of the previous level t-1. z Let represent the bias vector; then the submodal fusion feature h(t) of the current level t can be calculated by the following formula: h(t)=(1-z(t))⊙h(t-1)+z(t)⊙h′(t) ⊙ represents the element-wise multiplication operation, 1-z(t) represents the importance of the previous level of the current level, and h′(t) represents the weighted fusion feature of the submodal fusion feature of the current level, h′(t)=W a *x a (t)+W b *x b (t).
[0069] Different modalities receive different levels of attention during cross-modal fusion at different levels. For example, audio, as a low-level feature, has significant noise and redundant information; text, on the other hand, has a certain degree of abstraction and is advantageous in fusing more abstract features at higher levels. This embodiment of the present disclosure can determine the importance of different modalities during cross-modal fusion at different levels by setting specific gating modules for the input of multimodal features across different layers, thereby fusing the features of submodalities and effectively improving the accuracy of cross-modal feature fusion.
[0070] Step 105: Fuse the primary modal features and secondary modal features of each level to obtain multimodal features of each level, and determine the target features of the target media content based on the highest level multimodal features and the highest level features of each modality.
[0071] Multimodal features can be obtained by fusing all features of multiple modalities included in each level, as well as the multimodal features of the previous level. The multimodal features of each level are obtained by serial fusion, that is, by stacking the features of different modalities layer by layer. Target features can be a high-level fused feature representation obtained by fusing the cross-modal features of the final target media content.
[0072] Specifically, after the media content processing device determines the submodal fusion features of each level, it can fuse the submodal fusion features and main modal features of each level with the multimodal features of the previous level to obtain the multimodal features of that level. The multimodal features of each level are serially fused with the multimodal features of the previous level. Then, the highest level multimodal features and the highest level features of each modality can be extracted. After the target fusion operation, the target features can be obtained. The target fusion operation can include at least one of the following: splicing operation, fully connected layer fusion, and activation function. The specific operation can be determined according to business requirements.
[0073] For example, target features can be obtained using `F_fusion1 = Concatenate([F_f, F_text, F_image, F_audio])`, where `Concatenate` represents the concatenation operation, `F_fusion1` represents the target features obtained from the concatenation operation, `F_f` represents the highest-level multimodal features, `F_text` represents the highest-level text modality features, `F_image` represents the highest-level image modality features, and `F_audio` represents the highest-level audio modality features. Target features can also be further fused using a fully connected layer `F_fusion` on top of the concatenation operation. 2 = Dense(F_fusion1) is obtained, where Dense represents a fully connected layer, F_fusion1 represents the features obtained by the above concatenation operation, and F_fusion2 represents the target fusion operation, which is the target feature obtained by fusing the concatenation operation and the fully connected layer. The target feature can also be obtained by adding an activation function F_fusion3 = Activation(F_fusion2) on the basis of the above concatenation operation and fully connected layer fusion. F_fusion3 represents the target fusion operation, which is the target feature obtained by concatenation operation, fully connected layer fusion and activation function, and F_fusion2 represents the feature obtained by the above fully connected layer fusion.
[0074] The media content processing scheme provided in this embodiment acquires multimodal data of target media content; inputs the multimodal data of target media content into a multimodal model to extract multiple hierarchical features corresponding to each modality; utilizes the automatic fusion module in the multimodal model to determine the correlation degree between different hierarchical features of each submodality when fusion, based on the multiple hierarchical features and hierarchical information of each submodality, and determines the hierarchical fusion features of each submodality at each level based on the correlation degree; utilizes the gating module in the multimodal model to perform fusion processing on the hierarchical fusion features of each submodality at each level to determine the submodal fusion features of each level; fuses the primary modality features and submodal fusion features of each level to obtain the multimodal features of each level, and determines the target features of the target media content based on the highest level multimodal features and the highest level features of each modality. By employing the above technical solution, the automatic fusion module of the multimodal model can automatically determine the correlation between different levels of features in each modality when fusing multimodal data of the target media content. Based on this correlation, the hierarchical fusion features of each modality at each level are determined. Subsequently, the gating module is used to perform cross-modal fusion processing of the hierarchical fusion features of different modalities, which are then fused with the main modality features to obtain multimodal features at different levels, thereby obtaining the target features of the target media content. This not only enables refined fusion of cross-modal data of media content, but also considers the establishment of connections between hierarchical information and features at different levels during the fusion process. It achieves adaptive selection of more matching hierarchical features for cross-modal fusion, avoiding complicated design and experiments, and effectively improving the extraction effect and efficiency of multimodal features of media content.
[0075] In some embodiments, the media content processing method may further include: performing target operations based on target features of the target media content, wherein the target operations include at least one of the following: media content recommendation and media content classification. Here, target operations refer to subsequent business operations performed on the target media content, and may include at least one of media content recommendation, media content classification, etc., without specific limitation. After determining the target features of the target media content, the media content processing device can input the target features into a subsequent target operation model for target operations. Since the target features include features at various levels and multiple modalities of the target media content, the understanding of the media content is improved, thereby enhancing the accuracy of subsequent target operations.
[0076] The media content processing method of this disclosure embodiment will be further described below through specific examples. For example, Figure 3 This is a schematic diagram of the structure of a multimodal model provided in an embodiment of the present disclosure, such as... Figure 3As shown in the figure, taking target media content as video, primary modality as image, and secondary modalities including text and audio as an example, the multimodal model can include a feature extraction module, an automatic fusion module, a gating module, and a high-dimensional feature fusion module. The feature extraction and high-dimensional feature fusion modules are not shown in the figure. There are multiple automatic fusion modules, each corresponding to a level setting of a secondary modality. There are also multiple gating modules, each corresponding to a level setting. Multiple feature extraction modules extract multiple levels of features from image, text, and audio data in the multimodal data. N in the figure represents the number of levels. The cross-modal automatic fusion module establishes interdependent attention relationships between features at multiple levels, thereby achieving information interaction and fusion. The gating module, by learning adjustable gating signals, allows the network to selectively filter and pass features from other modalities. The high-dimensional feature fusion module assigns weights to high-level features of different modalities through model learning, and then sums them weighted to obtain the fused feature representation.
[0077] use Figure 3 Multimodal models can incorporate semantic information from other modalities into the high, medium, and low-level features extracted from images. They can automatically learn the weights (attention) at different levels between different modalities, thereby increasing the weight of important semantic information in different modalities while decreasing the weight of unimportant semantic information, thus suppressing the interference of redundant and invalid information on model learning.
[0078] For example, Figure 4 A schematic diagram illustrating the execution process of an automatic fusion module provided in this embodiment of the disclosure, as shown below. Figure 4As shown in the figure, the execution process of an automatic fusion module can be illustrated. The automatic fusion module can include a hierarchical position encoder (not shown in the figure) and a self-attention mechanism module. The input data of the automatic fusion module can include multiple hierarchical features and hierarchical information of the corresponding submodality. Each hierarchical feature includes multiple unit features. The hierarchical position encoder is used to determine the hierarchical position encoding information corresponding to each unit feature based on the original position encoding information, feature position, and hierarchical information of each unit feature of the submodality. The self-attention mechanism module is used to calculate the sub-association degree between each unit feature of the target level and the corresponding unit features of other levels, and to generate the representation information of the unit feature by weighted fusion of the corresponding positions of other levels according to the sub-association degree. Each segment in the figure represents a representation information. After combining the representation information of each unit feature of the target level, the corresponding hierarchical fusion feature is obtained by dimensionality reduction processing through an average pooling layer. By using additional hierarchical positional encoding, cross-modal hierarchical information is introduced into the learning of model parameters. This allows each position to simultaneously consider information from other layer positions, enabling the model to learn and fuse cross-modal features from semantically more matching layers through samples, while suppressing cross-modal features from mismatched layers, thereby better capturing the dependencies between contexts.
[0079] This solution provides a multimodal fusion processing approach for media content. Based on a large-scale multimodal dataset, it performs self-supervised training and learns the correlation and importance between different modalities to obtain the parameters of the automatic fusion module. It eliminates the need for manual design or specification of fusion strategies and extensive parameter search experiments, allowing the model to automatically select the most suitable fusion strategy and trade-off modalities. This significantly reduces the cost of manually designing fusion strategies and conducting numerous parameter search experiments, while improving the performance of media content understanding. By fusing information from different modalities, the model can obtain comprehensive semantic information from multiple angles and sources during feature extraction at different stages, thereby improving its ability to understand and process tasks.
[0080] The innovation of this scheme lies not only in its ability to achieve refined fusion between different modalities and establish interdependencies, but more importantly, in its ability to automatically establish connections between different levels using a unified automatic fusion module, eliminating the need for additional complex inter-layer information fusion structures. It highlights the importance of other modalities at different levels during fusion through additional hierarchical positional encoding and attention mechanisms. During the learning process at different levels of the model, the importance of high-level cross-modal information differs from that of low-level cross-modal information. The proposed structure can automatically select and retain shallow information during model training, avoiding information loss. It can also utilize high-dimensional semantic information to guide learning, maintaining the advantages of post-fusion.
[0081] Figure 5This is a schematic diagram of a media content processing apparatus provided in an embodiment of the present disclosure. The apparatus can be implemented by software and / or hardware and is generally integrated into an electronic device. Figure 5 As shown, the device includes:
[0082] The acquisition module 501 is used to acquire multimodal data of the target media content;
[0083] Feature module 502 is used to input the multimodal data of the target media content into the multimodal model and extract multiple hierarchical features corresponding to each modality;
[0084] The hierarchical module 503 is used to utilize the automatic fusion module in the multimodal model to determine the correlation degree between different hierarchical features of each modality when fusion is performed based on multiple hierarchical features and hierarchical information of each modality, and to determine the hierarchical fusion features of each modality at each level based on the correlation degree.
[0085] The first fusion module 504 is used to perform fusion processing on the hierarchical fusion features of each submodality at each level using the gating module in the multimodal model, and to determine the submodal fusion features at each level.
[0086] The second fusion module 505 is used to fuse the main modal features and submodal fusion features of each level to obtain multimodal features of each level, and to determine the target features of the target media content based on the highest level multimodal features and the highest level features of each modality.
[0087] Optionally, there may be multiple automatic fusion modules, with each automatic fusion module corresponding to a submodal and a hierarchical setting.
[0088] Optionally, hierarchical module 503 includes:
[0089] The first unit is used to input multiple hierarchical features and hierarchical information of the corresponding submodality of each automatic fusion module as input data, wherein one hierarchical feature includes multiple unit features;
[0090] The second unit is used to determine the hierarchical position encoding information corresponding to each unit feature based on the original position encoding information, feature position, and hierarchical information of each unit feature in the submodal using the hierarchical position encoder of the automatic fusion module.
[0091] The third unit is used to utilize the self-attention mechanism module of the automatic fusion module to determine the correlation degree between different levels of features of the submodality when fusion, based on the correlation degree parameter, the multiple unit features included in each level of the submodality, and the corresponding multiple level position encoding information, and to determine the level fusion feature of the target level of the submodality corresponding to the automatic fusion module according to the correlation degree.
[0092] Optionally, the correlation degree between the features of different levels of the submodality during fusion includes the sub-correlation degree between each unit feature of the submodality at the target level and the unit features at the corresponding positions of other levels, wherein the other levels are the levels other than the target level among the plurality of levels.
[0093] Optionally, the third unit is used for:
[0094] For the target level of the submodal corresponding to the automatic fusion module, the corresponding hierarchical fusion feature is obtained by weighted summation based on the multiple unit features included in each level feature, the hierarchical position encoding information corresponding to each unit feature and the sub-association degree.
[0095] Optionally, there may be multiple gate control modules, with each gate control module corresponding to a hierarchical setting;
[0096] The first fusion module 504 is used for:
[0097] Using each of the gating modules, the hierarchical fusion features of the current level in the next modality corresponding to the gating module and the previous modality fusion features of the previous level of the current level are used as input data.
[0098] Using the importance formula of the gating module, based on the hierarchical fusion features of the current level in the submodal, the weight matrix of the current level in the submodal, the weight matrix of the current level, the previous mode fusion features, and the bias vector, the current importance of the submodal of the current level and the previous importance of the previous level are determined. Then, based on the weighted fusion features of the previous mode fusion features, the previous importance, the weighted fusion features of the hierarchical fusion features of the current level in the submodal, and the current importance, a weighted calculation is performed to determine the current submodal fusion features of the current level corresponding to the gating module.
[0099] Optionally, the multimodal data includes image data, text data, and audio data, with the primary modality being images and the secondary modalities including text and / or audio.
[0100] The media content processing apparatus provided in this disclosure can execute the media content processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0101] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the media content processing method provided in any embodiment of this disclosure.
[0102] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present disclosure. See below for details. Figure 6 The diagram illustrates a structural schematic suitable for implementing the electronic device 600 in the embodiments of this disclosure. The electronic device 600 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0103] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0104] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0105] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined above in the media content processing method of embodiments of this disclosure.
[0106] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0107] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0108] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0109] The aforementioned computer-readable medium carries one or more programs. When the aforementioned one or more programs are executed by the electronic device, the electronic device causes the following: to acquire multimodal data of target media content; to input the multimodal data of the target media content into a multimodal model, and to extract multiple hierarchical features corresponding to each modality; to use the automatic fusion module in the multimodal model to determine the correlation degree between different hierarchical features of each submodality when fusion is performed based on the multiple hierarchical features and hierarchical information of each submodality, and to determine the hierarchical fusion features of each submodality at each level based on the correlation degree; to use the gating module in the multimodal model to perform fusion processing on the hierarchical fusion features of each submodality at each level, and to determine the submodal fusion features of each level; to fuse the primary modality features and submodal fusion features of each level to obtain the multimodal features of each level, and to determine the target features of the target media content based on the highest level multimodal features and the highest level features of each modality.
[0110] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0112] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0113] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0114] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0115] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0116] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0117] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0118] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A media content processing method, characterized in that, include: Acquire multimodal data of the target media content; The multimodal data of the target media content is input into the multimodal model, and multiple hierarchical features corresponding to each modality are extracted. The automatic fusion module in the multimodal model is used to determine the correlation degree between different level features of each modality when fusion is performed based on multiple level features and level information of each modality, and the level fusion features of each modality at each level are determined based on the correlation degree. The gating module in the multimodal model is used to perform fusion processing on the hierarchical fusion features of each submodality at each level to determine the submodal fusion features at each level. The primary modal features and secondary modal features of each level are fused to obtain multimodal features of each level. The target features of the target media content are determined based on the highest level multimodal features and the highest level features of each modality.
2. The method according to claim 1, characterized in that, The number of automatic fusion modules is multiple, and each automatic fusion module corresponds to a submodal and a hierarchical setting.
3. The method according to claim 2, characterized in that, The automatic fusion module in the multimodal model determines the correlation degree between different levels of features of each modality when fusion is performed based on multiple hierarchical features and hierarchical information of each modality, and determines the hierarchical fusion features of each modality at each level based on the correlation degree, including: For each of the automatic fusion modules, multiple hierarchical features and hierarchical information of the corresponding submodality of the automatic fusion module are input as input data, wherein one hierarchical feature includes multiple unit features; Using the hierarchical position encoder of the automatic fusion module, the hierarchical position encoding information corresponding to each unit feature is determined based on the original position encoding information, feature position, and hierarchical information of each unit feature in the sub-modality; Using the self-attention mechanism module of the automatic fusion module, based on the correlation parameter, the multiple unit features included in each level feature of the submodality, and the corresponding multiple level position encoding information, the correlation degree between different level features of the submodality is determined, and the level fusion feature of the target level of the submodality corresponding to the automatic fusion module is determined according to the correlation degree.
4. The method according to claim 3, characterized in that, The correlation degree between different levels of features of the submodality during fusion includes the sub-correlation degree between each unit feature of the submodality at the target level and the unit features at corresponding positions in other levels, wherein the other levels are the levels other than the target level among the plurality of levels.
5. The method according to claim 4, characterized in that, Based on the correlation degree, the hierarchical fusion features of the target level of the submodal corresponding to the automatic fusion module are determined, including: For the target level of the submodal corresponding to the automatic fusion module, the corresponding hierarchical fusion feature is obtained by weighted summation based on the multiple unit features included in each level feature, the hierarchical position encoding information corresponding to each unit feature and the sub-association degree.
6. The method according to claim 1, characterized in that, There are multiple gate control modules, and each gate control module corresponds to a hierarchical setting. The gating module in the multimodal model is used to fuse the hierarchical fusion features of each submodality at each level to determine the submodal fusion features at each level, including: Using each of the gating modules, the hierarchical fusion features of the current level in the next modality corresponding to the gating module and the previous modality fusion features of the previous level of the current level are used as input data. Using the importance formula of the gating module, based on the hierarchical fusion features of the current level in the submodal, the weight matrix of the current level in the submodal, the weight matrix of the current level, the previous mode fusion features, and the bias vector, the current importance of the submodal of the current level and the previous importance of the previous level are determined. Then, based on the weighted fusion features of the previous mode fusion features, the previous importance, the weighted fusion features of the hierarchical fusion features of the current level in the submodal, and the current importance, a weighted calculation is performed to determine the current submodal fusion features of the current level corresponding to the gating module.
7. The method according to any one of claims 1-6, characterized in that, The multimodal data includes image data, text data, and audio data, with the primary modality being images and the secondary modalities including text and / or audio.
8. A media content processing device, characterized in that, include: The acquisition module is used to acquire multimodal data of the target media content; The feature module is used to input the multimodal data of the target media content into the multimodal model and extract multiple hierarchical features corresponding to each modality; The hierarchical module is used to utilize the automatic fusion module in the multimodal model to determine the correlation degree between different hierarchical features of each modality when fusion is performed based on multiple hierarchical features and hierarchical information of each modality, and to determine the hierarchical fusion features of each modality at each level based on the correlation degree. The first fusion module is used to perform fusion processing on the hierarchical fusion features of each submodality at each level using the gating module in the multimodal model, and to determine the submodal fusion features at each level. The second fusion module is used to fuse the main modal features and submodal fusion features of each level to obtain multimodal features of each level, and to determine the target features of the target media content based on the highest level multimodal features and the highest level features of each modality.
9. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the media content processing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the media content processing method according to any one of claims 1-7.