Visual-Text Collaborative Summary Generation Method and System Based on Multimodal Learning

The multi-modal learning approach effectively integrates visual and textual understanding to generate coherent and accurate summaries by dynamically adjusting fusion strategies, addressing the limitations of single-modal methods.

CN119862861BActive Publication Date: 2025-07-15SHANDONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510352173.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing image description generation model is difficult to accurately capture the subtle semantic information of the image. The text summary generation model ignores the deep semantic association between the image and the text, resulting in the low quality of the generated digests. The existing multimodal digest generation method cannot dynamically adjust the information fusion strategy and cannot generate a summary matching the input text.

Method used

By introducing the visual understanding module, the text and visual understanding combination module and the summary generation module, adaptive spatiotemporal weight fusion, mixed convolution strategy, image-text comparison loss and image-text matching loss, dynamically adjust the fusion strategy between image and text to generate a summary that conforms to the context semantics.

Benefits of technology

The deep fusion of visual information and text information is achieved, the semantic richness and accuracy of the abstract is improved, semantic conflicts are avoided, and the generated abstract is more coherent and consistent in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862861B_ABST
    Figure CN119862861B_ABST
Patent Text Reader

Abstract

This application belongs to the cross - field of natural language processing, and specifically relates to a visual - text collaborative summary generation method and system based on multimodal learning, including a multimodal data receiving module for receiving multimodal input data in parallel, including text data and visual data; a visual semantic understanding module that uses a visual semantic understanding model to extract high - level semantic features of images and generate text descriptions; a semantic fusion module that uses a visual - text semantic fusion method based on direction consistency and adaptive semantic completion to fuse the original text with the generated image description at the semantic level; and a summary optimization module that uses a hybrid neural network architecture with multi - layer deep fusion to perform semantic reconstruction on the fused features and generate summary texts that conform to the context semantics and are accurately expressed. The advantages are: accurately aligning visual and text information, generating high - quality summaries, and being particularly suitable for scenarios that require cross - modal information fusion such as news reports, meeting records, video content analysis, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the cross - field of natural language processing, and specifically relates to a visual - text collaborative summary generation method and system based on multi - modal learning. Background Art

[0002] With the rapid development of artificial intelligence technology, the task of summary generation has become a hot topic in current research. However, traditional summary generation methods often rely on single - image description generation technology or text summary generation models, both of which have their own limitations. Image description generation models usually have difficulty accurately capturing the subtle semantic information in images, and the generated descriptions are often too brief or vague, lacking a deep understanding of the image content. This makes the image description unable to fully reflect the complex details of the image, thus affecting the quality of the summary. While text summary generation models perform well at the text level, they usually ignore the deep semantic relationship between images and texts, resulting in the generated summary may not be consistent enough with the image content, lacking semantic accuracy and context coherence.

[0003] To address the limitations brought by single - modality, multi - modal learning has emerged. Multi - modal learning aims to improve the model's understanding and generation ability by integrating data from different modalities (such as images, texts, videos, etc.). In visual - text tasks, images and texts, as two different types of modal information, need to be effectively fused to extract richer semantic information, thereby improving the quality of natural language generation tasks. However, existing multi - modal summary generation methods often cannot dynamically adjust the information fusion strategy between images and texts when dealing with complex situations, resulting in the inability to fully explore the potential semantic relationship between images and texts in some scenarios and unable to generate a summary that better matches the input text. Therefore, how to efficiently fuse visual and text information and improve the accuracy and semantic consistency of summary generation remains an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides a visual - text collaborative summary generation method based on multi - modal learning, aiming to effectively fuse image and text information and improve the accuracy and semantic consistency of summary generation. By introducing a visual understanding module, a text - visual understanding combination module, and a summary generation module, the present invention can generate a summary that conforms to the context semantics on the premise of ensuring no semantic conflict between the image and the text. Its technical solution is as follows:

[0005] A visual - text collaborative summary generation method based on multi - modal learning, comprising the following steps:

[0006] S1. Receive multi - modal data: including text data and visual data as the basis for subsequent summary generation;

[0007] S2. Visual semantic understanding: Input the visual data into the visual semantic understanding model. By extracting visual features, identify the key information in the image or video, and generate a text description of the content of the visual data;

[0008] S3. Combination of text and visual semantic understanding: Effectively fuse the semantic features of the text data in step S1 with the image description text or video parsing text generated in step S2 to provide semantic information for the summary generation model;

[0009] S4. Summary generation: Input the fused text in step S3 into the summary generation model. Through the model's understanding and processing of the input text, finally generate a text summary that conforms to the context semantics.

[0010] Preferably, the visual semantic understanding model in step S2 includes video key frame recognition. By calculating the importance weights of pixel regions through adaptive spatio-temporal weight fusion, identify the core changes in the video content, remove redundant information, and obtain a set of representative and semantically rich key frame images:

[0011] Adaptive spatio-temporal weight fusion: Based on the standard spatio-temporal attention, add a dynamic weight factor to adaptively adjust the attention in the time and space dimensions;

[0012] The time attention focuses on the relationship between different frames in the video sequence, reflecting the dynamic changes between frames; The time attention calculation formula is as follows:

[0013] ;

[0014] The attention in the space dimension focuses on the regional relationship within a single frame image, emphasizing the local and global structures of the object. The calculation formula is as follows:

[0015] ;

[0016] The adaptive spatio-temporal weight fusion calculation formula is as follows:

[0017] ;

[0018] is the time attention factor, learning the global importance of each frame, is the space attention factor, , are all learnable parameters, enabling the model to dynamically adjust the time and space weights. Through the sigmoid function for normalization to ensure that and do not appear negative or too large or too small.

[0019] Preferably, in step S2, the visual semantic understanding model includes single-modal visual feature encoding, which combines standard convolution and depthwise separable convolution using a hybrid convolution strategy. The single-modal visual feature encoding adopts a dual-branch structure, and global features are extracted through dilated convolution and local features are extracted through hybrid convolution via two parallel computing paths, and adaptive feature fusion is performed through dual-branch fusion.

[0020] Preferably, dilated convolution: By introducing a dynamic dilation factor strategy, dilation factors of different scales are used for multi-level information fusion, enabling the network to capture long-range pixel dependencies while maintaining local information integrity. The calculation formula is as follows:

[0021] ;

[0022] ;

[0023] where is the pixel value at position in the output feature map; represents the pixel value of the input feature at position , is the index of the convolution kernel, indicating the position of the convolution kernel in the sliding window; R represents the receptive field, the set of neighboring pixels involved in the convolution operation; represents the weight of the convolution kernel at position , is the dynamic dilation factor, which controls the sampling interval of the convolution kernel, is the local feature of the current feature map at position , and are parameters to be learned; when , it is a standard convolution, and when , there is an interval between the sampling points of the convolution kernel;

[0024] Depthwise separable convolution includes:

[0025] Pointwise convolution across channels:

[0026] ;

[0027] where represents the pixel value at position in the th channel of the input feature map, is the index of the convolution kernel, is the weight of the convolution kernel at position ;

[0028] Pointwise convolution:

[0029] ;

[0030] Among them, C is the number of input channels, is the output of per-channel convolution, is the convolution kernel, which is used for cross-channel weighted summation, and controls the number of output channels by adjusting the number of

[0031] Hybrid convolution formula:

[0032] ;

[0033] Among them, represents the output feature map of the hybrid convolution operation at position , is the balance coefficient, , which is used to control the weight between the standard convolution and the depthwise separable convolution;

[0034] Two-branch fusion formula:

[0035] ;

[0036] Among them, represents the fusion formula of the multi-scale dilated convolution and the hybrid convolution, is the dynamic fusion weight, which controls the proportion of global and local information.

[0037] Preferably, in step S2, the visual semantic understanding model includes an image-text contrast loss and an image-text matching loss;

[0038] Image-text contrast loss: By activating the single-modal visual feature encoder, aligning the visual and text feature spaces, encouraging positive sample image-text pairs to have similar representations, while negative sample pairs have a large distinction, enhancing the correlation between images and texts. The contrast loss formula is as follows:

[0039] ;

[0040] Among them, represents the similarity between the image feature and the text feature , is the logarithmic loss function, and N is the number of samples;

[0041] Image-text matching loss: By activating the image-based text encoder, learning to capture the fine-grained alignment between vision and language in the image-text multi-modal representation, using the image-text matching loss to predict whether the given image-text pair of multi-modal features matches. The matching loss formula is as follows:

[0042] ;

[0043] ;

[0044] Among them, is the label of the image-text pair, indicating that the image and the text match, indicating that the image and the text do not match, is the probability that the model predicts whether the image and the text match, and the logarithmic loss function is used to optimize the model, is the image feature and the text feature for the similarity calculation between them, is the sigmoid function, which is used to normalize the similarity value to to ensure that it can be used as a probability.

[0045] Preferably, the visual semantic understanding model in step S2 includes a language modeling loss. By activating the image-based text encoder, a text description of the given image is generated, and the likelihood of the generated text is maximized in an autoregressive manner. The calculation formula is as follows:

[0046] ;

[0047] Among them, is the probability of generating the next word under the given previous context and image information , and Q represents the length of the generated text sequence.

[0048] Preferably, the combination of text and visual semantic understanding in step S3 includes graphic-text consistency evaluation, text-visual semantic coordination detection, and context supplementation based on dynamic semantic fusion;

[0049] Graphic-text consistency evaluation: The direction consistency is used to measure the similarity between the input text and the image description text. The higher the direction consistency, the higher the similarity between the text description and the image. The calculation formula is as follows:

[0050] ;

[0051] Among them, and are the vector representations of the input text and the image description text respectively, and T is the transpose.

[0052] Text-visual semantic coordination detection: According to the similarity calculation result, it is judged whether there is a semantic conflict between the input text and the image description text. If the similarity between the two is low or there are obvious contradictions, it is marked as a conflict to avoid incorrect fusion;

[0053] Context supplementation based on dynamic semantic fusion includes keyword alignment, complementary weight calculation, and dynamic information fusion;

[0054] Keyword alignment: Extract keywords from the input text and the image description text, calculate their similarity, and evaluate the semantic alignment degree between the two;

[0055] Complementary weight calculation: Use an adaptive alignment algorithm to calculate the complementary weight between the text and the visual content. This weight measures the relative importance of the input text and the image description text, and the complementary weight is calculated through a gating mechanism, expressed as , where is the weight matrix, and are the encoded features of the input text and the image description text respectively, represents the complementary weight of the text and the image description:

[0056] ;

[0057] Dynamic information fusion: Based on the complementary weight , fuse the input text and the image description text. The fusion method is to adjust the weights of the two in proportion to generate an enhanced text representation , which contains the combination of image and text information, and its calculation formula is:

[0058] .

[0059] Preferably, in step S4, the abstract generation model includes a text encoding layer, a position encoding layer, a DeepFusion-Transformer layer, and an autoregressive decoding layer; Input the text fused in step S3 into the trained abstract generation model to generate a text abstract;

[0060] Text encoding layer: Convert each word in the text into a vector representation of a fixed dimension, and encode the semantics of each word and its context relationship, so as to obtain a high-quality representation of each word as the input of DeepFusion-Transformer;

[0061] Position encoding layer: Use sine and cosine functions to generate position encoding, so that the encoding of each position has different frequencies in different dimensions;

[0062] DeepFusion-Transformer layer: Adopt a stacked multi-layer structure, and each layer gradually learns higher-level features. The shallow Transformer layer focuses on capturing local semantic information, and the deep Transformer layer focuses on learning global semantic representations;

[0063] Autoregressive decoding layer: The decoder generates text step by step in an autoregressive manner. After generating a new word each time, the decoder updates its state, uses the generated word as the new input, and combines the current context information to continue the next generation until the complete summary text is generated.

[0064] Preferably, the formula for the autoregressive generation process is as follows:

[0065] ;

[0066] where, is the probability distribution of generating the word at the -th step given the word and input features at the -th step; is the probability distribution of generating the word at the is the current hidden state of the decoder, and are the weights and biases used to calculate the vocabulary distribution. The hidden state of the decoder is updated at each time step. Assuming the hidden state at the -th step is , then at the -th step, the new hidden state can be calculated recursively as follows:

[0067] ;

[0068] where, is the word generated at the -th step, is the hidden state at the previous moment, is the input multi-modal information.

[0069] A visual-text collaborative summary generation system based on multi-modal learning, including a multi-modal input module, a visual semantic understanding module, a text and visual understanding combination module, and a summary generation module;

[0070] Multi-modal input module: Accepts the input text data and visual data. The main function of this module is to provide the original input for subsequent visual understanding and text fusion;

[0071] Visual semantic understanding module: Inputs the visual data into the visual semantic understanding model. The model extracts visual features, identifies key information in the image or video, and generates a detailed description of the content of the visual data;

[0072] Text and visual understanding combination module: Effectively combines the input text with the image description text or video analysis text generated by the visual understanding model to obtain a fused representation of the text and image content;

[0073] Abstract Generation Module: Input the text and the text generated by the Visual Understanding and Integration Module into the generation model. Through the model's understanding and processing of the text, a text abstract that conforms to the context semantics is finally generated.

[0074] Compared with the prior art, the beneficial effects of the present application are as follows:

[0075] (1) Efficient integration of multimodal information: By introducing the Visual Understanding Module and the Text-Visual Understanding and Integration Module, the present invention effectively realizes the deep integration of visual information and text information. Compared with the traditional single-modal abstract generation method, the present invention can make full use of the complementarity of images and texts in the same abstract, improving the semantic richness and accuracy of the abstract.

[0076] (2) Dynamic semantic alignment and information supplementation: The dynamic semantic alignment algorithm proposed by the present invention can dynamically adjust the fusion strategy according to the semantic relationship between the text and the image, effectively avoiding semantic conflicts between the image description and the text content, and ensuring that the fused text abstract is more coherent and consistent. Compared with the static fusion method of the existing method, the present invention can flexibly adjust the information fusion strategy in different context scenarios, improving the quality of the generated abstract.

[0077] (3) Adapt to the requirements of complex scenarios: The abstract generation method of the present invention can handle complex application scenarios, such as long news, meeting records, and blind assistance technologies. Especially in the context where visual information and text information coexist, it can combine accurate image descriptions with the context to generate accurate and comprehensive text abstracts. When the prior art cannot effectively handle these scenarios, the present invention provides a more advanced and efficient solution. Brief Description of the Drawings

[0078] Figure 1 It is the overall flowchart of a visual-text collaborative abstract generation method based on multimodal learning disclosed in the embodiments of the present invention;

[0079] Figure 2 It is a schematic diagram of the visual semantic understanding stage in the embodiments of the present invention;

[0080] Figure 3 It is a schematic diagram of the text fusion stage in the embodiments of the present invention;

[0081] Figure 4 It is a schematic diagram of the abstract generation stage in the embodiments of the present invention. Detailed Description of the Embodiments

[0082] Next, in combination with the embodiments of the present invention and the accompanying drawings of the specification, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0083] The present invention provides a visual-text collaborative summary generation method based on multimodal learning, aiming to effectively fuse image and text information and improve the accuracy and semantic consistency of summary generation. By introducing a visual understanding module, a text-visual understanding combination module, and a summary generation module, the present invention can generate a summary that conforms to the context semantics on the premise of ensuring no semantic conflict between the image and the text.

[0084] As Figure 1 shown, the specific embodiments are as follows:

[0085] A visual-text collaborative summary generation method based on multimodal learning includes the following steps:

[0086] (1) Receive text data and visual data:

[0087] Receive multimodal input data containing text data and visual data. The text data can be an article, a report, or other text-form inputs, while the visual data can be the visual information of an image file, an image sequence, or a video. This step aims to provide the multimodal input data to the subsequent processing modules.

[0088] (2) Visual semantic understanding stage:

[0089] Video key frame recognition. Calculate the importance weights of pixel regions through an adaptive spatio-temporal attention mechanism, identify the core changes in the video content, remove redundant information, and obtain a set of representative and semantically rich key frame images. The calculation process is as follows:

[0090] Adaptive spatio-temporal weight fusion. On the basis of the standard spatio-temporal attention, add a dynamic weight factor to adaptively adjust the attention in the time and space dimensions. The time dimension focuses on the relationship between different frames in the video sequence, mainly reflecting the dynamic changes between frames, such as motion trajectories and scene switches, etc.; the space dimension focuses on the regional relationship within a single-frame image, emphasizing the local and global structures of objects, such as object shapes, textures, and boundaries, etc. The calculation formula is as follows:

[0091] Temporal attention:

[0092] ;

[0093] Spatial attention:

[0094] ;

[0095] Adaptive spatio-temporal attention:

[0096] ;

[0097] Among them, represents the attention matrix after weighting the importance of different frames in the time dimension, represents the attention matrix after weighting the importance of different regions in the spatial dimension; is the time attention factor, learning the global importance of each frame, is the spatial attention factor, , are both learnable parameters, enabling the model to dynamically adjust the time and space weights, and normalizing through the sigmoid function to ensure that and do not appear as negative values or values that are too large or too small, preventing unstable training.

[0098] Single-modal visual feature encoding. The present invention proposes an atrous convolution hybrid visual encoder, which enhances the ability to model global information through atrous convolution, and at the same time combines a hybrid convolution structure, using the combination of depthwise separable convolution and standard convolution to improve the compatibility of the model for multi-scale targets. Atrous convolution expands the receptive field by introducing an atrous factor, enabling the encoder to capture the dependencies between pixels at long distances in the image, and is suitable for processing targets with large spatial distributions. The hybrid convolution structure can effectively extract the local and global features of the image, thereby improving the expressiveness and efficiency of the model while reducing the computational complexity.

[0099] ;

[0100] ;

[0101] Among them, is the pixel value at position in the output feature map; represents the input feature at position in the pixel value, is the index of the convolution kernel, indicating the position of the convolution kernel in the sliding window; represents the weight of the convolution kernel at position , is the dynamic atrous factor, controlling the sampling interval of the convolution kernel, is the local feature of the current feature map at position , and are parameters to be learned; when it is a standard convolution, and when there is a gap between the sampling points of the convolution kernel;

[0102] Depthwise separable convolution includes:

[0103] Channel-wise convolution:

[0104] ;

[0105] Among them, represents the th channel of the input feature map, the pixel value at position , represents the index of the convolution kernel, is the weight of the convolution kernel at position ;

[0106] Pointwise convolution:

[0107] ;

[0108] Among them, is the output of the channel-wise convolution, is the convolution kernel, used for cross-channel weighted summation, and the number of output channels is controlled by adjusting ;

[0109] Hybrid convolution formula:

[0110] ;

[0111] Among them, represents the output feature map of the hybrid convolution operation at position , is the balance coefficient, , used to control the weight between the standard convolution and the depthwise separable convolution;

[0112] Dual-branch fusion formula:

[0113] ;

[0114] Among them, represents the fusion formula of the multi-scale dilated convolution and the hybrid convolution, is the dynamic fusion weight, controlling the proportion of global and local information.

[0115] Image-based text encoding. Convert the features extracted by the image encoder in the unimodal into text features, and inject visual information by introducing additional cross-attention between the self-attention layer and the feed-forward network layer to obtain a multimodal representation containing image content;

[0116] Image-based text decoding. The fused image-text information is further processed to generate a natural language description. By adopting a causal self-attention mechanism, the generated description text conforms to the context semantics and can generate an accurate description that matches the input image.

[0117] Image-text contrastive loss. By activating the unimodal visual feature encoder, the feature spaces of vision and text are aligned, encouraging positive sample image-text pairs to have similar representations while negative sample pairs have a large degree of discrimination, enhancing the correlation between images and text.

[0118] The contrastive loss formula is as follows:

[0119] ;

[0120] where, represents the image feature and the text feature The similarity between. The numerator calculates the similarity of positive samples. For a given image and the corresponding text , the similarity between the two is calculated. The denominator is the sum of the similarities between all candidate texts (including negative samples) and the image. is the logarithmic loss function, aiming to maximize the probability of positive samples and minimize the probability of negative samples. is the number of samples.

[0121] Image-text matching loss. By activating the image-based text encoder, an image-text multimodal representation that learns to capture fine-grained alignment between vision and language is learned, and the image-text matching loss is used to predict whether a given image-text pair of multimodal features matches.

[0122] The matching loss formula is as follows:

[0123] ;

[0124] ;

[0125] where, is the label of the image-text pair. indicates that the image and text match. indicates that the image and text do not match. is the probability that the model predicts whether the image and text match, and the logarithmic loss function is used to optimize the model. is the image feature and the text feature The similarity calculation between. is the sigmoid function, which is used to normalize similarity values to , ensuring that it can be used as a probability.

[0126] The language modeling loss. By activating the image-based text encoder, a text description of the given image is generated, and the likelihood of the generated text is maximized in an autoregressive manner to improve the quality of the text description and its semantic matching with the image.

[0127] The formula for the language modeling loss is as follows:

[0128] ;

[0129] where is the probability of generating the next word given the previous context and image information .

[0130] (3) Text fusion stage:

[0131] The steps of combining text and visual understanding include text-image consistency evaluation, text-visual semantic coordination detection, and context supplementation based on dynamic semantic fusion; the steps of accurately combining the input text with the text of visual understanding include:

[0132] Text-image consistency evaluation. Determine whether the input text and the text generated by visual understanding describe the same thing or concept. The direction consistency is used to measure the similarity between the input text and the text describing the image. The higher the direction consistency, the higher the similarity between the text description and the image. To avoid text redundancy, the similarity between the image description and the input text is calculated to determine whether to add new text. This method helps ensure that no description repeated with the original input content is introduced during the text fusion process. Its calculation formula is as follows:

[0133] ;

[0134] where and are the vector representations of the input text and the text describing the image respectively.

[0135] Text-visual semantic coordination detection. According to the similarity calculation result, determine whether there is a semantic conflict between the input text and the text describing the image. If the similarity between the two is low or there are obvious contradictions, it is marked as a conflict to avoid incorrect fusion;

[0136] Context supplementation based on dynamic semantic fusion. On the premise of ensuring that the input text does not conflict with the image description text, an information supplementation method based on dynamic semantic fusion and keyword alignment is proposed. This method uses an adaptive semantic alignment algorithm to dynamically adjust the information fusion strategy by calculating the complementary weights of the text and visual content, optimizing the generated text in terms of semantic integrity, context coherence, and description accuracy. Its steps include:

[0137] Keyword alignment: First, extract keywords from the input text and the image description text, calculate their similarity, and evaluate the semantic alignment degree between the two.

[0138] Complementary weight calculation: Use the adaptive alignment algorithm to calculate the complementary weights between the text and visual content. This weight measures the relative importance of the input text and the image description text, and the complementary weights are calculated through a gating mechanism, expressed as , where is the weight matrix, and are the encoded features of the input text and the image description text respectively, represents the complementary weights of the text and the image description:

[0139] .

[0140] Dynamic information fusion: Based on the complementary weights , fuse the input text and the image description text. The fusion method is to adjust the weights of the two in proportion to generate an enhanced text representation , which contains the combination of image and text information, and its calculation formula is:

[0141] .

[0142] (4) Abstract generation stage:

[0143] Text encoding layer. Convert each word in the text into a vector representation of a fixed dimension. These vectors can effectively capture the semantic information of the text and encode the semantic of each word and its context relationship, so as to obtain a high-quality representation of each word as the input of DeepFusion-Transformer.

[0144] Position Encoding Layer. Pure text encoding does not preserve the positional relationship of words in a sentence. It only presents the semantic representation of vocabulary and lacks the understanding of the word order within a sentence. Therefore, position encoding is introduced to enable the model to understand the relative and absolute positions between words. Through position encoding, the model can fully consider the order of words in the sequence during self-attention calculation, thereby improving the accuracy and fluency of text understanding. Specifically, the generation method of position encoding uses sine and cosine functions, making the encoding of each position have different frequencies in different dimensions, ensuring that the model can distinguish words at different positions.

[0145] DeepFusion-Transformer Layer. It adopts a stacked multi-layer structure, and each layer gradually learns more advanced features. The shallow Transformer layers focus on capturing local semantic information, while the deep Transformer layers focus on learning global semantic representations.

[0146] Autoregressive Decoding Layer. The decoder generates text step by step in an autoregressive manner. After generating a new word each time, the decoder updates its state, takes the generated word as the new input, and combines the current context information to continue the next generation until the complete summary text is generated.

[0147] The formula for the autoregressive generation process is as follows:

[0148] ;

[0149] where is the probability distribution of generating the word at the th step given the word and input features at the th step. is the current hidden state of the decoder, and are the weights and biases used to calculate the vocabulary distribution. Among them, the hidden state of the decoder is updated at each time step. Assuming the hidden state at the th step is , then at the th step, the new hidden state can be calculated recursively:

[0150] ;

[0151] where is the word generated at the th step, is the hidden state at the previous moment, and is the input multi-modal information.

[0152] ​A visual-text collaborative summary generation system based on multimodal learning, comprising:

[0153] Multimodal input module: Accepts input text data (such as news reports, meeting records, etc.) and visual data (such as images, image sequences, or videos related to the text content). The main function of this module is to provide the original input for subsequent visual understanding and text fusion;

[0154] Visual semantic understanding module: Inputs the visual data into the visual semantic understanding model. The model extracts visual features, identifies key information in the image or video, and generates a detailed description of the content of the visual data.

[0155] Text and visual understanding combination module: Effectively combines the input text with the image description text or video analysis text generated by the visual understanding model, thereby obtaining a fused representation of the text and image content.

[0156] Summary generation module: Inputs the text generated by the text and visual understanding combination module into the generation model. Through the model's understanding and processing of the text, a text summary that conforms to the context semantics is finally generated.

[0157] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visual-text collaborative summary generation method based on multimodal learning, characterized in that, Including the following steps: S1. Receive multimodal data: including text data and visual data as the basis for subsequent summary generation; S2. Visual semantic understanding: Input the visual data into the visual semantic understanding model. By extracting visual features, identify the key information in the image or video, and generate a text description of the content of the visual data; In step S2, the visual semantic understanding model includes video key frame recognition. By calculating the importance weights of pixel regions through adaptive spatio-temporal weight fusion, identify the core changes in the video content, remove redundant information, and obtain a set of representative and semantically rich key frame images: Adaptive spatio-temporal weight fusion: On the basis of standard spatio-temporal attention, add a dynamic weight factor to adaptively adjust the attention in the time and space dimensions; Temporal attention focuses on the relationship between different frames in the video sequence, reflecting the dynamic changes between frames. The formula for temporal attention is as follows: ; Spatial dimension attention focuses on the regional relationship within a single-frame image, emphasizing the local and global structure of the object. The formula is as follows: ; The formula for adaptive spatio-temporal weight fusion is as follows: ; is the temporal attention factor, which learns the global importance of each frame, is the spatial attention factor, which is normalized through the sigmoid function to ensure that and do not have negative values or values that are too large or too small; S3. Combine text and visual semantic understanding: Effectively fuse the semantic features of the text data in step S1 and the image description text or video parsing text generated in step S2, and provide semantic information for the summary generation model; In step S3, the combination of text and visual semantic understanding includes context supplementation based on dynamic semantic fusion. Context supplementation based on dynamic semantic fusion includes keyword alignment, complementary weight calculation, and dynamic information fusion; Complementary weight calculation: Using the adaptive alignment algorithm, calculate the complementary weight between the text and the visual content. This weight measures the relative importance of the input text and the image description text, and the complementary weight is calculated through a gating mechanism ; Dynamic Information Fusion: Based on Complementary Weights , fuse the input text and the image description text by adjusting the weights of the two in proportion to generate an enhanced text representation , which contains the combination of image and text information; S4. Summary generation: Input the fused text in step S3 into the summary generation model. Through the model's understanding and processing of the input text, finally generate a text summary that conforms to the context semantics.

2. The visual-text collaborative summary generation method based on multi-modal learning according to claim 1, wherein In step S2, the visual semantic understanding model includes unimodal visual feature encoding. A hybrid convolution strategy is adopted to combine standard convolution and depthwise separable convolution. The unimodal visual feature encoding adopts a two-branch structure. Global features are extracted through dilated convolution and local features are extracted through hybrid convolution through two parallel calculation paths, and adaptive feature fusion is performed through two-branch fusion.

3. The visual-text collaborative summary generation method based on multi-modal learning according to claim 2, wherein Dilated Convolution: By introducing a dynamic dilation factor strategy and using dilation factors of different scales for multi-level information fusion, the network can capture long-range pixel dependencies while maintaining the integrity of local information. The calculation formula is as follows: ; ; wherein, is the pixel value at position in the output feature map; represents the pixel value of the input feature at position ; is the index of the convolution kernel, indicating the position of the convolution kernel in the sliding window; represents the weight of the convolution kernel at position ; is the dynamic dilation factor, which controls the sampling interval of the convolution kernel; is the local feature of the current feature map at position ; and are parameters to be learned; when , it is a standard convolution, and when , there is an interval between the sampling points of the convolution kernel; Depthwise separable convolution includes: Pointwise convolution: ; Among them, represents the th channel of the input feature map, and the pixel value at position ; represents the index of the convolutional kernel, and is the weight of the convolutional kernel at position ; Depthwise convolution: ; Among them, is the output of per-channel convolution, is the convolution kernel, which is used for cross-channel weighted summation. By adjusting the number, the number of output channels is controlled; Hybrid convolution formula: ; Among them, represents the output feature map of the hybrid convolution operation at the position , is the balance coefficient, , which is used to control the weight between the standard convolution and the depthwise separable convolution; Two-branch fusion formula: ; Among them, represents the fusion formula of multi-scale dilated convolution and hybrid convolution, is the dynamic fusion weight, which controls the proportion of global and local information.

4. The visual-text collaborative summary generation method based on multi-modal learning according to claim 2, wherein In step S2, the visual semantic understanding model includes image-text contrast loss and image-text matching loss; Image-text contrast loss: By activating the unimodal visual feature encoder, align the feature spaces of vision and text, encourage positive sample image-text pairs to have similar representations, while negative sample pairs have a large degree of distinctiveness, enhance the correlation between images and text. The contrast loss formula is as follows: ; Among them, represents the similarity between the image features and the text features, is the logarithmic loss function, and N is the number of samples; Image-text matching loss: By activating the text encoder based on the image, learn to capture the image-text multimodal representation that finely aligns vision and language. Use the image-text matching loss to predict whether the given image-text pair of multimodal features matches. The matching loss formula is as follows: ; ; Among them, is the label of the image-text pair, indicating that the image and text match, indicating that the image and text do not match, is the probability that the model predicts whether the image and text match, using the logarithmic loss function to optimize the model, is the image feature and the text feature for calculating the similarity between them, is the sigmoid function, used to normalize the similarity value to to ensure that it can be used as a probability.

5. The visual-text collaborative summary generation method based on multi-modal learning according to claim 1, wherein In step S2, the visual semantic understanding model includes a language modeling loss. By activating the image-based text encoder, a text description of the given image is generated, and the likelihood of the generated text is maximized in an autoregressive manner. The calculation formula is as follows: ; Among them, is the probability of generating the next word given the previous context and image information .

6. The visual-text collaborative summary generation method based on multimodal learning according to claim 1, characterized in that In step S3, the combination of text and visual semantic understanding also includes graphic-text consistency evaluation and text-visual semantic coordination detection; Graphic-text consistency evaluation: The direction consistency is used to measure the similarity between the input text and the image description text. The higher the direction consistency, the higher the similarity between the text description and the image. The calculation formula is as follows: ; Among them, and are the vector representations of the input text and the image description text, respectively; Text-visual semantic coordination detection: According to the similarity calculation result, it is judged whether there is a semantic conflict between the input text and the image description text. If the similarity between the two is low or there are obvious contradictions, it is marked as a conflict to avoid incorrect fusion; Keyword alignment: Extract keywords from the input text and the image description text, and calculate their similarity to evaluate the semantic alignment degree between the two; Complementary weight calculation: ; is the weight matrix, and are the encoded features of the input text and the image description text respectively, represents the complementary weight of the text and the image description; The formula for dynamic information fusion is: 。 7. The visual-text collaborative summary generation method based on multi-modal learning according to claim 1, wherein In step S4, the model for abstract generation includes a text encoding layer, a position encoding layer, a DeepFusion-Transformer layer, and an autoregressive decoding layer; the text fused in step S3 is input into the trained abstract generation model to generate a text abstract; Text encoding layer: Each word in the text is converted into a vector representation of a fixed dimension, and the semantics of each word and its context relationship are encoded to obtain a high-quality representation of each word as the input of the DeepFusion-Transformer; Position encoding layer: Sine and cosine functions are used to generate position encodings so that the encodings of each position have different frequencies in different dimensions; DeepFusion-Transformer layer: A stacked multi-layer structure is adopted, and each layer gradually learns higher-level features. The shallow Transformer layer focuses on capturing local semantic information, and the deep Transformer layer focuses on learning global semantic representations; Autoregressive decoding layer: The decoder generates text step by step in an autoregressive manner. After generating a new word each time, the decoder updates its state, takes the generated word as the new input, and combines the current context information to continue the next generation until the complete abstract text is generated.

8. The method for generating a visual-text collaborative summary based on multimodal learning according to claim 7, wherein The formula for the autoregressive generation process is as follows: ; Among them, is the probability distribution of generating the vocabulary of the step given the words and input features of the previous step; is the hidden state of the current decoder, and are the weights and biases used to calculate the vocabulary distribution; where the hidden state of the decoder is updated at each time step. Assuming the hidden state at the step is ht-1, then at the t-th step, the new hidden state ht can be calculated recursively as follows: ; Among them, is the word generated in the previous time step, and is the input multi-modal information.

9. A visual-text collaborative summary generation system based on multimodal learning, characterized in that, The method described in any one of claims 1-8 is adopted, including a multimodal input module, a visual semantic understanding module, a text-visual understanding combination module, and an abstract generation module; Multimodal input module: Accepts the input text data and visual data. The main function of this module is to provide the original input for subsequent visual understanding and text fusion; Visual semantic understanding module: Inputs the visual data into the visual semantic understanding model. The model extracts visual features, identifies key information in the image or video, and generates a detailed description of the content of the visual data; Text-visual understanding combination module: Effectively combines the input text with the image description text or video parsing text generated by the visual understanding model to obtain a fused representation of the text and image content; Abstract Generation Module: Input the text generated by the Text and Visual Understanding Integration Module into the generation model. Through the model's understanding and processing of the text, finally generate a text abstract that conforms to the context semantics.

Citation Information

Patent Citations

  • Multi-modal text abstract system based on dependence gating fusion mechanism

    CN113609285A

  • Multi-modal model and method for fusing characters, images and audios

    CN118861988A