News-oriented multi-modal summary generation method

Through a multimodal summary generation method, the ViT-B/32 and CLIP models are used to extract image and text features, and the Transformer encoder and decoder are combined for multimodal interaction. This solves the problems of missing image description text and inconsistent evaluation indicators in multimodal summary generation, generates high-quality image and text summaries, and improves the user reading experience and information acquisition efficiency.

CN119621959BActive Publication Date: 2025-10-14SOUTHEAST UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411673674.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-10-14
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

Existing news summary generation technology has the problem of single modality extraction and ignores image information in multimedia data, resulting in incomplete summary content and a single user reading experience. In addition, existing multimodal summary generation methods fail to effectively utilize image description text and the objective function of the training stage is inconsistent with the evaluation indicators of the evaluation stage.

Method used

A multimodal summary generation method is adopted. The ViT-B/32 model is used to extract image and text features. The CLIP model is used for cross-modal matching. The Transformer encoder and decoder are combined for multimodal feature interaction. A summary scoring model is constructed to select the optimal summary. A cross-attention layer is added to the decoder to generate image-text matching information. The image selection module is integrated to improve the summary quality.

Benefits of technology

The accuracy and richness of multimodal summaries have been improved, the user reading experience has been greatly enhanced, news topics and sources can be customized according to personal interests, efficient graphic and text summaries can be generated, and the efficiency and convenience of information acquisition can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621959B_ABST
    Figure CN119621959B_ABST
Patent Text Reader

Abstract

The method for generating multi-modal summary of news customization comprises the following steps: first, inputting picture and text data into multi-modal encoding based on BART, and realizing multi-modal interaction in the encoding process; meanwhile, a picture-text matching module is added to select the corresponding sentence of each picture, and a cross-attention layer is added in the decoder; then, an abstract scoring model is constructed, the similarity between the candidate abstract and the news text is calculated as the score, and the optimal abstract is selected; the ROUGE score between the sentence corresponding to the picture and the optimal abstract is calculated as the text similarity score, the picture-text similarity score is obtained by combining the picture features after multi-modal interaction, the picture selection probability is obtained, the picture with the highest probability is selected as the model selected picture, and a multi-modal news abstract is generated. The system adopts B / S architecture, uses a lightweight Web framework Flask to build, and uses the Layui open source framework to realize the construction of the interactive interface. The application effectively improves the accuracy and richness of the multi-modal summary generation of news, and has strong robustness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of customized news generation, in particular to a multi-modal summary generation method for news customization. BACKGROUND

[0002] With the rapid expansion of the Internet, people are facing unprecedented information challenges. Due to the large amount of information, it is difficult to filter out the truly useful and relevant information. This situation is particularly prominent in Internet news, where a large number of news reports are intertwined, making it difficult for people to quickly understand the core and key points of the event. Therefore, a method that can help people quickly integrate news information can greatly improve the efficiency and experience of users in obtaining news information on the Internet.

[0003] However, there are still some problems in the current way of helping people quickly obtain news information. First of all, search engines, through the keywords input by the user, find news related to the keywords, and at the same time provide a brief summary of the search results. This summary mainly relies on the text in the news page matching the search keywords, ignoring the overall content and theme of the news page, so the summary may deviate from the actual content of the news, thereby affecting the user's reading experience and judgment. In addition, search engines usually select the picture in the summary according to the established rules, which is usually the first picture in the searched web page. This approach may result in a mismatch between the news content and the selected picture, confusing the user. Another way is the RSS (Really Simple Syndication) technology, which allows users to obtain website content in an RSS reader through subscribing to the RSS link provided by the website, realizing the real-time update of website content, but it also has obvious limitations. First of all, not all websites support RSS, which means that once a website does not support RSS, the user cannot obtain information from the website through the RSS reader. Secondly, the RSS subscription link provided by most news websites is usually the finest granularity of the channel, which makes it impossible for users to subscribe to news according to their own interest themes. Finally, the RSS technology simply presents the content of the news web page in the subscription link in the RSS reader without providing any form of summary, which undoubtedly increases the user's reading burden.

[0004] Therefore, there is an urgent need for a better way to quickly and accurately obtain the required information. Among them, dimensionality reduction processing of news reports is an effective way, that is, by simplifying the complex news content into a concise and clear summary, ensuring that the original meaning is preserved while eliminating redundant and insignificant details, which greatly reduces the user's browsing burden and frees people from tedious and redundant information.

[0005] Traditional text summarization typically relies on manual effort, requiring the writer to invest significant time and effort in reading the entire article, understanding its content, and extracting key information. This process is not only time-consuming and labor-intensive, but also limited by the writer's individual understanding and expertise. Different writers may extract different key information from the same article, resulting in discrepancies in the summary. In the era of big data, news information is experiencing explosive growth, making it unrealistic to manually summarize every piece of news. Furthermore, manual summarization carries potential issues such as inaccurate comprehension, omissions, and errors, which can compromise the completeness and accuracy of information. Therefore, relying solely on manual summarization fails to meet human needs. However, the advancement of natural language processing (NLP) technology is driving the automation and intelligentization of summary generation. In supervised methods, text summarization is treated as a binary classification task, using neural networks to learn the correspondence between sentences and their labels. Zhou et al. proposed a new scoring method that uses sentence gains as a scoring method by recording past sentence extractions, taking into account the interrelationships between sentences. Liu et al. first applied pre-trained language models to the field of extractive summarization and proposed BERTSUM. This model, based on the Bert model, adds a [CLS] token to the front of each sentence to obtain the characteristics of each sentence, and then generates a summary through the summary judgment layer. Zhong et al. considered the relationship between sentences in the article, extracting summaries by combining candidate sentences, and using the similarity between the combined sentences of the candidate sentences and the original document to judge the document summary. Bahdanau et al. proposed applying the attention mechanism to the original Seq2Seq to address Seq2Seq's poor ability to handle long sequences. The decoder uses the attention mechanism to dynamically extract the encoded information to generate a summary. Compared with manually written summaries, automatic text summarization technology is more efficient and accurate, greatly improving the efficiency and quality of information acquisition.

[0006] The rapid development of internet technology has also brought about a massive increase in multimedia data. Most current news web pages are multimedia documents, with the most common non-textual data being images interspersed between text. These images, with their vivid presentation and intuitive visual effects, help people better understand the news content. However, current automatic text summarization technologies suffer from a common problem of single-modality extraction when processing news data. Most of these technologies are limited to processing pure text content, neglecting the rich image information in news. This single-modality summary generation approach not only limits the comprehensiveness of the summary content but also fails to fully utilize the multimodal information in multimedia data. Furthermore, current automatic text summarization technologies only output text summaries, which gives people an overly limited sensory experience.

[0007] Therefore, researchers in academia have begun to focus on the task of multi-modal summary generation based on multimedia text data, aiming to break through the limitations of traditional summary techniques and achieve more comprehensive and accurate information extraction and presentation. This task involves making full use of text and image information in the data, achieving organic integration of text features and image features, semantic complementarity, and jointly generating a text-image summary to enable users to better understand the content of the original multimedia document and improve the user's reading experience. Existing research has shown that, compared to text-only summaries, multi-modal summaries can improve the quality of generated summaries by using information from the visual modality. And compared to text-only output, multi-modal output (text and images) increases user satisfaction by 12.4%. Current multi-modal summary generation methods mostly focus on how to interact between image data and text data across modalities, but in fact, there are significant differences between image and text data of different modalities, making their interaction extremely challenging. Lin et al. proposed a BART-based multi-modal sentence summary generation framework (BART-MMSS), which can effectively extract and utilize key information from images by introducing a prompt-guided image encoding module, and fuse it with text information to generate more accurate and rich summaries. Some researchers have introduced auxiliary tasks to further improve the quality of summaries in the multi-modal summary generation task. Zhu et al. proposed a new task, multi-modal summary task with multi-modal output (MSMO), which adds a picture selection task in addition to the text summary generation task. However, previous research has ignored the description text of the picture, which to some extent carries the information of the picture, connects the picture and the text, and provides a bridge for cross-modal semantics, and can be used as enhanced information for summary generation and picture selection in multi-modal output. However, not all pictures have description text, which limits the use of this information, so how to solve this problem is also a new challenge for multi-modal summary generation. At the same time, there is a problem in the current summary generation model that the objective function in the training phase does not match the evaluation criteria in the evaluation phase, that is, during training, the difference between the summary generated by the model and the reference summary at each position is measured, while the evaluation criteria usually measure the overall similarity between the summary generated by the model and the reference summary. This can cause a problem, where the model selects the sequence with the highest generation probability as the summary, but it may not be the best summary.

[0008] The present application is distinguished from the prior art as follows:

[0009] Comparison with the technology of patent CN106844341A "News summary extraction method and device based on artificial intelligence"

[0010] Patent CN106844341A provides an AI-based news summary extraction method, including: selecting core news from all news items related to the same news event, selecting core news from all news items related to the core news event, splitting all news items included in the news cluster into sentences, obtaining semantic similarity between each pair of sentences, selecting important sentences from the core news items based on their importance to form a summary, and concatenating them in the order of the original text, thereby avoiding logical confusion and semantic inconsistencies. This patent inputs multimodal image-text data into a multimodal encoder for multimodal interaction, simultaneously uses an image-text matching module to select the sentence corresponding to each image, and adds a cross-attention layer to the decoder to construct a summary scoring model. The optimal summary is selected by calculating the similarity between the candidate summaries and the news text as a score. The ROUGE score between the sentence corresponding to the image and the optimal summary is calculated as a text similarity score. The image-text similarity score is then combined with the image features after multimodal interaction to obtain the image selection probability. The image with the highest probability is selected as the model selection image, thereby generating a multimodal news summary. The two inventions differ in their content.

[0011] Technical comparison with patent CN114880461A "A Chinese news text summarization method combining contrastive learning and pre-training technology"

[0012] Patent CN114880461A provides a method for summarizing Chinese news text that combines contrastive learning and pre-training techniques. The method includes: constructing contrastive learning input data, using a BERT pre-trained model fine-tuned with a Chinese news corpus to obtain a contextual vector representation of the news text, classifying and scoring sentences in the text, extracting candidate sentences containing key information to obtain a candidate sentence set, inputting the candidate sentence set into an MT5 model fine-tuned with the Chinese news corpus, generating a summary result, and finally implementing end-to-end training of both extractive and generative models by combining the AECLoss loss function. This patent allows for simultaneous input of image and text data, and simultaneously selects the image that best matches the summary, achieving multimodal news summary generation and building a fully functional system. The system adopts a B / S architecture, is built using the lightweight web framework Flask, and uses the Layui open source framework to implement the interactive interface, enabling accurate and efficient completion of multimodal summary generation tasks. The two inventions differ fundamentally in their implementation methods and technical approaches.

[0013] Therefore, in combination with the above research background, the present application aims at the phenomenon of multimedia news information overload in the current Internet, and uses a multi-modal summary generation technology to alleviate this problem. In view of the limitation of the absence of picture description text in the current multi-modal summary generation technology, the present application proposes to use the sentence in the original text that is most relevant to the picture as a substitute for the description text of the picture, and integrates it into a sequence-to-sequence model as enhancement information for summary generation and picture selection in multi-modal output. In addition, multiple summaries are generated when generating summaries, and these summaries are input into a summary scoring model trained based on contrast learning. The summary scoring model, after training, can ensure that candidate summaries with higher evaluation indicators obtain higher scores. Finally, the summary with the highest score is selected as the output of the model. Finally, a news customization system is realized based on the model, which not only allows users to customize news topics and sources according to personal interests, but also automatically generates previews containing picture-text summaries for search results, helping users quickly grasp the news content and greatly improving the efficiency and convenience of news information acquisition. SUMMARY

[0014] To solve the above technical problems, the present application proposes a multi-modal summary generation method for news customization, which can effectively improve the accuracy and richness of news multi-modal summary generation and has strong robustness.

[0015] To achieve the above purpose, the technical solution adopted by the present application is:

[0016] The multi-modal summary generation method for news customization comprises the following steps:

[0017] (1) Text-picture multi-modal feature extraction:

[0018] The original picture is input into ViT-B / 32 for feature extraction to obtain a picture feature sequence. The original text is divided into sentences and then input into an embedding layer to be converted into a word vector to obtain a text feature sequence. Finally, the picture feature sequence and the text feature sequence are input into a Transformer encoder to realize multi-modal feature extraction through an attention mechanism;

[0019] (2) Text-picture matching model:

[0020] According to the picture feature sequence and the text feature sequence obtained in step (1), the CLIP model is used to convert them into feature vectors, and the cosine similarity of the picture features and the text features is calculated. The contrast learning is used to make the text and the picture have similar representations in the same feature space. The sentence with the maximum cosine similarity is considered as the sentence most relevant to the picture, thereby realizing cross-modal matching of the picture and the text;

[0021] (3) Multi-modal summary generation with integrated picture-text matching information:

[0022] According to the output of the picture-text matching module in step (2) and the output of the Transformer encoder obtained in step (1), they are taken as enhanced information, and a summary is generated through a multi-layer Transformer decoder by using a multi-layer cross-attention mechanism, wherein the first cross-attention layer is used to input the text feature vector corresponding to the picture into the decoder in advance and make the decoder pay more attention to the text vector, and the second cross-attention layer is used to obtain all text vector features after the multi-modal encoder interacts to obtain complete information and generate candidate summaries;

[0023] (4) The summary scoring model based on contrast learning:

[0024] According to the candidate summaries generated in step (3), they and the news text features are converted into vectors through an embedding layer and then input into the encoder of the summary scoring model again, and after semantic encoding, the [CLS] token in the last hidden state of the encoder output is taken as the vector representation of the candidate summary, the cosine similarity between the candidate summary vector and the news text vector is calculated to evaluate the quality of the candidate summary, and the candidate summary with the highest similarity to the news text is selected as the best summary.

[0025] (5) The picture selection module fusing picture-text matching information:

[0026] According to the outputs of steps (3) and (4), the picture-text similarity score and the text similarity score are obtained, which are input into a linear layer after splicing operation and normalized by applying a Sigmoid function, and finally the picture selection score is obtained.

[0027] (6) System function display.

[0028] As a further improvement of the application, in step (1), the features are extracted from the original picture and the original text by ViT-B / 32, which specifically includes the following steps:

[0029] (1-1) The pre-trained ViT-B / 32 model is used to extract the features of the original picture, specifically the ViT-B / 32 model pre-trained on the JFT-3B video dataset is used to extract the features of the original picture and the original text: the ViT-B / 32 model extracts the features of N input pictures to obtain a picture feature sequence For a text with a sequence length of M, a word segmentation operation is performed, and [CLS] tokens and [SEP] tokens are added before and after each sentence, and then the text is input into an embedding layer to convert it into a word vector to obtain a text feature sequence Before e v and e t are input into the Transformer encoder, e v and e t need to be processedv and e t Each of the vectors is added to a learnable modal type flag vector, and then the picture feature and the text feature are spliced and added to the position encoding to obtain e = [e v ; e t ];

[0030] (1-2) According to the obtained picture feature sequence e v and the text feature sequence e t , before e v and e t are input into the Transformer encoder, each of e v and e t needs to be added to a learnable modal type flag vector e type , e type corresponding to the picture and the text are different, and the dimension of e type is the same as the dimension of the picture feature and the text feature, so that the model can better learn the interaction between the modal, and considering that the position encoding has been added to the picture feature when passing through the Vision Transformer, therefore, only the text vector is added to the position encoding in the multi-modal encoder. The calculation formula is as follows:

[0031]

[0032] Where e type_img represents the type encoding corresponding to the picture, e type_text represents the type encoding corresponding to the text, e pos represents the position encoding, and the picture feature sequence and the text feature sequence are spliced as follows:

[0033] e = Concat (e v , e t ) (2)

[0034] Where Concat(·) represents the splicing operation;

[0035] (1-3) After obtaining e, it is input into the Transformer encoder for processing. The Transformer encoder is composed of N stacked Transformer encoding sublayers. Each encoding sublayer contains a multi-head attention part and a feedforward neural network. If it is the first layer encoding sublayer, the input of the first layer is H 0 = e = {e v ; e t}, otherwise, assuming that it is the jth layer encoding sublayer, its input is the hidden state sequence of the output of the last layer Where and respectively v and e t The hidden state output by the j-1 layer, for each encoding layer, the model first performs multi-head self-attention calculation on the input sequence H of the current layer, and the self-attention output is calculated by the interaction between the query matrix Q, the key matrix K and the value matrix V. In the calculation of self-attention, Q, K and V are all input H. On each head, Q and K are calculated by point multiplication and then divided by the square root of the dimension of the key matrix K After scaling, the attention weights are obtained by the Softmax operation, and these weights are multiplied by V to obtain the self-attention output of the head. Finally, the self-attention outputs of all heads are spliced and may be subjected to a linear transformation to obtain the final output result.

[0036] The multi-head attention mechanism is represented by the following formula:

[0037]

[0038] Where W O is the weight matrix to be trained, head i is the result obtained by the i-th head self-attention. After the multi-head attention calculation is completed, the result obtained by the multi-head attention needs to be subjected to residual connection and layer normalization. Let the output after the multi-head attention be H mul , the input of the encoding sublayer is H, and the result obtained after the residual connection and layer normalization is H'. Then H' can be calculated by the following formula:

[0039] H' = LayerNorm(H + H mul ) (4)

[0040] Where LayerNorm represents layer normalization. For input H, it can be represented by the following formula:

[0041]

[0042] Where μ and σ represent the mean and standard deviation of the hidden representation sequence subjected to normalization, ∈ is a very small constant to avoid the denominator being 0, α is a scaling factor, and β is a bias quantity to scale and translate the standardized vector. After obtaining the output result H', the H' is input into the feedforward neural network and subjected to residual connection and layer normalization again to obtain the hidden representation output by the current encoding sublayer. Let the hidden representation output by the j-th layer be H j , H j and the calculation formula of FFN is as follows: where W1, W2, b1 and b2 are all learnable parameters.

[0043]

[0044] wherein W1, W2, b1, b2 are all learnable parameters;

[0045] (1-4) Model training and classification using attention mechanism, the encoder constantly inputs the hidden state output by the previous encoding sublayer into the next encoding sublayer until the last encoding sublayer, and takes the hidden state h of the last layer of the encoder h v ; h t ], h is the visual language representation obtained after multi-modal interaction, wherein For the obtained h = [h v ; h t ], take h v in as the feature vector representing each picture to obtain Take h t in as the feature vector representing each sentence of the original text to obtain wherein L is the number of sentences after the original text is divided into sentences, and is the output of the multi-modal encoder, is the picture feature after interacting with the text content, let be the feature of the i-th picture after multi-modal interaction, perform linear transformation on it to reduce the dimension to 1, and then map it to the interval [0, 1] through the Sigmoid activation function, that is, the similarity score of the i-th picture and the text is obtained Finally, the following is obtained The calculation formula of the picture-text similarity score is as follows:

[0046]

[0047] During model training, binary classification cross-entropy is used as the loss function, denoted as L1, and the calculation formula is as follows:

[0048]

[0049] wherein y i represents the label of the real sample, represents the score output by the model, and N represents the number of pictures.

[0050] As a further improvement of the present application, step (2) specifically comprises:

[0051] With the CLIP model, the text and image are made to have similar representations in the same feature space through contrastive learning, thereby realizing cross-modal matching. The L input sentences are converted into L feature vectors by the text encoder of the CLIP, the N input pictures are converted into N feature vectors by the picture encoder of the CLIP, for each picture, the feature vector of the picture is calculated with the feature vectors of the L sentences to obtain the cosine similarity, the sentence with the maximum cosine similarity is regarded as the most relevant sentence of the picture, then the sequence number of the sentence in the input sentence sequence is taken, N pictures obtain N sentence sequence numbers, and finally the feature vector corresponding to the sequence number is taken from the output of the multi-modal encoder in step (1) . That is, the output of the image-text matching module.

[0052] As a further improvement of the application, step (3) specifically comprises:

[0053] According to the output of the multi-modal summary encoder in step (1) and the output of the image-text matching module , it is taken as the enhanced information and the final summary is generated, the input of the summary generation encoder is moved one position backward at each position during training and a special start token is inserted at the starting position, after obtaining the text vector features and the image-text matching features of the corresponding text of the picture after the multi-modal interaction, the decoder needs to generate a summary according to the above features, the structure of the decoder is similar to that of the encoder, and it is also stacked by N decoding layers, if the first layer of the decoder is currently processed, the input of this layer is the result of the initial input of the decoder after the embedding layer, otherwise, assuming that the jth decoding layer is currently processed, the input thereof will come from the hidden state sequence H j-1 output by the (j-1)th decoding layer.

[0054] Unlike the traditional Transformer decoding layer which only has one cross-attention layer, the decoding layer of the multi-modal summary generation model proposed by the application increases one cross-attention layer because it also needs to receive the sentence vector features of the sentences corresponding to the picture. The main function of the first cross-attention layer is to input the sentence vector of the sentence corresponding to the picture into the decoder in advance, so that the decoder pays more attention to these sentence vectors, and focuses more on these sentence vectors when calculating the cross-attention with the output of the encoder . The second cross-attention layer is to obtain all the sentence vector features after the multi-modal encoder interaction, so as to obtain complete information to generate a summary.

[0055] Assume that the current decoding layer is the jth layer, the input of the current layer is H, and the output hidden state is H j , then the calculation process of the decoding layer can be expressed as follows:

[0056]

[0057] where Q H =K H =V H =H, Represents the sentence vector of each sentence corresponding to each picture output by the image-text matching module, Represents all sentence vector features output by the multimodal encoder. Unlike the encoder, the decoder has a MaskedMultiHeadAttention layer. This is because when the decoder generates sentences, it must follow the order from left to right, that is, it can only consider the words in the current position and the previous position. In order to achieve this order, a mask mechanism is introduced when calculating self-attention. The specific operation is to calculate the dot product of matrix Q and matrix K and add a specific attention mask matrix when performing the calculation of formula 3. The lower left corner of the matrix M is all 0, and the upper right corner is all negative infinity. Since the upper right corner of the matrix M is negative infinity, after adding this mask, the corresponding position values ​​of the result matrix will become extremely small. When the Softmax operation is performed, the weight results corresponding to these minimum values ​​will be close to 0. Therefore, when calculating self-attention, the model actually only calculates attention for the current position and the previous position. The calculation formula of attention at this time is as follows:

[0058]

[0059] Finally, take the hidden state H output by the last layer of the decoder Decoder ={h1, h2, ..., h n}, where h i is the hidden state of the ith position, n is the length of the generated summary, and h i Input to the last linear layer, the output dimension of the linear layer is the size of the vocabulary, and then a Softmax calculation is performed on the output of the linear layer to obtain the probability distribution P(y i,p );

[0060] During training, cross entropy is used as the loss function, denoted as L2. The calculation formula of the loss function is as follows:

[0061]

[0062] Among them, yi P(y | x) represents the character distribution in the reference abstract, P(y i,p ) is the character probability distribution output by the model. The total loss formula of the multi-modal abstract generation model is as follows:

[0063] L=L1+L2 (12)

[0064] Wherein, L1 is the loss function of picture selection, and L2 is the loss function of abstract generation.

[0065] As a further improvement of the application, step (4) specifically comprises:

[0066] Using the n candidate abstracts generated in step (3), the candidate abstracts and the news text are converted into vectors through an embedding layer, and then input into the encoder of the abstract scoring model. After semantic encoding, the [CLS] token in the last layer of the encoder output is taken as the sentence vector of the candidate abstract and the document vector representation of the news text. The cosine similarity between the candidate abstract vector and the news text vector is used to evaluate the quality of the candidate abstract. The candidate abstract with the highest similarity to the news text is the best abstract. The calculation formula of this process is as follows:

[0067]

[0068] Wherein, s i is the sentence vector representation of the i-th candidate abstract, 1≤i≤n, D is the document vector representation of the news text, cosine(-) represents the calculation of the cosine similarity between the inputs, and argmax(sim i ) represents the subscript when the cosine similarity is maximum. The candidate abstract corresponding to the subscript is the output abstract.

[0069] The news text, candidate abstract and reference abstract are mapped to the same semantic space, and the similarity between the reference abstract and the news text should be greater than or equal to the similarity between each candidate abstract and the news text. Then the calculation formula of the loss function L ref is as follows:

[0070]

[0071] Wherein, n is the number of candidate abstracts, represents the sentence vector representation of the reference abstract obtained by the encoder of the abstract scoring model;

[0072] The candidate abstract generated by the multi-modal abstract generation model is used as a sample for contrast learning, the ROUGE score is calculated for each generated candidate abstract and the reference abstract, the average of the ROUGE-1, ROUGE-2 and ROUGE-L is taken as the comprehensive evaluation index of the candidate abstract, and the candidate abstracts are sorted from high to low according to the evaluation index, the candidate abstracts ranked in the front should have higher similarity with the news text, otherwise, the similarity is lower, and meanwhile, the similarity difference between two candidate abstracts with large sorting difference should be larger, so the loss function L cand The calculation formula is as follows:

[0073]

[0074] Wherein, λ is the margin, and is a hyperparameter, for the reference abstract ranked in a certain position, the similarity of the reference abstract after the position with the news text should be at least lower than the similarity of the reference abstract in the position with the news text by λ;

[0075] Finally, the loss function calculation formula of the abstract scoring model is as follows:

[0076] Loss=L ref +L cand (16)。

[0077] As a further improvement of the application, step (5) specifically comprises:

[0078] According to the image-text similarity score s1 obtained in step (1) and the sequence number of the picture corresponding sentence in step (2), the picture is one-to-one corresponding to the sentence text input into the model through the sequence number, and the ROUGE score calculation is performed between the picture corresponding sentence and the optimal abstract filtered out by the abstract scoring model in step (4);

[0079] Specifically, the average of ROUGE-1, ROUGE-2 and ROUGE-L is used as the text similarity score s2. This score reflects the text similarity between the picture corresponding sentence and the optimal abstract;

[0080] After the image-text similarity score s1 and the text similarity score s2 are spliced, they are input into a linear layer. The linear layer first performs dimension increasing operation, then performs dimension reduction, and applies Sigmoid function for normalization processing, and finally obtains the picture selection score s. The dimension increasing operation can extract more potential features in the input data, helping the model to capture more complex and subtle information in the data.

[0081] As a further improvement of the application, step (6) specifically comprises:

[0082] The system function display in step (6) includes a news crawling module, a news customization module, a hot search module, a browsing history module, a personal center module, a news website management module and a user management module;

[0083] The news crawling module is used for dynamically crawling the latest information on the news website in real time;

[0084] The news customization module is used for loading a preselected two-stage multi-modal summary generation model trained to provide a service of generating a preview of a text-image summary for a user;

[0085] The hot search module is used for recording and displaying the current search popularity in real time, and a user can call the news customization module to generate a multi-modal summary by clicking a hot search keyword;

[0086] The browsing history module is used for recording the news browsing record of a user to ensure that the user behavior data can be completely saved;

[0087] The personal center module is used for providing a personal information management service for a user, and the user can modify personalized information such as a nickname and a personal description and view the personal browsing record;

[0088] The news website management module is used for a system administrator to manage the news website provided by the system;

[0089] The user management module is used for a system administrator to manage user permissions and provide a user password reset function.

[0090] Advantages: Compared with the prior art, the present application has the following advantages:

[0091] (1) In view of the problem of missing picture description text in web news, the present application proposes to use the sentence most related to the picture as a substitute for the description text, and realizes multi-modal interaction in the encoding process, so as to ensure that the information of the sentences closely related to the picture is fully considered during summary generation to enhance the effect of summary generation;

[0092] (2) In view of the inconsistency between the objective function in the training stage of the summary generation model and the evaluation index in the evaluation stage, the present application constructs a summary scoring model. The model calculates the similarity between the candidate summary and the news text as the score to ensure that the summary with a high evaluation index has a higher score, so as to select the optimal summary;

[0093] (3) The present application also considers integrating the text-image matching information into the picture selection task, obtaining the sentence sequence number corresponding to each picture through the text-image matching module, and further selecting the candidate pictures according to the optimal summary to obtain the pictures most consistent with the news summary, so that the reader can obtain the best summary while viewing the matching pictures, greatly improving the reading experience of the reader in multiple modalities. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 It is the overall framework diagram of the present invention;

[0095] Figure 2 is a schematic diagram of the multimodal encoder structure;

[0096] Figure 3 This is a structural diagram of the image-text matching module;

[0097] Figure 4 This is a schematic diagram of the abstract generation decoder structure;

[0098] Figure 5 It is a schematic diagram of the structure of the summary scoring model;

[0099] Figure 6 It is the overall framework diagram of the system of the present invention;

[0100] Figure 7 It is the system news customization interface of the present invention;

[0101] Figure 8 It is the personal center interface of the system of the present invention;

[0102] Figure 9 It is the system user management interface of the present invention. DETAILED DESCRIPTION

[0103] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0104] The following is only one embodiment of the present invention. The present invention has many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications based on the present invention. These corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

[0105] like Figure 1 As shown, the news customized multimodal summary generation method of the present invention includes the following steps:

[0106] 1. Multimodal feature extraction of text and images

[0107] ViT-B / 32 is used to extract features from the input original image to obtain an image feature sequence. The original text is divided into sentences and then input into the embedding layer to be converted into word vectors to obtain a text feature sequence. Finally, the image feature sequence and the text feature sequence are input into the Transformer encoder together. The multimodal feature extraction is realized through the attention mechanism. The structural diagram of the multimodal encoder is shown in the figure. Figure 2 shown.

[0108] (1) The original picture is extracted using a pre-trained ViT-B / 32 model, specifically a ViT-B / 32 model pre-trained on a JFT-3B video dataset to extract features of the original picture and the original text: the ViT-B / 32 model extracts features of N input pictures to obtain a picture feature sequence For a text with a sequence length of M, a word segmentation operation is performed, and a [CLS] token and a [SEP] token are added before and after each sentence, and then input into an embedding layer to convert into a word vector to obtain a text feature sequence In the process of obtaining e v and e t , the picture feature sequence e v and the text feature sequence e t are obtained by inputting the original picture and the original text into the pre-trained ViT-B / 32 model, and then adding a learnable position encoding vector to each vector in e v and e t , and then concatenating the picture feature and the text feature to obtain e v = Concat(e t , e v )

[0109] (2) According to the obtained picture feature sequence e v and the text feature sequence e t , before inputting e v and e t into the Transformer encoder, a learnable modal type flag vector e v needs to be added to each vector in e t , e type of the picture and e type of the text are different, and the dimension of e type is the same as that of the picture feature and the text feature, so that the model can better learn the interaction between the modal, and considering that the position encoding has been added to the picture feature when passing through the Vision Transformer, therefore, only the text vector is added in the multi-modal encoder Position encoding calculation formula is as follows:

[0110]

[0111] Where e type_img represents the type encoding corresponding to the picture, e type_text represents the type encoding corresponding to the text, e pos represents the position encoding, and the picture feature sequence and the text feature sequence are concatenated as follows:

[0112] e = Concat(e v , e t ) (2)

[0113] Where Concat(-) represents the concatenation operation.

[0114] (3) After obtaining e, it is input into the Transformer encoder for processing. The Transformer encoder is composed of N stacked Transformer encoding sublayers. Each encoding sublayer contains a multi-head attention part and a feedforward neural network. If the current layer is the first encoding sublayer, then the input of the first layer is H 0 =e={e v ;e t Otherwise, let the current layer be the j-th coding sublayer, and its input is the hidden state sequence output by the previous layer. in and e v and e t The hidden state of the j-1 layer output. For each encoding layer, the model first performs a multi-head self-attention calculation on the input sequence H of the current layer. The self-attention output is obtained by the interaction between the query matrix Q, the key matrix K and the value matrix V. When calculating the self-attention, Q, K and V are all input H. On each head, Q and K are dot-producted and then divided by the square root of the dimension of the key matrix K. After scaling, the attention weights are obtained through the Softmax operation, and these weights are multiplied by V to obtain the self-attention output of the head. Finally, the self-attention outputs of all heads are spliced ​​together and may undergo a linear transformation to obtain the final output. The multi-head attention mechanism can be expressed as the following formula:

[0115]

[0116] Where W O is the weight matrix to be trained, head i is the result obtained by the i-th self-attention. After the multi-head attention calculation is completed, the result obtained by the multi-head attention needs to be residually connected and layer normalized. Let the output of the multi-head attention be H mul , the input of the encoding sublayer is H, and the result after residual connection and layer normalization is H′, then H′ can be calculated by the following formula:

[0117] H′=LayerNorm(H+H mul ) (4)

[0118] Among them, LayerNorm represents layer normalization. For input H, it can be expressed as the following formula:

[0119]

[0120] where μ and σ represent the mean and standard deviation of the hidden representation sequence for normalization, ∈ is a tiny constant to avoid the denominator being zero, α is a scaling factor, and β is a bias for scaling and shifting the normalized vector. After obtaining the output H', the H' is input into a Feed-Forward Neural Network (FFN) and after a residual connection and layer normalization again, the hidden representation of the current encoding sub-layer output is obtained, and the hidden representation of the output of the jth layer is denoted as H j , H j , and the calculation formula of the FFN is as follows:

[0121]

[0122] where W1, W2, b1, and b2 are all learnable parameters.

[0123] (4) The model is trained and classified using an attention mechanism, and the encoder continuously inputs the hidden state of the output of the previous encoding sub-layer into the next encoding sub-layer until the last encoding sub-layer, and the hidden state h = [h v ; h t ] of the last layer of the encoder is taken, and h is the visual language representation obtained after multi-modal interaction, where For the obtained h = [h v ; h t ], the v in h t is taken as the feature vector representing each picture to obtain The i in h j-1 is taken as the feature vector representing each sentence of the original text to obtain where L is the number of sentences after the original text is divided into sentences, and are the outputs of the multi-modal encoder. is the picture feature after interaction with the text content. Let be the feature of the ith picture after multi-modal interaction, perform linear transformation to reduce the dimension to 1, and then map it to the interval [0, 1] through the Sigmoid activation function, that is, the similarity score of the ith picture and the text is obtained Finally, the picture-text similarity score is obtained. The calculation formula of the picture-text similarity score is as follows:

[0124]

[0125] During model training, binary cross-entropy is used as the loss function, denoted as L1. The calculation formula is as follows:

[0126]

[0127] where y i represents the label of the real sample, represents the score output by the model, and N represents the number of pictures.

[0128] 2. Text-picture matching model

[0129] As shown in Figure 3 , the CLIP (Contrastive Language-Image Pre-training) model is used to make the text and image have similar representations in the same feature space through contrastive learning, thereby realizing cross-modal matching. The input L sentences are converted into L feature vectors by the text encoder of CLIP, and the input N pictures are converted into N feature vectors by the picture encoder of CLIP. For each picture, the feature vector of the picture is calculated with the feature vectors of the L sentences to obtain the cosine similarity, and the sentence with the maximum cosine similarity is regarded as the most relevant sentence to the picture. Then, the sequence number of this sentence in the input sentence sequence is taken, and N picture sequence numbers are obtained. Finally, the feature vector corresponding to the sequence number is taken from the output of the multi-modal encoder in step (1) , which is the output of the text-picture matching module.

[0130] 3. Multi-modal summary generation fusing text-picture matching information

[0131] According to the output of the multi-modal summary encoder in step (1) and the output of the text-picture matching module , they are used as enhanced information to generate the final summary. The structure diagram of the summary generation decoder is shown in Figure 4 . The input of the summary generation encoder is a reference summary in which each position is moved back by one bit and a special start token (decoder_start_token) is inserted at the starting position. After obtaining the text vector feature and the text-picture matching feature of the picture after multi-modal interaction, the decoder needs to generate a summary according to the above features. The structure of the decoder is similar to that of the encoder, and it is also stacked by N decoding layers. If the first layer of the decoder is being processed, the input of this layer is the result of the initial input of the decoder after the embedding layer. Conversely, assuming that the jth decoding layer is being processed, its input will come from the previous layer, i.e., the hidden state sequence H​j-1 .

[0132] Unlike the traditional Transformer decoding layer which only has one cross attention layer, the decoding layer of the multimodal summary generation model proposed in this paper also needs to receive the sentence vector features of the sentence corresponding to the image, so a cross attention layer is added to the decoding layer. The main function of the first cross attention layer is to pre-attract the sentence vector of the sentence corresponding to the image. Input to the decoder, so that the decoder pays more attention to these sentence vectors, and then compares them with the encoder output The calculation of cross attention focuses more on these sentence vectors. The second cross attention layer is to obtain all the sentence vector features after the interaction of the multimodal encoder Used to obtain complete information to generate a summary. Assume that the current decoding layer is the jth layer, the input of the current layer is H, and the output hidden state is H j , then the calculation process of the decoding layer can be expressed as follows:

[0133]

[0134] where Q H =K H =V H =H, Represents the sentence vector of each sentence corresponding to each picture output by the image-text matching module, Represents all sentence vector features output by the multimodal encoder. Unlike the encoder, the decoder has a MaskedMultiHeadAttention layer. This is because when the decoder generates sentences, it must follow the order from left to right, that is, it can only consider the words in the current position and the previous position. In order to achieve this order, a mask mechanism is introduced when calculating self-attention. The specific operation is to calculate the dot product of matrix Q and matrix K and add a specific attention mask matrix when performing the calculation of formula 3. The lower left corner of the matrix M is all zeros, and the upper right corner is all negative infinity. Since the upper right corner of the matrix M is negative infinity, after adding this mask, the corresponding position values ​​in the result matrix will become extremely small. When performing the Softmax operation, the weights corresponding to these minimum values ​​will be close to 0. Therefore, when calculating self-attention, the model actually only calculates attention for the current position and the previous position. The calculation formula for attention at this time is as follows:

[0135]

[0136] Finally, take the hidden state H output by the last layer of the decoder Decoder={h1, h2, ..., h n}, where h i is the hidden state of the ith position, and n is the length of the generated summary. i Input to the last linear layer, the output dimension of the linear layer is the size of the vocabulary, and then a Softmax calculation is performed on the output of the linear layer to obtain the probability distribution P(y i,p ).

[0137] During training, cross entropy is used as the loss function, denoted as L2. The calculation formula of the loss function is as follows:

[0138]

[0139] Among them, y i represents the character distribution in the reference abstract, P(y i,p ) is the character probability distribution of the model output. The total loss formula of the multimodal summary generation model is as follows:

[0140] L=L1+L2 (12)

[0141] It includes two parts: L1 is the image selection loss function obtained in step (1), and L2 is the loss function for summary generation.

[0142] 4. Summary Scoring Model Based on Contrastive Learning

[0143] Using the n candidate summaries generated in step (3), the candidate summaries and news text are converted into vectors through the embedding layer and then input into the encoder in the summary scoring model. After semantic encoding, the [CLS] token in the last hidden state of the encoder output is taken as the sentence vector of the candidate summary and the document vector representation of the news text. The structure of the summary scoring model is as follows: Figure 5 The cosine similarity between the candidate summary vector and the news text vector is used to evaluate the quality of the candidate summary. The best summary is the one with the highest similarity to the news text. The calculation formula for this process is as follows:

[0144]

[0145] where s i is the sentence vector representation of the i-th candidate summary (1≤i≤n), D is the document vector representation of the news text, cosine(·) represents the cosine similarity between the calculated inputs, argmax(sim i ) represents the subscript when the cosine similarity is maximized, and the candidate summary corresponding to the subscript is the output summary.

[0146] The news text, the candidate abstract and the reference abstract are mapped to the same semantic space, the similarity between the reference abstract and the news text should be greater than or equal to the similarity between each candidate abstract and the news text, and the loss function L ref The calculation formula of the loss function L

[0147]

[0148] Wherein, n is the number of candidate abstracts, The reference abstract is obtained by the encoder of the abstract scoring model.

[0149] The candidate abstracts generated by the multi-modal abstract generation model are used as samples for contrast learning, the ROUGE score is calculated between each generated candidate abstract and the reference abstract, the average value of ROUGE-1, ROUGE-2 and ROUGE-L is taken as the comprehensive evaluation index of the candidate abstract, and the candidate abstracts are sorted from high to low according to the evaluation index, the candidate abstracts ranked in the front should have higher similarity with the news text, and vice versa. The similarity difference between the two candidate abstracts with large sorting gap and the news text should also be large, so the loss function L cand The calculation formula of the loss function L

[0150]

[0151] Wherein, λ is the margin, and is a hyperparameter, for the reference abstract ranked in a certain position, the similarity between the reference abstract after the position and the news text should be at least lower than the similarity between the reference abstract in the position and the news text.

[0152] Finally, the loss function of the abstract scoring model is calculated as follows:

[0153] Loss=L ref +L cand (16)

[0154] 5、Picture selection module fusing picture-text matching information

[0155] According to the picture-text similarity score s1 obtained in step (1) and the sequence number of the picture corresponding sentence in step (2), the picture and the sentence text input by the model are one-to-one corresponding through the sequence number, and the ROUGE score is calculated between the picture corresponding sentence and the optimal abstract filtered by the abstract scoring model in step (4). Specifically, the average value of ROUGE-1, ROUGE-2 and ROUGE-L is used as the text similarity score s2. This score reflects the text similarity between the picture corresponding sentence and the optimal abstract.

[0156] The image-text similarity score s1 and the text similarity score s2 are spliced and input into a linear layer. The linear layer first performs dimension increasing, then dimension reduction, and normalizes by applying a Sigmoid function, and finally obtains the picture selection score s. The dimension increasing operation can extract more potential features in the input data, helping the model capture more complex and subtle information in the data.

[0157] 6. System function display

[0158] The news customization system designed by the application can be divided into five parts of a presentation layer, a service layer, an algorithm layer, a data layer and a persistence layer, and a system overall architecture diagram is as shown in Figure 6 The presentation layer directly interacts with the system user, provides a use entrance of the system for the user, and shows the request processing result to the user. The presentation layer in the system is divided into a user end and an administrator end. The service layer is a service provided by the system to the user, and completes specific business logic processing. The layer mainly provides seven function modules of a personal center, a browsing history, news crawling, news customization, today's hot search, news website management and user management. The algorithm layer provides algorithm support for obtaining a news customization result, and generates a related image-text summary for the crawled news. The data layer includes user data and browsing history data. The persistence layer provides data storage. In the system, there are two ways of MySQL and Redis, which are applied to different application scenarios.

[0159] The seven function modules of the news customization-oriented multi-modal summary generation system include a news crawling module, a news customization module, a today's hot search module, a browsing history module, a personal center module, a news website management module and a user management module.

[0160] The news crawling module is used for dynamically crawling the latest information on a news website in real time. In view of the fact that the system needs to dynamically crawl the latest information on the news website in real time, in order to improve the efficiency of the crawler in processing web pages, a distributed crawler is realized by using the Docker container technology, so that the packaging, deployment and management of the application program can be easily realized, and the environment configuration and version control between different hosts become consistent and convenient. Through Docker, multiple crawler instances can be quickly deployed, and the size of the crawler cluster can be easily expanded to adapt to different sizes of crawling tasks. The distributed crawler can integrate the computing power of multiple hosts to jointly complete a single crawling task, thereby significantly improving the overall performance and efficiency of the crawler.

[0161] The news customization module is used to load a pre-selected two-stage multi-modal summary generation model to provide a service of generating a preview of a picture-text summary for a user, the user selects a news website to browse, and inputs a keyword of personal interest, then the system retrieves the query interface, related parameters and URL of the website from the MySQL database according to the news website selected by the user, obtains the text and pictures of the news through the news crawling module, inputs the multi-modal summary generation model, and then generates a summary and selects the most relevant picture as the picture-text summary of the news. The effect display of the module is as shown in Figure 7 ;

[0162] The hot search module is used to record and display the current search popularity in real time. When the user enters the search page, the system will take out the top ten keywords in the hotSearch table according to the score from high to low, and the user can call the news customization module to generate a multi-modal summary by clicking the hot search keyword;

[0163] The browsing history module is used to record the user's news browsing record, and ensure that the user behavior data is saved completely. When the user clicks the news pushed by the news customization module, the system backend will respond immediately to automatically capture and record the URL, title, text summary generated by the model, picture link selected by the model, and the unique ID of the current user, and add these data to the browsing record table as a record;

[0164] The personal center module is used to provide personal information management services for users. The user can modify the nickname, personal description and other personalized information here, and view the personal browsing record to meet the personal display needs. At the same time, the module also provides the function of modifying the password to ensure the security of the user account. In addition, the personal center also integrates the function of browsing history display. When the user enters the personal center, the system will query the user's browsing history from the browsing record table, and present the results to the user in real time, so that the user can access his own browsing record at any time to review the previous news browsing situation. The effect display of the module is as shown in Figure 8 ;

[0165] The news website management module is used for system administrator to manage the news website provided by the system. When the crawling rule of the news website is invalid, the system administrator can modify the crawling rule to ensure the normal operation of the news crawling function and maintain the stability of the system. In addition, in order to further improve the use experience of the administrator, the news website management module specially adds a news website test function. The administrator only needs to click the test button behind a specific news website, and the system will try to grab the URL, title, body and picture of the news according to the related rules of the website, and the results will be fed back to the front end in real time. This function greatly facilitates the administrator to quickly locate and solve problems, and improves the convenience and efficiency of news website management;

[0166] The user management module is used for system administrator to manage user permissions and provide user password reset function, and the access permission is limited to the system administrator role. The system administrator can promote the authority of the ordinary user according to the actual demand, and perform the administrator's duty. The effect of the module is shown as Figure 9

[0167] The above is only a preferred embodiment of the present application, not any other form of limitation on the present application, and any modification or equivalent change made according to the technical essence of the present application still belongs to the scope of the present application.​

Claims

1. A news-customized multimodal summary generation method, characterized by: The method comprises the following steps: (1) Multimodal feature extraction of text and images: ViT-B / 32 is used to extract features from the input original image to obtain an image feature sequence. The original text is segmented into sentences and then input into the embedding layer to convert it into word vectors to obtain a text feature sequence. Finally, the image feature sequence and text feature sequence are input into the Transformer encoder together to realize multimodal feature extraction through the attention mechanism. (2) Text-image matching model: According to the image feature sequence and text feature sequence obtained in step (1), they are converted into feature vectors through the CLIP model, and the cosine similarity between the image features and the text features is calculated. By using contrastive learning, the text and the image have similar representations in the same feature space. The sentence with the largest cosine similarity will be regarded as the sentence most relevant to the image, thereby achieving cross-modal matching between images and texts. (3) Multimodal summary generation integrating image-text matching information: According to the output of the image-text matching module in step (2) and the output of the Transformer encoder obtained in step (1), they are used as enhanced information and a summary is generated through a multi-layer Transformer decoder using a multi-layer cross-attention mechanism. The first cross-attention layer is used to input the text feature vector corresponding to the image into the decoder in advance and make the decoder pay more attention to the text vector, while the second cross-attention layer is used to obtain all the text vector features after the multimodal encoder interaction to obtain complete information and thus generate a candidate summary; (4) Summary scoring model based on contrastive learning: Based on the candidate summaries generated in step (3), they are converted into vectors together with the news text features through the embedding layer and then input into the encoder of the summary scoring model again. After semantic encoding, the [CLS] token in the last hidden state of the encoder output is taken as the vector representation of the candidate summary. The cosine similarity between the candidate summary vector and the news text vector is calculated to evaluate the quality of the candidate summary. The candidate summary with the highest similarity to the news text is selected as the best summary; (5) Image selection module integrating image-text matching information: According to the output of step (3) and step (4), the image-text similarity score and the text similarity score are obtained. After the splicing operation, they are input into a linear layer and normalized by the Sigmoid function to finally obtain the image selection score; (6) System function display.

2. The news-customized multimodal summary generation method according to claim 1 is characterized by: In step (1), feature extraction is performed using ViT-B / 32 based on the original image and original text, specifically including the following steps: (1-1) Use the pre-trained ViT-B / 32 model to extract the original image. Specifically, the ViT-B / 32 model pre-trained on the JFT-3B video dataset is used to extract the features of the original image and original text: ViT-B / 32 extracts features from the input N images and obtains the image feature sequence For the text with a sequence length of M, perform word segmentation and add [CLS]token and [SEP]token before and after each sentence, then input it into the embedding layer to convert it into a word vector to obtain the text feature sequence In the v and e t Before inputting into the Transformer encoder, e v and e t Each vector in is added with a learnable modality type flag vector, and then the image features and text features are concatenated and positionally encoded to obtain e = [e v ;e t ]; (1-2) According to the obtained image feature sequence e v and text feature sequence e t , in the e v and e t Input to Transformer Before the encoder, you need to v and e t Each vector in is added with a learnable modality type flag vector e type , e corresponding to the picture and text type Different, e type The dimension of is the same as that of the image features and text features, so that the model can better learn the interaction between modalities. Considering that the image features have been positionally encoded when passing through VisionTransformer, only the text vector is positionally encoded in the multimodal encoder. The calculation formula is as follows: where e type_img Represents the type code corresponding to the picture, e type_text Represents the type code corresponding to the text, e pos Represents position encoding and concatenates the image feature sequence and text feature sequence: e=Concat(e v ,And t )(2) Where Concat(·) represents the concatenation operation; (1-3) After obtaining e, it is input into the Transformer encoder for processing. The Transformer encoder is composed of N stacked Transformer encoding sublayers. Each encoding sublayer contains a multi-head attention part and a feedforward neural network. If the current layer is the first encoding sublayer, then the input of the first layer is H 0 =e={e v ;e t Otherwise, if the current layer is the j-th coding sublayer, its input is the hidden state sequence output by the previous layer. in and e v and e t After the hidden state of the j-1th layer output, for each encoding layer, the model first performs multi-head self-attention calculation on the input sequence H of the current layer. The self-attention output is obtained by the interaction between the query matrix Q, the key matrix K and the value matrix V. When calculating the self-attention, Q, K and V are all input H; on each head, Q and K are dot-producted and then divided by the square root of the dimension of the key matrix K. After scaling, the attention weights are obtained through the Softmax operation, and these weights are multiplied by V to obtain the self-attention output of the head. Finally, the self-attention outputs of all heads are spliced ​​together and may undergo a linear transformation to obtain the final output result; The multi-head attention mechanism is expressed as the following formula: Where W O is the weight matrix to be trained, head i is the result obtained by the i-th self-attention; After the multi-head attention calculation is completed, the results obtained by the multi-head attention need to be residually connected and layer normalized. Let the output of the multi-head attention be H mul , the input of the encoding sublayer is H, and the result after residual connection and layer normalization is H', then H' can be calculated by the following formula: H'=LayerNorm(H+H mul ) (4) Among them, LayerNorm represents layer normalization. For input H, it can be expressed as the following formula: Where μ and σ represent the mean and standard deviation of the normalized implicit representation sequence, ∈ is a very small constant used to avoid the case where the denominator is 0, α is the scaling factor, and β is the bias used to scale and translate the normalized vector. After obtaining the output result H', H' is input into the feedforward neural network and a residual connection and layer normalization are performed again to obtain the implicit representation of the current encoding sublayer output. Let the implicit representation of the j-th layer output be H j , H j The calculation formula of FFN is as follows: where W1, W2, b1, and b2 are all learnable parameters; Among them, W1, W2, b1, and b2 are all learnable parameters; (1-4) Use the attention mechanism for model training and classification. The encoder continuously inputs the hidden state output by the previous encoding sublayer into the next encoding sublayer until the last encoding sublayer. The hidden state h of the last layer of the encoder is taken as h = [h v ;h t ], h is the visual language representation obtained after multimodal interaction, where For the obtained h=[h v ;h t ], take h v in As the feature vector representing each image, we get Take h t in As the feature vector representing each sentence of the original text, we get Where L is the number of sentences in the original text after sentence segmentation, and is the output of the multimodal encoder, is the image feature after interacting with the text content, let The feature of the i-th picture after multimodal interaction is linearly transformed to reduce the dimension to 1, and then mapped to the [0,1] interval through the Sigmoid activation function, that is, the similarity score between the i-th picture and the text is obtained. Finally got The formula for calculating the image-text similarity score is as follows: When training the model, the binary cross entropy is used as the loss function, denoted as L1, and the calculation formula is as follows: where y i represents the label of the real sample, Represents the score of the model output, and N represents the number of images.

3. The news-customized multimodal summary generation method according to claim 1 is characterized by: Step (2) specifically includes: Using the CLIP model, text and images are made to have similar representations in the same feature space through contrastive learning, thereby achieving cross-modal matching. The input L sentences are converted into L feature vectors through the CLIP text encoder, and the input N pictures are converted into N feature vectors through the CLIP picture encoder. For each picture, the cosine similarity is calculated between the feature vector of this picture and the feature vectors of the L sentences respectively. The sentence with the largest cosine similarity will be regarded as the sentence most relevant to this picture. Then, the sequence number of this sentence in the input sentence sequence is taken. N pictures get N sentence sequence numbers. Finally, the output of the multimodal encoder in step (1) is obtained. Take out the feature vector of the corresponding sequence number This is the output of the image-text matching module.

4. The news-customized multimodal summary generation method according to claim 1 is characterized by: Step (3) specifically includes: According to the output of the multimodal summary encoder in step (1) Output of the image-text matching module It is used as enhanced information to generate the final summary. The input of the summary generation encoder is shifted back one position at each position during training and a special start marker is inserted at the starting position. After obtaining the text vector features after multimodal interaction, And the image-text matching features of the corresponding text After that, the decoder needs to generate a summary based on the above features. The structure of the decoder is similar to that of the encoder, and is also composed of N decoding layers stacked together. If the first layer of the decoder is currently being processed, then the input of this layer is the result of the initial input of the decoder passing through the embedding layer. Conversely, assuming that the jth layer of the decoder is currently being processed, then its input will come from the previous layer, that is, the implicit state sequence H output by the j-1th layer of the decoder. j-1 ; Assume that the current decoding layer is the jth layer, the input of the current layer is H, and the output hidden state is H j , then the calculation process of the decoding layer can be expressed as follows: where Q H =K H =V H =H, Represents the sentence vector of each sentence corresponding to each picture output by the image-text matching module, Represents all sentence vector features output by the multimodal encoder. Unlike the encoder, the decoder has a MaskedMultiHeadAttention layer. This is because when the decoder generates sentences, it must follow the order from left to right, that is, it can only consider the words in the current position and the previous position. In order to achieve this order, a mask mechanism is introduced when calculating self-attention. The specific operation is to calculate the dot product of matrix Q and matrix K and add a specific attention mask matrix when performing the calculation of formula 3. The lower left corner of the matrix M is all 0, and the upper right corner is all negative infinity. Since the upper right corner of the matrix M is negative infinity, after adding this mask, the corresponding position values ​​of the result matrix will become extremely small. When the Softmax operation is performed, the weight results corresponding to these minimum values ​​will be close to 0. Therefore, when calculating self-attention, the model actually only calculates attention for the current position and the previous position. The calculation formula of attention at this time is as follows: Finally, take the hidden state H output by the last layer of the decoder Decoder ={h1,h2,…,h n }, where h i is the hidden state of the ith position, n is the length of the generated summary, and h i Input to the last linear layer, the output dimension of the linear layer is the size of the vocabulary, and then a Softmax calculation is performed on the output of the linear layer to obtain the probability distribution P(y i,p ); During training, cross entropy is used as the loss function, denoted as L2. The calculation formula of the loss function is as follows: Among them, y i represents the character distribution in the reference abstract, P(y i,p ) is the character probability distribution of the model output; the total loss formula of the multimodal summary generation model is as follows: L=L1+L2 (12) It includes two parts: L1 is the loss function for image selection, and L2 is the loss function for summary generation.

5. The news-customized multimodal summary generation method according to claim 1 is characterized by: Step (4) specifically includes: Using the n candidate summaries generated in step (3), the candidate summaries and news text are converted into vectors through the embedding layer and then input into the encoder in the summary scoring model. After semantic encoding, the [CLS] token in the last hidden state of the encoder output is taken as the sentence vector of the candidate summary and the document vector of the news text. The cosine similarity between the candidate summary vector and the news text vector is used to evaluate the quality of the candidate summary. The candidate summary with the highest similarity to the news text is the best summary. The calculation formula of this process is as follows: where s i is the sentence vector representation of the i-th candidate summary 1≤i≤n, D is the document vector representation of the news text, cosine(·) represents the cosine similarity between the calculated inputs, argmax(sim i ) represents the subscript when the cosine similarity is maximized, and the candidate summary corresponding to the subscript is the output summary; Map the news text, candidate abstracts, and reference abstracts to the same semantic space, so that the similarity between the reference abstract and the news text should be greater than or equal to the similarity between each candidate abstract and the news text. Then the loss function L ref The calculation formula is as follows: Where n is the number of candidate summaries, Represents the sentence vector representation obtained by the encoder of the summary scoring model after the reference summary is passed; The candidate summaries generated by the multimodal summary generation model are used as samples for comparative learning. The ROUGE score is calculated for each generated candidate summary and the reference summary. The average of ROUGE-1, ROUGE-2 and ROUGE-L is taken as the comprehensive evaluation index of the candidate summary. The candidate summaries are ranked from high to low according to the evaluation index. The candidate summaries ranked higher should have a higher similarity with the news text, and vice versa. At the same time, the difference in similarity between the two candidate summaries with a large ranking difference and the news text should also be large. Therefore, the loss function L cand The calculation formula is as follows: Where λ is the margin, which is a hyperparameter. For a reference summary ranked at a specific position, the similarity between the reference summary at the next position and the news text should be at least λ lower than the similarity between the reference summary at that position and the news text. The loss function calculation formula of the final summary scoring model is as follows: Loss=L ref +L cand (16)。 6. The news-customized multimodal summary generation method according to claim 1 is characterized by: Step (5) specifically includes: According to the image-text similarity score s1 obtained in step (1) and the serial number of the sentence corresponding to the image in step (2), the image is matched with the sentence text input by the model through this serial number, and the ROUGE score is calculated for the sentence corresponding to each image and the optimal summary selected by the summary scoring model in step (4); Specifically, the average of ROUGE-1, ROUGE-2, and ROUGE-L is used as the text similarity score s2; this score reflects the text similarity between the sentence corresponding to the image and the optimal summary; After the image-text similarity score s1 and the text similarity score s2 are concatenated, they are input into a linear layer. This linear layer first performs a dimensionality increase operation, then a dimensionality reduction operation, and applies a Sigmoid function for normalization. Finally, the image selection score s is obtained. The dimensionality increase operation can extract more potential features from the input data and help the model capture more complex and subtle information in the data.

7. The news-customized multimodal summary generation method according to claim 1 is characterized by: Step (6) specifically includes: In step (6), the system function display includes news crawling module, news customization module, today's hot search module, browsing history module, personal center module, news website management module and user management module; The news crawling module is used to dynamically crawl the latest information on the news website in real time; The news customization module is used to load a pre-selected and trained two-stage multimodal summary generation model to provide users with a service for generating a graphic summary preview; The Today's Hot Search module is used to record and display the current search popularity in real time. Users can call the news customization module to generate a multimodal summary by clicking on the hot search keyword; The browsing history module is used to record the user's news browsing history to ensure that the user behavior data is fully preserved; The personal center module is used to provide users with personal information management services, where users can modify personalized information such as nicknames, personal descriptions, and view their personal browsing history; The news website management module is used by the system administrator to manage the news websites provided by the system; The user management module is used by the system administrator to manage user permissions and provide a user password reset function.

Citation Information

Patent Citations

  • News summary extraction method and device based on artificial intelligence

    CN106844341A

  • Query-based text summarization using cosine similarity and nmf

    KR100751295B1

  • Generating summary content using supervised sentential extractive summarization

    US20210042391A1