A model training method, an information recommendation method, and a device
By training a generative model and combining text and multimedia information to generate summaries, the problem of users being unable to vividly and intuitively understand the key points of recommended content in existing technologies is solved, resulting in a better user experience.
Patent Information
- Application Number
- CN202210621979.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-06-01
AI Technical Summary
In existing technologies, summaries generated solely from text information cannot vividly convey the key information of recommended content, resulting in a reduced user experience.
By acquiring target samples containing text and multimedia information, a generative model is used to determine the relevance between multimedia information and text information. Target multimedia information is selected, and a text summary is generated. The model is trained with the optimization objective of maximizing the similarity between multimedia information and labeled information, and the similarity between text summary and labeled summary.
It improves users' understanding of the key information in recommended content, thus enhancing the user experience.
Smart Images

Figure CN115033787B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method for model training, a method for information recommendation, and an apparatus. Background Technology
[0002] Currently, when various media platforms recommend content of interest to users, they often need to extract a summary from the text of the recommended content to summarize the main idea, so that users can understand the important information in the recommended content after reading the summary. However, a summary consisting only of text cannot vividly and clearly convey the key information in the recommended content, thus reducing the user experience.
[0003] Therefore, how to determine a summary of a vivid image is an urgent problem to be solved. Summary of the Invention
[0004] This specification provides a method, apparatus, storage medium, and electronic device for model training, in order to partially solve the aforementioned problems existing in the prior art.
[0005] The following technical solution is adopted in this specification:
[0006] This manual provides a method for model training, including:
[0007] Obtain target sample information, which includes text information and at least one multimedia information;
[0008] The target sample information is input into the generative model to be trained, so as to determine the correlation between the multimedia information and the text information contained in the target sample information for each multimedia information contained in the target sample information, and use it as the correlation corresponding to the multimedia information.
[0009] Based on the relevance of each multimedia information, target multimedia information is selected from each multimedia information, and based on the text information contained in the target sample information, a text summary corresponding to the target sample information is generated.
[0010] The generation model is trained with the optimization objective of maximizing the similarity between the target multimedia information and the labeled multimedia information, as well as the similarity between the text summary and the labeled text summary.
[0011] Optionally, for each piece of multimedia information contained in the target sample information, the correlation between the multimedia information and the text information contained in the target sample information is determined as the correlation corresponding to the multimedia information, specifically including:
[0012] For each multimedia information contained in the target sample information, the original multimedia features corresponding to the multimedia information are processed to obtain the processed features corresponding to the multimedia information. In addition, the original text features corresponding to the text information contained in the target sample information are processed to determine the processed features corresponding to the text information. The processed features corresponding to the multimedia information include multimedia query features and multimedia key features. The processed features corresponding to the text information include text query features and text key features.
[0013] Based on the correlation between the multimedia query features corresponding to the multimedia information and the text key features corresponding to the text information contained in the target sample information, and the correlation between the text query features corresponding to the text information contained in the target sample information and the multimedia key features corresponding to the multimedia information, the correlation between the multimedia information and the text information contained in the target sample information is determined, and this correlation is taken as the correlation between the multimedia information and the target sample information.
[0014] Optionally, the processed features corresponding to the multimedia information further include: multimedia value features;
[0015] Based on the relevance of each multimedia information, target multimedia information is selected from the multimedia information, specifically including:
[0016] For each multimedia information, the attention features corresponding to that multimedia information are determined based on its relevance and multimedia value features.
[0017] Based on the attention characteristics corresponding to each multimedia message, determine the comprehensive multimedia features;
[0018] Target multimedia information is selected from each multimedia information based on the similarity between the original multimedia features corresponding to each multimedia information and the comprehensive multimedia features.
[0019] Optionally, the processed features corresponding to the text information further include: text value features, wherein the text information contains several sub-texts;
[0020] Based on the text information contained in the target sample information, a text summary corresponding to the target sample information is generated, specifically including:
[0021] Based on the relevance of each multimedia information and the text value features of the text information contained in the target sample information, the attention features corresponding to the text information contained in the target sample information are determined.
[0022] For each subtext in the text information contained in the target sample information, the comprehensive text features corresponding to the subtext are determined based on the correlation between the subtext and other subtexts in the text information contained in the target sample information.
[0023] Based on the comprehensive text features corresponding to each subtext and the attention features corresponding to the text information contained in the target sample information, a text summary corresponding to the target sample information is generated.
[0024] Optionally, for each subtext in the text information contained in the target sample information, based on the correlation between the subtext and other subtexts in the text information contained in the target sample information, the comprehensive text features corresponding to the subtext are determined, specifically including:
[0025] For each subtext in the text information contained in the target sample information, the original subtext features corresponding to the subtext are processed to obtain the processed features corresponding to the subtext. The processed features corresponding to the subtext include: subtext query features, subtext key features, and subtext value features.
[0026] For each other subtext in the text information contained in the target sample information, the correlation between the subtext query feature corresponding to the subtext and the subtext key feature corresponding to the other subtext in the text information contained in the target sample information is taken as the correlation of the other subtext.
[0027] Based on the relevance of the other subtext and the subtext value features of the other subtext, determine the attention features corresponding to the other subtext;
[0028] Based on the attention features corresponding to each other subtext, determine the comprehensive text features corresponding to that subtext.
[0029] Optionally, obtain target sample information, specifically including:
[0030] Obtain information for each initial sample. Each initial sample contains an annotated text summary, text information contained in the initial sample, and several multimedia information.
[0031] The initial sample information is input into a pre-trained annotation model. Based on the explanatory text corresponding to each multimedia information contained in the text information of the initial sample information, multimedia information for annotation is selected from the multimedia information contained in the initial sample information to annotate the initial sample information, thereby obtaining the target sample information corresponding to the initial sample information.
[0032] Optionally, training the labeled model includes:
[0033] Obtain standard sample information, which includes annotated multimedia information, annotated text summaries, text information contained in the standard sample information, and several multimedia information items;
[0034] The standard sample information is input into the annotation model to be trained, so as to select the multimedia information to be verified corresponding to the standard sample information from each multimedia information based on the similarity between the explanatory text corresponding to each multimedia information contained in the text information in the standard sample information and the annotated text summary contained in the standard sample information.
[0035] The annotation model is trained with the optimization objective of minimizing the deviation between the multimedia information to be verified corresponding to the standard sample information and the annotated multimedia information contained in the standard sample information.
[0036] Optionally, based on the similarity between the descriptive text corresponding to each multimedia information contained in the text information of the standard sample information and the annotated text summary contained in the standard sample information, the multimedia information to be verified corresponding to the standard sample information is selected from each multimedia information, specifically including:
[0037] Based on at least one of the following: the similarity between the explanatory text corresponding to each multimedia information contained in the text information of the standard sample information and the annotated text summary contained in the standard sample information; and the position information of each multimedia information contained in the standard sample information in the text information, the multimedia information to be verified corresponding to the standard sample information is selected from each multimedia information.
[0038] This specification provides a method for information recommendation, including:
[0039] Retrieve recommendation information that includes several multimedia and text information;
[0040] The recommendation information, which contains several multimedia information and text information, is input into the generation model to generate a text summary and determine the target multimedia information. The generation model is trained using the model training method described above.
[0041] The recommended information is displayed to the user using the target multimedia information and the text summary.
[0042] This specification provides a model training apparatus, comprising:
[0043] An acquisition module is used to acquire target sample information, which includes text information and at least one multimedia information.
[0044] The determination module is used to input the target sample information into the generative model to be trained, so as to determine the correlation between the multimedia information and the text information contained in the target sample information for each multimedia information contained in the target sample information, and use it as the correlation corresponding to the multimedia information.
[0045] The generation module is used to select target multimedia information from each multimedia information according to the relevance of each multimedia information, and to generate a text summary corresponding to the target sample information according to the text information contained in the target sample information.
[0046] The training module is used to train the generative model with the optimization objective of maximizing the similarity between the target multimedia information and the labeled multimedia information, and the similarity between the text summary and the labeled text summary.
[0047] This specification provides an information recommendation device, comprising:
[0048] The acquisition module is used to acquire recommendation information that contains several multimedia information and text information;
[0049] The generation module is used to input the recommendation information containing several multimedia information and text information into the generation model to generate a text summary and determine the target multimedia information. The generation model is trained by the above-mentioned model training method.
[0050] The display module is used to present the recommendation information to the user using the target multimedia information and the text summary.
[0051] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for model training or information recommendation.
[0052] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method for model training or information recommendation.
[0053] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0054] In the model training method provided in this specification, target sample information is obtained, which includes text information and at least one multimedia information. Next, the target sample information is input into the generative model to be trained. For each multimedia information contained in the target sample information, the relevance between that multimedia information and the text information contained in the target sample information is determined, and this relevance is used as the corresponding relevance for that multimedia information. Then, based on the relevance corresponding to each multimedia information, target multimedia information is selected from the multimedia information, and a text summary corresponding to the target sample information is generated based on the text information contained in the target sample information. Finally, the generative model is trained with the optimization objective of maximizing the similarity between the target multimedia information and the labeled multimedia information, and the similarity between the text summary and the labeled text summary.
[0055] As can be seen from the model training method described above, this method can select target multimedia information from each multimedia information based on the correlation between each multimedia information and the text information contained in the target sample information. This allows users to vividly and figuratively understand the main idea of the text information contained in the target sample information through the target multimedia information, thereby improving the user experience. Attached Figure Description
[0056] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0057] Figure 1 A schematic flowchart illustrating the model training method provided in the embodiments of this specification;
[0058] Figure 2 A schematic diagram of a model structure provided for an embodiment of this specification;
[0059] Figure 3 A flowchart illustrating the method for recommending information provided in the embodiments of this specification;
[0060] Figure 4 A schematic diagram of the device structure for model training provided in the embodiments of this specification;
[0061] Figure 5 A schematic diagram of the recommended device structure for the information provided in the embodiments of this specification;
[0062] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this specification. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0064] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0065] In the embodiments of this specification, when generating text summaries based on recommendation information containing multiple multimedia and textual information, and determining the target multimedia information, a pre-trained generative model is required. The process of training the generative model will be described below. Figure 1 As shown.
[0066] Figure 1 The flowchart of the model training method provided in the embodiments of this specification is shown in the figure, which specifically includes the following steps:
[0067] S100: Obtain target sample information, which includes text information and at least one multimedia information.
[0068] In the embodiments of this specification, the execution subject of the model training method provided in this specification can be a server or an electronic device such as a desktop computer. For ease of description, the following description will only use a server as the execution subject to illustrate the model training method provided in this specification.
[0069] In the embodiments of this specification, the server can obtain target sample information, which includes text information and at least one multimedia information. The text information mentioned here may consist of several paragraphs, sentences, or words. The multimedia information mentioned here may refer to images, videos, etc. The target sample information mentioned here may refer to text information containing several multimedia elements. For example, a news article containing several images. Another example is an article containing several videos.
[0070] If the multimedia information is video, the server can extract several images from the video for subsequent model training.
[0071] S102: Input the target sample information into the generative model to be trained, so as to determine the correlation between the multimedia information and the text information contained in the target sample information for each multimedia information contained in the target sample information, and use it as the correlation corresponding to the multimedia information.
[0072] Currently, existing summary generation methods typically determine the text summary corresponding to the recommended content based on the text information within the recommended content. However, most recommended content often contains both multimedia and text information. Determining only the text summary may prevent users from intuitively understanding the key information in the recommended content, thus reducing their interest in it. Therefore, the server can determine the target multimedia information corresponding to the recommended content, allowing users to understand the key information in a more vivid and engaging way.
[0073] In the embodiments of this specification, the server can input target sample information into the generative model to be trained, so as to determine the correlation between the multimedia information and the text information contained in the target sample information for each multimedia information contained in the target sample information, and use the correlation corresponding to the multimedia information.
[0074] Specifically, the server can input the text information from the target sample into the text encoder of the generative model to be trained, determining the original text features corresponding to the text information in the target sample. Simultaneously, the server can input the multimedia information from the target sample into the multimedia encoder of the generative model to be trained, determining the original multimedia features corresponding to the multimedia information in the target sample.
[0075] The text encoder mentioned above can be a model that can encode text information, such as a bidirectional and auto-regressive transformer (BART) or a bidirectional encoder representation from transformers (BERT). This specification does not limit the specific form of the text encoder.
[0076] The multimedia encoders mentioned above can be models such as VGG19 and VGG16 that can encode multimedia information. This specification does not limit the specific form of the multimedia encoder.
[0077] Furthermore, the server can process the original multimedia features corresponding to each multimedia information contained in the target sample information to obtain the processed features corresponding to the multimedia information. It can also process the original text features corresponding to the text information contained in the target sample information to determine the processed features corresponding to the text information. The processed features corresponding to the multimedia information include multimedia query features, multimedia key features, and multimedia value features. The processed features corresponding to the text information include text query features, text key features, and text value features.
[0078] Specifically, the server can process the original multimedia features corresponding to the multimedia information, multiplying the original multimedia features by the multimedia query matrix, multimedia key matrix, and multimedia value matrix respectively to obtain the multimedia query features, multimedia key features, and multimedia value features. The specific formulas are as follows:
[0079]
[0080] In the above formula, X v It can be used to characterize the original multimedia features corresponding to multimedia information. It can be used to represent the multimedia query matrix corresponding to multimedia information. Q v It can be used to characterize multimedia query features corresponding to multimedia information. It can be used to represent the multimedia key matrix corresponding to multimedia information. K v It can be used to characterize the multimedia key features corresponding to multimedia information. It can be used to represent multimedia value matrices corresponding to multimedia information. V v It can be used to characterize the multimedia value features corresponding to multimedia information.
[0081] The server can process the original text features corresponding to the text information contained in the target sample information, and multiply the original text features by the text query matrix, text key matrix, and text value matrix respectively to obtain the text query features, text key features, and text value features. The specific formulas are as follows:
[0082]
[0083] In the above formula, X t It can be used to characterize the original text features corresponding to text information. It can be used to represent the text query matrix corresponding to text information. Q t It can be used to characterize the text query features corresponding to text information. It can be used to represent the text key matrix corresponding to text information. K t It can be used to characterize the text key features corresponding to text information. It can be used to represent the text value matrix corresponding to text information. t It can be used to characterize the text value features corresponding to text information.
[0084] Then, the server can determine the correlation between the multimedia information and the text information contained in the target sample information based on the correlation between the multimedia query features corresponding to the multimedia information and the text key features corresponding to the text information contained in the target sample information, as well as the correlation between the text query features corresponding to the text information contained in the target sample information and the multimedia key features corresponding to the multimedia information. This correlation is then used as the correlation of the multimedia information.
[0085] Finally, for each multimedia message, the server can determine the attention feature corresponding to that multimedia message based on its relevance and multimedia value features.
[0086] The specific formula for determining the attention features corresponding to this multimedia information is as follows:
[0087]
[0088] In the above formula, Q t It can be used to characterize the text query features corresponding to text information. K v T It can be used to characterize the transpose of multimedia key features corresponding to multimedia information. V v It can be used to characterize the multimedia value features corresponding to multimedia information. It can be used to characterize the square root of a vector dimension. Q t ·K v T This formula can be used to characterize the correlation between text query features corresponding to text information contained in target sample information and multimedia key features corresponding to multimedia information. It can be seen that this formula determines the weights corresponding to multimedia information through softmax calculation, and weights the multimedia value features corresponding to multimedia information to determine the attention features corresponding to multimedia information.
[0089] S104: Select target multimedia information from each multimedia information based on the relevance of each multimedia information, and generate a text summary corresponding to the target sample information based on the text information contained in the target sample information.
[0090] In the embodiments of this specification, the server can select target multimedia information from each multimedia information based on the relevance of each multimedia information.
[0091] Specifically, for each multimedia message, the server can determine its attention features based on its relevance and multimedia value features. Secondly, the server can determine comprehensive multimedia features based on the attention features of each multimedia message. Finally, the server can select the target multimedia message from among the multimedia messages based on the similarity between the original multimedia features and the comprehensive multimedia features.
[0092] There are several ways for the server to determine the comprehensive multimedia features based on the attention features corresponding to each multimedia message. For example, the server can concatenate the attention features corresponding to each multimedia message to obtain the comprehensive multimedia features. Another example is that the server can sum the data for each dimension of the attention features corresponding to each multimedia message to obtain the comprehensive multimedia features. This specification does not limit the specific form of determining the comprehensive multimedia features.
[0093] There are several ways for a server to select target multimedia information from various multimedia information sources. For example, the server can select multimedia information whose original multimedia features and comprehensive multimedia features have a similarity greater than a set threshold as target multimedia information. Another example is that the server can select multimedia information whose original multimedia features and comprehensive multimedia features have the highest similarity as target multimedia information. This specification does not limit the specific form in which target multimedia information is selected from various multimedia information sources.
[0094] In the embodiments of this specification, the server can generate a text summary corresponding to the target sample information based on the text information contained in the target sample information.
[0095] In practical applications, a summary text that encapsulates the main idea of the recommended content is often selected solely from the textual information within the recommended content, ignoring the multimedia information. This means the selected text summary is unrelated to the target multimedia message, making it difficult for users to understand the connection between the target multimedia information and the text summary. Therefore, the text summary pushed to users by the server should not only summarize the main idea of the textual information but also reflect the content of the target multimedia information.
[0096] In the embodiments of this specification, the server can determine the attention features corresponding to the text information contained in the target sample information based on the relevance of each multimedia information and the text value features corresponding to the text information contained in the target sample information.
[0097] The specific formula for the attention feature corresponding to the text information contained in the target sample information is as follows:
[0098]
[0099] In the above formula, Q v It can be used to characterize multimedia query features corresponding to multimedia information. K t T It can be used to represent the transpose of text key features corresponding to text information. V t It can be used to characterize the text value features corresponding to text information. It can be used to characterize the square root of a vector dimension. Q v ·K t T This can be used to characterize the correlation between multimedia query features corresponding to multimedia information and text key features corresponding to text information contained in the target sample information. It can be seen that the above formula can determine the weights corresponding to the text information through softmax calculation, and then weight the text value features corresponding to the text information to obtain the attention features corresponding to the text information contained in the target sample information.
[0100] Secondly, the text information contains several sub-texts. For each sub-text in the text information contained in the target sample information, the server can determine the comprehensive text features corresponding to that sub-text based on the correlation between that sub-text and other sub-texts in the text information contained in the target sample information.
[0101] Finally, the server can generate a text summary of the target sample information based on the comprehensive text features corresponding to each sub-text and the attention features corresponding to the text information contained in the target sample information. For example, the server can concatenate the comprehensive text features corresponding to each sub-text and the attention features corresponding to the text information contained in the target sample information, and input them into the text decoder of the generative model to be trained to generate a text summary of the target sample information.
[0102] Specifically, the server can process the original subtext features corresponding to each subtext in the text information contained in the target sample information to obtain the processed features corresponding to that subtext. The processed features of the subtext include: subtext query features, subtext key features, and subtext value features. The specific formula is as follows:
[0103]
[0104] In the above formula, X ti It can be used to characterize the original text features corresponding to the i-th sub-text in the text information contained in the target sample information. It can be used to represent the text query matrix corresponding to the i-th sub-text in the text information contained in the target sample information. Q ti It can be used to characterize the text query features corresponding to the i-th sub-text in the text information contained in the target sample information. It can be used to characterize the text key matrix corresponding to the i-th sub-text in the text information contained in the target sample information. K ti It can be used to characterize the text key features corresponding to the i-th sub-text in the text information contained in the target sample information. It can be used to represent the text value matrix corresponding to the i-th sub-text in the text information contained in the target sample information. V ti It can be used to characterize the text value features corresponding to the i-th sub-text in the text information contained in the target sample information.
[0105] Secondly, for each other subtext in the text information contained in the target sample information, the server can use the correlation between the subtext query feature corresponding to the subtext and the subtext key feature corresponding to the other subtext in the text information contained in the target sample information as the correlation of the other subtext.
[0106] Then, the server can determine the attention features corresponding to the other subtext based on the relevance of the other subtext and the subtext value features of the other subtext.
[0107] The specific formula for determining the attention features corresponding to the other subtext is as follows:
[0108]
[0109] In the above formula, Q ti It can be used to characterize the subtext query feature corresponding to the i-th subtext in the text information contained in the target sample information. K tn T It can be used to characterize the transpose of the subtext key features corresponding to the nth other subtext in the text information contained in the target sample information, where the nth other subtext refers to all subtexts except the ith subtext. V tn It can be used to characterize the subtext value feature corresponding to the nth other subtext in the text information contained in the target sample information. It can be used to represent the square root of the vector dimension. It can be seen that the above formula can use softmax to calculate the weights corresponding to the nth other sub-text, and then weight the sub-text value features corresponding to the nth other sub-text to obtain the attention features corresponding to the nth other sub-text in the text information contained in the target sample.
[0110] Finally, the server can determine the comprehensive text features corresponding to each of the other subtexts based on the attention features they correspond to.
[0111] There are several methods by which the server determines the comprehensive text features corresponding to a given sub-text. For example, the server can concatenate the attention features corresponding to all other sub-texts to obtain the comprehensive text features corresponding to that sub-text. Another example is that the server can sum the data for each dimension of the attention features corresponding to other sub-texts to obtain the comprehensive text features corresponding to that sub-text. This specification does not limit the specific form in which the comprehensive text features corresponding to the sub-text are determined.
[0112] S106: The generation model is trained with the optimization objective of maximizing the similarity between the target multimedia information and the labeled multimedia information, and the similarity between the text summary and the labeled text summary.
[0113] In the embodiments of this specification, the server may train the generative model with the optimization objective of maximizing the similarity between the target multimedia information and the labeled multimedia information, as well as the similarity between the text summary and the labeled text summary.
[0114] In practical applications, most sample information only contains labeled text summaries, while sample information containing both labeled multimedia information and labeled text summaries is relatively rare. Training the generative model with a small number of samples containing both labeled multimedia information and labeled text summaries leads to low accuracy in the model's results. Therefore, the server needs to select a target multimedia information from among the multimedia information in the unlabeled multimedia information samples and use this target multimedia information as the labeled multimedia information in the sample information.
[0115] In the embodiments of this specification, the server can obtain each initial sample information. For each initial sample information, the initial sample information includes an annotated text summary, the text information contained in the initial sample information, and several multimedia information.
[0116] Secondly, the server can input the initial sample information into a pre-trained annotation model. Based on the descriptive text corresponding to each multimedia information contained in the text information of the initial sample information, the model selects multimedia information from the multimedia information contained in the initial sample information to annotate it, thereby obtaining the target sample information corresponding to the initial sample information. The descriptive text corresponding to the multimedia information mentioned here can refer to text used to describe the content of the multimedia information.
[0117] Before inputting the initial sample information into the pre-trained labeled model, the server needs to process the initial sample information. Data processing methods include: removing sensitive words from the initial sample information, filling in missing data, removing duplicate images, standardizing and normalizing the data, and removing data other than multimedia information. For example, removing audio, website links, and other data from the initial sample information. Sensitive words mentioned here can refer to words with sensitive political leanings, violent tendencies, unhealthy connotations, or uncivilized language.
[0118] The server can use the above method to annotate the initial sample information of the labeled multimedia information in order to obtain a large number of target sample information with the labeled multimedia information.
[0119] In the embodiments of this specification, before the server annotates the initial sample information using the annotation model, it needs to rely on a pre-trained annotation model. The process of training the annotation model will be described below.
[0120] First, the server can obtain standard sample information, which includes annotated multimedia information, annotated text summaries, text information contained in the standard sample information, and several multimedia information items.
[0121] Secondly, the server can input standard sample information into the annotation model to be trained, and select the multimedia information to be verified from each multimedia information based on the similarity between the explanatory text corresponding to each multimedia information contained in the text information of the standard sample information and the annotated text summary contained in the standard sample information.
[0122] Finally, the server can train the annotation model with the optimization objective of minimizing the deviation between the multimedia information to be verified corresponding to the standard sample information and the annotated multimedia information contained in the standard sample information.
[0123] Specifically, there are several methods a server can use to determine the similarity between the descriptive text corresponding to each multimedia message and the annotated text summary contained in the standard sample information. For example, the server can use the ROUGE-2 method to determine similarity. ROUGE-2 refers to the proportion of shared bigram tokens between the descriptive text corresponding to the multimedia message and the annotated text summary contained in the standard sample information. Of course, the server can also use methods such as ROUGE-1, ROUGE-3, and ROUGE-X to determine similarity.
[0124] For example, the server can determine the similarity between the explanatory text corresponding to each multimedia message and the annotated text summary contained in the standard sample information based on the text features of the explanatory text corresponding to each multimedia message and the text features of the annotated text summary contained in the standard sample information.
[0125] In practical applications, most recommended content, in order to attract readers from the outset, often concisely places the core and important information at the beginning. Therefore, the importance of multimedia information can be determined based on its order of appearance within the recommended content. Based on this, the server can label multimedia information according to its position within the recommended content.
[0126] In the embodiments of this specification, the server can select the multimedia information to be verified corresponding to the standard sample information from each multimedia information based on at least one of the following: the similarity between the explanatory text corresponding to each multimedia information contained in the text information of the standard sample information and the annotated text summary contained in the standard sample information, and the position information of each multimedia information contained in the text information.
[0127] Furthermore, the server can select the multimedia information to be verified corresponding to the standard sample information from the multimedia information based on at least one of the following: the similarity between the explanatory text corresponding to each multimedia information contained in the text information and the annotated text summary contained in the standard sample information; the position information of each multimedia information contained in the text information; the position information of the explanatory text corresponding to each multimedia information contained in the standard sample information; and the text length of the explanatory text corresponding to each multimedia information contained in the standard sample information.
[0128] Figure 2 This is a schematic diagram of a model structure provided for an embodiment of this specification.
[0129] exist Figure 2In the process, the server inputs text information into the text encoder of the generative model to be trained, determining the original text features corresponding to the text information in the target sample information. Simultaneously, the server can input multimedia information from the target sample information into the multimedia encoder of the generative model to be trained, determining the original multimedia features corresponding to the multimedia information in the target sample information.
[0130] Secondly, the server can input the original text features corresponding to the text information in the target sample information, and the original multimedia features corresponding to the multimedia information in the target sample information, into the first cross-attention network in the generative model to be trained. Based on the relevance of the multimedia information and the multimedia value features corresponding to the multimedia information, the server determines the attention features corresponding to the multimedia information. Based on the relevance of each multimedia information and the text value features corresponding to the text information contained in the target sample information, the server determines the attention features corresponding to the text information contained in the target sample information.
[0131] Simultaneously, the server can input the original sub-text features corresponding to each sub-text in the text information contained in the target sample information into the second cross-attention network in the generative model to be trained. For each sub-text in the text information contained in the target sample information, the server determines the comprehensive text features corresponding to the sub-text based on the correlation between the sub-text and other sub-texts in the text information contained in the target sample information.
[0132] Then, the server can input the attention features corresponding to each multimedia information into the multimedia decoder in the generative model to be trained, determine the comprehensive multimedia features, and select the target multimedia information from each multimedia information based on the similarity between the original multimedia features and the comprehensive multimedia features corresponding to each multimedia information.
[0133] Meanwhile, the server can input the attention features corresponding to the text information contained in the target sample information, as well as the comprehensive text features corresponding to each sub-text, into the text decoder in the generative model to be trained, to generate a text summary corresponding to the target sample information.
[0134] As can be seen from the above process, this method can select target multimedia information from each multimedia information based on the correlation between each multimedia information and the text information contained in the target sample information. This allows users to vividly and figuratively understand the main idea of the text information contained in the target sample information through the target multimedia information, thereby improving the user experience.
[0135] The embodiments in this specification, after the generative model has been trained, can generate text summaries and determine target multimedia information, such as... Figure 3 As shown.
[0136] Figure 3 A flowchart illustrating the method for recommending information provided in the embodiments of this specification is shown, specifically including:
[0137] S300: Obtain recommendation information containing several multimedia and text information.
[0138] S302: The recommendation information containing several multimedia information and text information is input into the generation model to generate a text summary and determine the target multimedia information. The generation model is trained by the above-mentioned model training method.
[0139] S304: Display the recommendation information to the user using the target multimedia information and the text summary.
[0140] In the embodiments described in this specification, the server can obtain recommendation information containing several multimedia information items and text information. Next, the server can input the recommendation information containing the multimedia information items and text information into a generative model to generate a text summary and determine the target multimedia information. Finally, the server can display the recommendation information to the user using the target multimedia information and the text summary.
[0141] The server displays the target multimedia information and text summary corresponding to the recommended information to the user. After browsing the target multimedia information and text summary, the user can click on the corresponding target multimedia information and text summary. Upon receiving the click, the server then displays the recommended information to the user.
[0142] The above describes a model training method provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding model training apparatus, such as... Figure 4 As shown.
[0143] Figure 4 A schematic diagram of the device structure for model training provided in the embodiments of this specification, specifically including:
[0144] The acquisition module 400 is used to acquire target sample information, which includes text information and at least one multimedia information.
[0145] The determination module 402 is used to input the target sample information into the generative model to be trained, so as to determine the correlation between the multimedia information and the text information contained in the target sample information for each multimedia information contained in the target sample information, and use it as the correlation corresponding to the multimedia information.
[0146] The generation module 404 is used to select target multimedia information from each multimedia information according to the relevance of each multimedia information, and to generate a text summary corresponding to the target sample information according to the text information contained in the target sample information.
[0147] The training module 406 is used to train the generative model with the optimization objective of maximizing the similarity between the target multimedia information and the labeled multimedia information, and the similarity between the text summary and the labeled text summary.
[0148] Optionally, the determining module 402 is specifically configured to: process the original multimedia features corresponding to each multimedia information contained in the target sample information to obtain the processed features corresponding to the multimedia information; and process the original text features corresponding to the text information contained in the target sample information to determine the processed features corresponding to the text information. The processed features corresponding to the multimedia information include multimedia query features and multimedia key features, and the processed features corresponding to the text information include text query features and text key features. Based on the correlation between the multimedia query features corresponding to the multimedia information and the text key features corresponding to the text information contained in the target sample information, and the correlation between the text query features corresponding to the text information contained in the target sample information and the multimedia key features corresponding to the multimedia information, the correlation between the multimedia information and the text information contained in the target sample information is determined as the correlation corresponding to the multimedia information.
[0149] Optionally, the processed features corresponding to the multimedia information further include: multimedia value features;
[0150] The generation module 404 is specifically used to: for each multimedia information, determine the attention features corresponding to the multimedia information based on the relevance and multimedia value features corresponding to the multimedia information; determine the comprehensive multimedia features based on the attention features corresponding to each multimedia information; and select the target multimedia information from each multimedia information based on the similarity between the original multimedia features corresponding to each multimedia information and the comprehensive multimedia features.
[0151] Optionally, the processed features corresponding to the text information further include: text value features, wherein the text information contains several sub-texts;
[0152] The generation module 404 is specifically used to: determine the attention features corresponding to the text information contained in the target sample information based on the relevance of each multimedia information and the text value features corresponding to the text information contained in the target sample information; for each sub-text in the text information contained in the target sample information, determine the comprehensive text features corresponding to the sub-text based on the relevance between the sub-text and other sub-texts in the text information contained in the target sample information; and generate a text summary corresponding to the target sample information based on the comprehensive text features corresponding to each sub-text and the attention features corresponding to the text information contained in the target sample information.
[0153] Optionally, the generation module 404 is specifically configured to: process the original subtext features corresponding to each subtext in the text information contained in the target sample information to obtain the processed features corresponding to the subtext; the processed features corresponding to the subtext include: subtext query features, subtext key features, and subtext value features; for each other subtext in the text information contained in the target sample information, determine the correlation between the subtext query features corresponding to the subtext and the subtext key features corresponding to the other subtext in the text information contained in the target sample information as the correlation of the other subtext; determine the attention features corresponding to the other subtext based on the correlation and the subtext value features corresponding to the other subtext; and determine the comprehensive text features corresponding to the subtext based on the attention features corresponding to each other subtext.
[0154] Optionally, the acquisition module 400 is specifically used to acquire each initial sample information. For each initial sample information, the initial sample information contains an annotated text summary, text information contained in the initial sample information, and several multimedia information. The initial sample information is input into a pre-trained annotation model to select multimedia information for annotation of the initial sample information from the multimedia information contained in the initial sample information according to the explanatory text corresponding to each multimedia information contained in the text information of the initial sample information, so as to annotate the initial sample information and obtain the target sample information corresponding to the initial sample information.
[0155] Optionally, the acquisition module 400 is specifically used to acquire standard sample information, which includes labeled multimedia information, labeled text summaries, text information contained in the standard sample information, and several multimedia information. The standard sample information is input into the annotation model to be trained. Based on the similarity between the descriptive text corresponding to each multimedia information contained in the text information in the standard sample information and the labeled text summaries contained in the standard sample information, the multimedia information to be verified corresponding to the standard sample information is selected from each multimedia information. The goal is to minimize the deviation between the multimedia information to be verified corresponding to the standard sample information and the labeled multimedia information contained in the standard sample information, and the annotation model is trained.
[0156] Optionally, the acquisition module 400 is specifically used to select the multimedia information to be verified corresponding to the standard sample information from each multimedia information based on at least one of the following: the similarity between the explanatory text corresponding to each multimedia information contained in the text information of the standard sample information and the annotated text summary contained in the standard sample information, and the position information of each multimedia information contained in the standard sample information in the text information.
[0157] Figure 5 The recommended device structure diagrams for the embodiments provided in this specification specifically include:
[0158] The acquisition module 500 is used to acquire recommendation information containing several multimedia information and text information;
[0159] The generation module 502 is used to input the recommendation information containing several multimedia information and text information into the generation model to generate a text summary and determine the target multimedia information. The generation model is trained by the above-mentioned model training method.
[0160] The display module 504 is used to display the recommendation information to the user through the target multimedia information and the text summary.
[0161] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The provided model training method or the above Figure 3 The information provided is recommended by the method.
[0162] This instruction manual also provides Figure 6 The diagram shows the structure of the electronic device. Figure 6At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 1 The model training method described above or the above Figure 3 The information provided recommends certain methods. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0163] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.
[0164] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0165] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0166] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0167] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0168] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0169] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0170] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0171] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0172] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0173] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0174] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0175] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0176] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0177] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0178] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0179] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for training a model, characterized in that, include: Obtain target sample information, which includes text information and at least one multimedia information; The target sample information is input into the generative model to be trained, so as to determine the correlation between the multimedia information and the text information contained in the target sample information for each multimedia information contained in the target sample information, and use it as the correlation corresponding to the multimedia information. Based on the relevance of each multimedia information, target multimedia information is selected from each multimedia information, and a text summary corresponding to the target sample information is generated based on the text information contained in the target sample information; the generation model is trained with the optimization objective of maximizing the similarity between the target multimedia information and the labeled multimedia information, and the similarity between the text summary and the labeled text summary. For each multimedia information contained in the target sample information, the correlation between the multimedia information and the text information contained in the target sample information is determined as the correlation corresponding to the multimedia information. Specifically, this includes: processing the original multimedia features corresponding to each multimedia information contained in the target sample information to obtain the processed features corresponding to the multimedia information; and processing the original text features corresponding to the text information contained in the target sample information to determine the processed features corresponding to the text information. The processed features corresponding to the multimedia information include multimedia query features and multimedia key features, and the processed features corresponding to the text information include text query features and text key features. Based on the correlation between the multimedia query features corresponding to the multimedia information and the text key features corresponding to the text information contained in the target sample information, and the correlation between the text query features corresponding to the text information contained in the target sample information and the multimedia key features corresponding to the multimedia information, the correlation between the multimedia information and the text information contained in the target sample information is determined as the correlation corresponding to the multimedia information.
2. The method as described in claim 1, characterized in that, The processed features corresponding to the multimedia information also include: multimedia value features; selecting target multimedia information from each multimedia information based on the relevance of each multimedia information, specifically including: for each multimedia information, determining the attention features corresponding to the multimedia information based on the relevance of the multimedia information and the multimedia value features corresponding to the multimedia information; determining comprehensive multimedia features based on the attention features corresponding to each multimedia information; and selecting target multimedia information from each multimedia information based on the similarity between the original multimedia features corresponding to each multimedia information and the comprehensive multimedia features.
3. The method as described in claim 1, characterized in that, The processed features corresponding to the text information further include: text value features, wherein the text information contains several sub-texts; generating a text summary corresponding to the target sample information based on the text information contained in the target sample information specifically includes: determining the attention features corresponding to the text information contained in the target sample information based on the relevance of each multimedia information and the text value features corresponding to the text information contained in the target sample information; for each sub-text in the text information contained in the target sample information, determining the comprehensive text features corresponding to the sub-text based on the relevance between the sub-text and other sub-texts in the text information contained in the target sample information; and generating a text summary corresponding to the target sample information based on the comprehensive text features corresponding to each sub-text and the attention features corresponding to the text information contained in the target sample information.
4. The method as described in claim 3, characterized in that, For each subtext in the text information contained in the target sample information, the comprehensive text features corresponding to the subtext are determined based on the correlation between the subtext and other subtexts in the text information contained in the target sample information. Specifically, this includes: processing the original subtext features corresponding to each subtext in the text information contained in the target sample information to obtain the processed features corresponding to the subtext, wherein the processed features corresponding to the subtext include: subtext query features, subtext key features, and subtext value features; for each other subtext in the text information contained in the target sample information, the correlation between the subtext query features corresponding to the subtext and the subtext key features corresponding to the other subtext in the text information contained in the target sample information is used as the correlation of the other subtext; based on the correlation and the subtext value features corresponding to the other subtext, the attention features corresponding to the other subtext are determined; and based on the attention features corresponding to each other subtext, the comprehensive text features corresponding to the subtext are determined.
5. The method as described in claim 1, characterized in that, Obtaining target sample information specifically includes: obtaining initial sample information, each initial sample containing a labeled text summary, text information contained in the initial sample, and several multimedia information; inputting the initial sample information into a pre-trained annotation model, and selecting multimedia information for annotation based on the explanatory text corresponding to each multimedia information contained in the text information of the initial sample, thereby annotating the initial sample information and obtaining the target sample information corresponding to the initial sample information.
6. The method as described in claim 5, characterized in that, The training of the annotation model specifically includes: acquiring standard sample information, which contains annotated multimedia information, annotated text summaries, text information contained in the standard sample information, and several multimedia information items; inputting the standard sample information into the annotation model to be trained, and selecting the multimedia information to be verified corresponding to the standard sample information from each multimedia information item based on the similarity between the explanatory text corresponding to each multimedia information item contained in the text information in the standard sample information and the annotated text summaries contained in the standard sample information; and training the annotation model with the optimization objective of minimizing the deviation between the multimedia information to be verified corresponding to the standard sample information and the annotated multimedia information contained in the standard sample information.
7. The method as described in claim 6, characterized in that, Based on the similarity between the descriptive text corresponding to each multimedia information contained in the text information of the standard sample information and the annotated text summary contained in the standard sample information, the multimedia information to be verified corresponding to the standard sample information is selected from each multimedia information. Specifically, this includes selecting the multimedia information to be verified corresponding to the standard sample information from each multimedia information based on at least one of the following: the similarity between the descriptive text corresponding to each multimedia information contained in the text information of the standard sample information and the annotated text summary contained in the standard sample information, and the position information of each multimedia information contained in the standard sample information in the text information.
8. A method for information recommendation, characterized in that, include: Retrieve recommendation information that includes several multimedia and text information; The recommendation information, which contains several multimedia information and text information, is input into the generation model to generate a text summary and determine the target multimedia information. The generation model is trained by the method described in any one of claims 1 to 5. The recommendation information is then displayed to the user using the target multimedia information and the text summary.
9. A device for model training, characterized in that, include: An acquisition module is used to acquire target sample information, which includes text information and at least one multimedia information. The determination module is used to input the target sample information into the generative model to be trained, so as to determine the correlation between the multimedia information and the text information contained in the target sample information for each multimedia information contained in the target sample information, and use it as the correlation corresponding to the multimedia information. The generation module is used to select target multimedia information from each multimedia information according to the relevance of each multimedia information, and to generate a text summary corresponding to the target sample information according to the text information contained in the target sample information. The training module is used to train the generative model with the optimization objective of maximizing the similarity between the target multimedia information and the labeled multimedia information, and the similarity between the text summary and the labeled text summary. For each multimedia information contained in the target sample information, the correlation between the multimedia information and the text information contained in the target sample information is determined as the correlation corresponding to the multimedia information. Specifically, this includes: processing the original multimedia features corresponding to each multimedia information contained in the target sample information to obtain the processed features corresponding to the multimedia information; and processing the original text features corresponding to the text information contained in the target sample information to determine the processed features corresponding to the text information. The processed features corresponding to the multimedia information include multimedia query features and multimedia key features, and the processed features corresponding to the text information include text query features and text key features. Based on the correlation between the multimedia query features corresponding to the multimedia information and the text key features corresponding to the text information contained in the target sample information, and the correlation between the text query features corresponding to the text information contained in the target sample information and the multimedia key features corresponding to the multimedia information, the correlation between the multimedia information and the text information contained in the target sample information is determined as the correlation corresponding to the multimedia information.
10. An information recommendation device, characterized in that, include: The acquisition module is used to acquire recommendation information that contains several multimedia information and text information; The generation module is used to input the recommendation information containing several multimedia information and text information into the generation model to generate a text summary and determine the target multimedia information. The generation model is trained by the method described in any one of claims 1 to 5. The display module is used to display the recommendation information to the user using the target multimedia information and the text summary.
11. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 8.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method of Generating Graph and Text Abstracts
CN109508400A
System and method for associating textual summaries with content media
CN110377789A