Title generation method and apparatus, electronic device, computer readable medium

By acquiring multimodal corpora, performing type matching, and calculating popularity and relevance scores, personalized titles related to the target site are generated, solving the problem of lengthy video titles on e-commerce platforms and improving user experience.

CN114329206BActive Publication Date: 2026-01-16BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111635525.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2026-01-16
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The titles of product introduction videos on e-commerce platforms are often lengthy and fail to effectively describe the video content, thus failing to attract user interest.

Method used

By acquiring multimodal corpora, performing type matching, calculating popularity and relevance scores, and generating personalized titles related to the target site.

Benefits of technology

The title was made more descriptive and attractive, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329206B_ABST
    Figure CN114329206B_ABST
Patent Text Reader

Abstract

The present disclosure provides a title generation method and device, relating to the technical fields of deep learning, image processing, speech recognition and the like. The specific implementation scheme is: obtaining a plurality of multi-modal corpora to be recognized; matching the multi-modal corpora with historical corpora of different types obtained by pre-clustering to determine the types of the multi-modal corpora; calculating a set of heat degree scores of the multi-modal corpora in the types, each heat degree score in the set of heat degree scores being used to represent the similarity between each multi-modal corpus and the historical corpora; calculating a set of correlation scores of the multi-modal corpora based on the title of a target site, each correlation score in the set of correlation scores being used to represent the correlation between each multi-modal corpus and the title of the target site; and generating a title corresponding to each multi-modal corpus based on the set of correlation scores, the set of heat degree scores and the multi-modal corpora. This embodiment realizes personalized recommendation of the title.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, in particular to the technical field of deep learning, image processing, speech recognition and the like, and especially relates to a title generation method and device, electronic equipment, computer readable medium and computer program product. BACKGROUND

[0002] In the process of selling on an e-commerce platform, in order to attract the interest of buyers and to display the related attributes of the goods more richly, the e-commerce platform supports sellers to upload introduction videos related to the goods. In the process of selecting a topic for the video, the seller will usually select a too lengthy title of the goods as the title of the video. However, the title of the video cannot well describe the content of the video on the one hand, and cannot arouse the interest of users on the other hand. SUMMARY

[0003] A title generation method and device, electronic equipment, computer readable medium and computer program product are provided.

[0004] According to a first aspect, a title generation method is provided. The method comprises: obtaining a multi-modal corpus to be recognized, the multi-modal corpus corresponding to a corpus of at least one modality information; matching the multi-modal corpus with historical corpora of different types obtained by pre-clustering to determine the type of the multi-modal corpus; calculating a set of heat degree scores of the multi-modal corpus in the type, each heat degree score in the set of heat degree scores being used to represent the similarity between each multi-modal corpus and the historical corpus; based on the title of a target site, calculating a set of relevance scores of the multi-modal corpus, each relevance score in the set of relevance scores being used to represent the relevance between each multi-modal corpus and the title of the target site; and generating a title corresponding to the multi-modal corpus based on the set of relevance scores, the set of heat degree scores and the multi-modal corpus.

[0005] According to a second aspect, a title generation device is provided. The device comprises: a corpus obtaining unit configured to obtain a multi-modal corpus to be recognized, the multi-modal corpus corresponding to a corpus of at least one modality information; a type matching unit configured to match the multi-modal corpus with historical corpora of different types obtained by pre-clustering to determine the type of the multi-modal corpus; a heat calculating unit configured to calculate a set of heat degree scores of the multi-modal corpus in the type, each heat degree score in the set of heat degree scores being used to represent the similarity between each multi-modal corpus and the historical corpus; a relevance calculating unit configured to calculate a set of relevance scores of the multi-modal corpus based on the title of a target site, each relevance score in the set of relevance scores being used to represent the relevance between each multi-modal corpus and the title of the target site; and a generating unit configured to generate a title corresponding to the multi-modal corpus based on the set of relevance scores, the set of heat degree scores and the multi-modal corpus.

[0006] According to a third aspect, an electronic device is provided, comprising at least one processor; and a memory connected with the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any implementation of the first aspect.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to cause a computer to perform the method according to any implementation of the first aspect.

[0008] According to a fifth aspect, a computer program product is provided, comprising a computer program, the computer program being used to implement the method according to any implementation of the first aspect when executed by a processor.

[0009] The title generation method and device provided by the embodiments of the present disclosure first acquire a plurality of modalities of corpus to be identified; secondly, the plurality of modalities of corpus are matched with historical corpus of different types obtained by pre-clustering to determine the type of the plurality of modalities of corpus; thirdly, a set of heat degree scores of the plurality of modalities of corpus is calculated, each heat degree score in the set of heat degree scores being used to represent the similarity between each plurality of modalities of corpus and the historical corpus; fourthly, a set of correlation scores of the plurality of modalities of corpus is calculated based on the title of the target site, each correlation score in the set of correlation scores being used to represent the correlation between each plurality of modalities of corpus and the title of the target site; and finally, a title corresponding to the plurality of modalities of corpus is generated based on the set of correlation scores, the set of heat degree scores, and the plurality of modalities of corpus. Thus, based on the heat degree scores and the correlation scores of the plurality of modalities of corpus, the title corresponding to the plurality of modalities of corpus related to the title of the target site and the historical type is generated based on the heat degree of the historical corpus and the title, the personalized recommendation of the title is realized, and the user experience is improved.

[0010] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:

[0012] Figure 1 is a flowchart of an embodiment of the title generation method according to the present disclosure;

[0013] Figure 2 is a flowchart of an embodiment of the semantic classification model training method in the present disclosure;

[0014] Figure 3 is a flowchart of an embodiment of a method for generating titles corresponding to multi-modal corpora in the present disclosure;

[0015] Figure 4 is a structural schematic diagram of a method for generating titles corresponding to multi-modal corpora in the present disclosure;

[0016] Figure 5 is a structural schematic diagram of an embodiment of a title generation apparatus according to the present disclosure;

[0017] Figure 6 is a block diagram of an electronic device for implementing the title generation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] Exemplary embodiments of the present disclosure are described herein below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.

[0019] Figure 1 A flowchart 100 of an embodiment of a title generation method according to the present disclosure is shown, the title generation method comprising the following steps:

[0020] In step 101, a plurality of multi-modal corpora to be recognized are obtained.

[0021] In this embodiment, the multi-modal corpora are corpora corresponding to at least one type of modal information, where the modal information has various forms in different applications or websites, for example, the modal information can include one or more of text, video, picture, and audio. By analyzing different modal information, the corpora corresponding to the modal information can be obtained, for example, by performing entity and relationship extraction on text on a webpage to obtain the multi-modal corpora corresponding to the text. For another example, by performing title extraction on an introduction video on an application to obtain the multi-modal corpora corresponding to the video. For another example, in the search field, by searching for all inputs to a search engine to obtain the multi-modal corpora corresponding to the search engine.

[0022] In this embodiment, there are a plurality of multi-modal corpora, and by performing removal, word segmentation, entity extraction, and the like on the plurality of multi-modal corpora, the titles corresponding to the multi-modal corpora can be obtained. When the number of multi-modal corpora is relatively large, the multi-modal corpora with high scores can be determined by performing hotness and correlation scoring on each multi-modal corpus, and the titles can be obtained based on the multi-modal corpora with high scores.

[0023] In this embodiment, the plurality of multi-modal corpora to be identified can be corpora on a relevant website or application extracted at a certain time point, the title of a user interested site is obtained, and the multi-modal corpora are processed based on the obtained title of the user interested site, so that the title of the user most interested in and related to the user interested site can be obtained.

[0024] Optionally, the plurality of multi-modal corpora to be identified can also be corpora obtained after processing of initial historical corpora, wherein the initial historical corpora are original corpora obtained on an application or a website, and the initial historical corpora are also corpora without type division. For example, for the search field, the multi-modal corpora are obtained by sorting the questions from the user's seller's business scope of the target merchant portal according to the search frequency.

[0025] Further, the plurality of multi-modal corpora to be identified can also be some corpora randomly extracted from different types of historical corpora obtained by pre-clustering. The randomly extracted corpora belong to different types, and the types of the multi-modal corpora can be determined by the different types of historical corpora obtained by pre-clustering.

[0026] Step 102, matching the multi-modal corpora with the different types of historical corpora obtained by pre-clustering to determine the type of the multi-modal corpora.

[0027] In this embodiment, the different types of historical corpora can be a large amount of multi-modal historical corpora collected on a target site in a preset time period, for example, the historical corpora are questions searched by users in the past three months obtained through a search engine. The obtained historical corpora are clustered by a clustering algorithm, and different types of historical corpora can be obtained. The types of the historical corpora can be divided based on the functions or attributes or features of the historical corpora. For example, the types of the historical corpora include introduction type, detail type, and question type.

[0028] In this embodiment, the types of the historical corpora can also be obtained by querying a corpus type table, and the corpus type table is used to represent the corresponding relationship between the historical corpora and the types.

[0029] In this embodiment, matching the multi-modal corpora with the different types of historical corpora obtained by pre-clustering includes: matching each multi-modal corpus in the plurality of multi-modal corpora to be identified with the different types of historical corpora obtained by pre-clustering respectively, when any one multi-modal corpus is the same as or similar to a historical corpus, the type of the matched historical corpus is taken as the type of the multi-modal corpus.

[0030] In some optional implementations of the embodiment, the matching of the multi-modal corpus with the pre-clustered historical corpora of different types to determine the type of the multi-modal corpus comprises: matching the multi-modal corpus with the historical corpora; and in response to the successful matching of the multi-modal corpus with any type of historical corpus, taking the type of the successfully matched historical corpus as the type of the multi-modal corpus.

[0031] In the optional implementation, the type of the successfully matched historical corpus is taken as the type of the multi-modal corpus after the successful matching of the multi-modal corpus with any type of historical corpus, which provides an optional implementation for determining the type of the multi-modal corpus and ensures the reliability of the question generation.

[0032] Optionally, the matching of the multi-modal corpus with the pre-clustered historical corpora of different types to determine the type of the multi-modal corpus comprises: matching each of the plurality of multi-modal corpora to be identified with the historical corpora; and in response to the completion of the matching of all the multi-modal corpora with the historical corpora, taking the types of all the successfully matched historical corpora as the type of the multi-modal corpus.

[0033] In the optional implementation, the types of all the successfully matched historical corpora are taken as the type of the multi-modal corpus, which improves the diversity of the types of the multi-modal corpora.

[0034] Step 103: calculating a set of heat scores of the multi-modal corpus in the type.

[0035] In the embodiment, each heat score in the set of heat scores is used to represent the similarity between the multi-modal corpus and the historical corpus.

[0036] In the embodiment, after obtaining the type of the multi-modal corpus, the execution subject on which the title generation method runs can match all the historical corpora corresponding to the type of the multi-modal corpus with the multi-modal corpus respectively to calculate the set of heat scores of the multi-modal corpus in the type.

[0037] In the embodiment, the calculation of the set of heat scores of the multi-modal corpus in the type comprises: obtaining all the historical corpora corresponding to the type, calculating a similarity value between each historical corpus and the multi-modal corpus, multiplying the obtained similarity value with the normalized heat value of each historical corpus, and then adding them to obtain the set of heat scores of the multi-modal corpus in the type. It should be noted that there are multiple multi-modal corpora, and after the calculation of the heat scores of each multi-modal corpus by adding the heat scores of all the historical corpora corresponding to the type and each multi-modal corpus, the heat scores of all the multi-modal corpora are combined to obtain the set of heat scores of the multi-modal corpus in the type.

[0038] In this embodiment, the calculation method of the similarity value can be various, such as the Euclidean distance algorithm, the cosine of the included angle algorithm, etc.

[0039] In some optional implementations of this embodiment, the above calculating the heat score set of the multi-modal corpus in the type includes: in response to at least one historical corpus matched successfully with the multi-modal corpus, calculating the cosine distance value of the vector of each multi-modal corpus of the multi-modal corpus and the vector of each historical corpus in the type, and adding the product of the cosine distance value and the normalized heat value of each historical corpus to obtain the heat score of the multi-modal corpus, and combining all heat scores of the multi-modal corpus to obtain the heat score set of the multi-modal corpus in the type, wherein the normalized heat value is used to represent the popularity of the modal information corresponding to the historical corpus.

[0040] In this optional implementation, the normalized heat value can be calculated by obtaining the video click volume and mapping it into a fixed interval by using a function.

[0041] In this optional implementation, for each multi-modal corpus in the plurality of multi-modal corpora to be recognized, the cosine distance value of the vector of each multi-modal corpus of the multi-modal corpus and the vector of each historical corpus in the type is calculated, and the product of all cosine distance values corresponding to the multi-modal corpus and the normalized heat value of each historical corpus is added to obtain the heat score of the multi-modal corpus.

[0042] As shown in formula (1), the vector q t of the multi-modal corpus is calculated. n The cosine distance value is calculated. The cosine distance value is multiplied by the normalized heat value h n of the historical corpus to obtain the heat score corresponding to the historical corpus, and the heat scores corresponding to all historical corpora in the type are added to obtain the heat score of the multi-modal corpus.

[0043]

[0044] In this optional implementation, the similarity between the multi-modal corpus and the historical corpus is determined by calculating the cosine distance value of the vector of each multi-modal corpus of the multi-modal corpus and the vector of the historical corpus, which is simple and convenient and ensures the reliability of the calculation of the heat score of the multi-modal corpus.

[0045] In step 104, the relevance score set of the multi-modal corpus is calculated based on the title of the target site.

[0046] In this embodiment, each relevance score in the relevance score set is used to represent the correlation degree between each multi-modal corpus and the title of the target site.

[0047] In this embodiment, the target site is a site where a target title is to be generated, which can be a web page, or an application, etc. By calculating the relevance score set of the multi-modal corpus, the relevance degree of the multi-modal corpus and the title of the target site can be determined, thereby providing a reliable basis for generating the title corresponding to the multi-modal corpus.

[0048] The above calculating the relevance score set of the multi-modal corpus based on the title of the target site includes: determining a similarity value of all titles of the target site and the multi-modal corpus, and taking the similarity value as the relevance score in response to the similarity value being greater than a preset threshold.

[0049] Optionally, in some optional implementations of this embodiment, the above calculating the relevance score set of the multi-modal corpus based on the title of the target site includes: calculating a cosine distance value between the vector of each multi-modal corpus of the multi-modal corpus and the vector of each title of the target site, and adding the product of the cosine distance value and the normalized heat value of each title to obtain the relevance score of the multi-modal corpus, and combining all relevance scores of the multi-modal corpus to obtain the relevance score set of the multi-modal corpus.

[0050] In this optional implementation, for each multi-modal corpus of the plurality of multi-modal corpora to be recognized, the cosine distance value is calculated between the vector of each multi-modal corpus of the multi-modal corpus and the vector of each title of the target site, and the product of all cosine distance values corresponding to the multi-modal corpus and the normalized heat value of each title is added to obtain the heat score of the multi-modal corpus.

[0051] As shown in formula (2), the vector q t of the multi-modal corpus is calculated with the vector t n of one title of the target site to obtain the cosine distance value. The cosine distance value is multiplied by the normalized heat value h n of the title to obtain the relevance score corresponding to the title, the cosine distance value is calculated between the vector of the multi-modal corpus and the vector of another title of the target site, and the cosine distance value is multiplied by the normalized heat value of the another title to obtain the relevance score corresponding to the another title, thereby adding the relevance scores corresponding to all titles of the target site to obtain the relevance score of the multi-modal corpus.

[0052]

[0053] In this optional implementation, the relevance degree of the multi-modal corpus and the title of the target site is determined by calculating the cosine distance value between the vector of each multi-modal corpus of the multi-modal corpus and the vector of each title of the target site, which is simple and convenient, and ensures the reliability of the calculation of the relevance score set of the multi-modal corpus.

[0054] In step 105, a title corresponding to the multi-modal corpus is generated based on the relevance score set, the hotness score set, and the multi-modal corpus.

[0055] In this embodiment, each multi-modal corpus in the plurality of multi-modal corpora to be identified has a relevance score and a hotness score. The relevance score set and the hotness score of each multi-modal corpus in the plurality of multi-modal corpora to be identified are weighted and summed to obtain a score of the multi-modal corpus. The score of the multi-modal corpus can reflect the importance of the generated title.

[0056] The title generation method provided by the embodiments of the present disclosure first acquires a multi-modal corpus to be identified. Then, the multi-modal corpus is matched with different types of historical corpora obtained by pre-clustering to determine the type of the multi-modal corpus. Next, a hotness score set of the multi-modal corpus is calculated. Each hotness score in the hotness score set is used to represent the similarity between each multi-modal corpus and the historical corpus. Then, a relevance score set of the multi-modal corpus is calculated based on the title of the target site. Each relevance score in the relevance score set is used to represent the relevance between each multi-modal corpus and the title of the target site. Finally, a title corresponding to the multi-modal corpus is generated based on the relevance score set, the hotness score set, and the multi-modal corpus. Thus, based on the hotness score and the relevance score of the multi-modal corpus, the title corresponding to the multi-modal corpus related to the title of the target site and the historical type is generated based on the hotness of the historical corpus and the hotness of the title, thereby realizing personalized recommendation of the title and improving user experience.

[0057] In another embodiment of the present disclosure, a pre-trained semantic classification model can be used to cluster the initial historical corpus to obtain different types of historical corpora. Specifically, Figure 2 An embodiment of the flow 200 of the semantic classification model training method is shown. The semantic classification model training method includes the following steps:

[0058] In step 201, the association relationship between the sample corpus and the sample title in a preset time period is obtained.

[0059] In this embodiment, the preset time period can be determined based on development requirements. For example, the preset time period is half a year. The sample corpus can be a corpus corresponding to multiple different modal information generated when a user accesses an application or a webpage related to the title to be generated. For example, when a user searches for a question on a search webpage, a text corpus is obtained. For another example, when a user watches a video, an evaluation text corresponding to the video is output. The sample title is a title corresponding to different modal information. For example, when a user searches for a webpage, the text title of the webpage is the sample title. For another example, when a user watches a video, the title of the video is the sample title.

[0060] Optionally, in the search field, a user initiates a search behavior by inputting a question to a search engine and then clicking a relevant article containing a title. All the search behaviors of the user converge to obtain that one question corresponds to multiple titles, and one title also corresponds to multiple questions. The question of the user is a sample corpus, the multiple titles corresponding to the question are sample titles, and the one question corresponding to multiple titles or one title corresponding to multiple questions is an association relationship between the sample corpus and the sample title.

[0061] In step 202, based on the association relationship, the sample corpus and the sample title corresponding to the sample corpus are organized into a bipartite graph.

[0062] In this embodiment, the sample corpus and the sample title corresponding to the sample corpus are organized into a bipartite graph based on the association relationship between the sample corpus and the sample title obtained in step 201.

[0063] A bipartite graph, also known as a bicolored graph, is a special model in graph theory. Let G=(V,E) be an undirected graph. If the vertices V can be divided into two disjoint subsets (A,B), and each edge (i,j) in the graph is associated with two vertices i and j belonging to the two different vertex sets (i in A,j in B), then the graph G is called a bipartite graph.

[0064] Optionally, the bipartite graph can be input into a pre-training model based on Ernie (Enhanced Representation from Knowledge Integration) to obtain the sample corpus and the type of the sample corpus, wherein the pre-training model is a model constructed by big data and trained for unsupervised tasks.

[0065] In step 203, the bipartite graph is clustered to obtain a clustered bipartite graph.

[0066] In this embodiment, different bipartite graphs can be clustered by using a clustering method according to different features of the bipartite graphs, to obtain bipartite graphs of multiple types. Different types of bipartite graphs can be obtained based on the type of historical corpus. For example, if the historical corpus is an introduction type, the clustered bipartite graph includes an introduction type bipartite graph.

[0067] In step 204, based on the clustered bipartite graph, the knowledge enhanced semantic representation model is trained to obtain a trained semantic classification model.

[0068] In this embodiment, the semantic classification model is a model obtained by improving a traditional knowledge-enhanced semantic representation model. The input of the semantic classification model is historical corpus, and the output of the semantic classification model is the type of the historical corpus. Compared with the traditional knowledge-enhanced semantic representation model, the semantic classification model is optimized in input during training, a bipartite graph with types is used, and the reliability of the output of the semantic classification model is improved.

[0069] The training process of the semantic classification model is as follows: obtaining a training set, the training set is a bipartite graph obtained according to user search questions filtered by a large search engine, manually labeling the category of each bipartite graph, clustering each bipartite graph, performing context prediction and masking word prediction using the clustered types, optimizing the knowledge-enhanced semantic representation model, until the knowledge-enhanced semantic representation model meets the training completion condition, and obtaining the trained semantic classification model. The training completion condition includes at least one of the following: the number of training iterations reaches a predetermined iteration threshold, and the loss value of the knowledge-enhanced semantic representation model is less than a predetermined loss value threshold. For example, the number of training iterations reaches 5,000. The loss value is less than 0.05.

[0070] In this embodiment, the Ernie knowledge-enhanced semantic representation model combines big data pre-training with multi-source rich knowledge, continuously absorbs knowledge of vocabulary, structure, and semantics in massive text data through continuous learning technology, and continuously evolves the model effect. Ernie has achieved good results in accumulating more than 40 typical NLP (Natural Language Processing) tasks.

[0071] The semantic classification model training method provided in this embodiment obtains the association relationship between sample corpus and sample title in a preset time period, organizes the bipartite graph of historical corpus and historical title based on the association relationship, trains the knowledge-enhanced semantic representation model after clustering the bipartite graph, and obtains the semantic classification model. Therefore, by clustering the bipartite graph, bipartite graphs of the same type can be divided together, and retraining the knowledge-enhanced semantic representation model using the clustered bipartite graph can improve the reliability of the semantic classification model in classifying historical corpus.

[0072] Figure 3 Flow 300 illustrating one embodiment of a method of generating a title corresponding to a multi-modal corpus according to the present disclosure is shown, and the method of generating a title corresponding to a multi-modal corpus includes the following steps:

[0073] In step 301, the relevance score and the heat score of each multi-modal corpus of the multi-modal corpus are added to obtain the score of each multi-modal corpus.

[0074] In the embodiment, each of the plurality of multi-modal corpora to be identified has a correlation score and a hotness score. For each multi-modal corpus, the correlation score and the hotness score of the multi-modal corpus are added to provide a score for the multi-modal corpus, which can reflect the reliability of the multi-modal corpus in generating a title.

[0075] At step 302, the multi-modal corpora are processed based on the scores of the multi-modal corpora to generate a title corresponding to the multi-modal corpora.

[0076] In the embodiment, each of the plurality of multi-modal corpora to be identified has a score, which reflects the reliability of the title generation. The multi-modal corpus with the highest score is selected, and the title corresponding to the plurality of multi-modal corpora to be identified can be obtained by performing entity extraction and splicing on the multi-modal corpus with the highest score, thereby providing an effective way for the user to generate a reliable title.

[0077] In some optional implementations of the embodiment, the processing of the multi-modal corpora based on the scores of the multi-modal corpora to generate a title corresponding to the multi-modal corpora includes: sorting the scores of the multi-modal corpora in descending order to obtain a score sequence of the multi-modal corpora; selecting the multi-modal corpora at a preset position in the score sequence as initial corpora; performing word segmentation on the initial corpora, extracting and splicing the segmented entities to obtain a title corresponding to the multi-modal corpora.

[0078] In the optional implementation, the preset position can be determined according to the title production demand and accuracy. When the title accuracy requirement is not high and the demand is not high, the preset position can be selected to be a small value, for example, the preset position is 10.

[0079] In the optional implementation, the scores of the multi-modal corpora are sorted in descending order, and the multi-modal corpora at the preset position are selected to provide a plurality of preferred corpora for generating a title, thereby ensuring the reliability of the title generation.

[0080] The method for generating a title corresponding to the multi-modal corpora provided in the embodiment adds the correlation score and the hotness score to obtain the score of the multi-modal corpora, and processes the multi-modal corpora based on the score of the multi-modal corpora to obtain a title corresponding to the multi-modal corpora, thereby providing a reliable implementation for the title generation and ensuring the reliability of the title generation.

[0081] In the example scenario of the embodiment, as shown in FIG. 1, a plurality of multi-modal corpora to be identified are obtained, and each of the plurality of multi-modal corpora to be identified has a correlation score and a hotness score. Figure 4As shown, from the initial historical corpus, select the historical corpus input pre-training completed semantic classification model to obtain different types of historical corpus, select part of the corpus from the initial historical corpus as the multi-modal corpus, which can be video corresponding corpus, image corresponding corpus, text corresponding corpus, etc. Through the title generator, the multi-modal corpus with high browsing volume and strong relevance to the title of the target site is recalled, and the recalled multi-modal corpus is used to guide the generation of a new title. The title generator can select part of the multi-modal corpus as the initial corpus from the recalled multi-modal corpus, perform word segmentation on the initial corpus, extract and splice or replace the segmented entities, and obtain the title T corresponding to the multi-modal corpus.

[0082] Further referring to Figure 5 , as an implementation of the method shown in the above figures, the disclosure provides an embodiment of a title generation device, which corresponds to the method embodiment shown in Figure 1 .

[0083] As shown in Figure 5 , the title generation device 500 provided in this embodiment includes a corpus acquisition unit 501, a type matching unit 502, a heat calculation unit 503, a correlation calculation unit 504, and a generation unit 505. The corpus acquisition unit 501 can be configured to acquire a plurality of multi-modal corpora to be identified, and the multi-modal corpus is a corpus corresponding to at least one type of modal information. The type matching unit 502 can be configured to match the multi-modal corpus with the historical corpus of different types obtained by pre-clustering, and determine the type of the multi-modal corpus. The heat calculation unit 503 can be configured to calculate a set of heat scores of the multi-modal corpus in the type, and each heat score in the set of heat scores is used to represent the similarity between each multi-modal corpus and the historical corpus. The correlation calculation unit 504 can be configured to calculate a set of correlation scores of the multi-modal corpus based on the title of the target site, and each correlation score in the set of correlation scores is used to represent the correlation between each multi-modal corpus and the title of the target site. The generation unit 505 can be configured to generate a title corresponding to the multi-modal corpus based on the set of correlation scores, the set of heat scores, and the multi-modal corpus.

[0084] In this embodiment, the specific processing of the corpus acquisition unit 501, the type matching unit 502, the heat calculation unit 503, the correlation calculation unit 504, and the generation unit 505 in the title generation device 500 and the technical effects brought by the specific processing can be respectively referred to the related description of step 101, step 102, step 103, step 104, and step 105 in the corresponding embodiment, which will not be repeated here. Figure 1 The corresponding embodiment in the corresponding embodiment, step 101, step 102, step 103, step 104, step 105, and step 105 are not described here.

[0085] In some optional implementations of the present embodiment, the type matching unit 502 includes a corpus matching module (not shown in the figure), and a type matching module (not shown in the figure). The corpus matching module can be configured to match the multi-modal corpus with the historical corpus. The type matching module can be configured to, in response to the multi-modal corpus matching any type of the historical corpus successfully, take the type of the historical corpus that matches successfully as the type of the multi-modal corpus.

[0086] In some optional implementations of the present embodiment, the heat calculation unit 503 is further configured to, in response to the historical corpus that matches the multi-modal corpus successfully being at least one, calculate the cosine distance value between the vector of each multi-modal corpus of the multi-modal corpus and the vector of each historical corpus in the type, and add the product of the cosine distance value and the normalized heat value of each historical corpus to obtain the heat score of the multi-modal corpus, and combine all the heat scores of the multi-modal corpus to obtain the heat score set of the multi-modal corpus in the type, wherein the normalized heat value is used to represent the popularity of the modal information corresponding to the historical corpus.

[0087] In some optional implementations of the present embodiment, the apparatus 500 further includes a semantic classification model, which is used to cluster the initial historical corpus to obtain the historical corpus, and the semantic classification model is obtained by training the following units: an association obtaining unit (not shown in the figure), an organization unit (not shown in the figure), a clustering unit (not shown in the figure), and an input unit (not shown in the figure). The association obtaining unit can be configured to obtain the association relationship between the sample corpus and the sample title in a preset time period. The organization unit can be configured to organize the sample corpus and the sample title corresponding to the sample corpus into a bipartite graph based on the association relationship. The clustering unit can be configured to cluster the bipartite graph to obtain a clustered bipartite graph. The input unit can be configured to train the knowledge-enhanced semantic representation model based on the clustered bipartite graph to obtain the trained semantic classification model.

[0088] In some optional implementations of the present embodiment, the correlation calculation unit 504 is further configured to calculate the cosine distance value between the vector of each multi-modal corpus of the multi-modal corpus and the vector of each title of the target site, and add the product of the cosine distance value and the normalized heat value of each title to obtain the correlation score of the multi-modal corpus, and combine all the correlation scores of the multi-modal corpus to obtain the correlation score set of the multi-modal corpus.

[0089] In some optional implementations of the present embodiment, the generation unit 505 includes an adding module (not shown in the figure), a generation module (not shown in the figure). The adding module can be configured to add the relevance score and the hotness score of each multi-modal corpus in the multi-modal corpus to obtain the score of each multi-modal corpus. The generation module can be configured to process the multi-modal corpus based on the score of each multi-modal corpus to generate the title corresponding to the multi-modal corpus.

[0090] In some optional implementations of the present embodiment, the generation module includes a sorting submodule (not shown in the figure), a selection submodule (not shown in the figure), and a segmentation submodule (not shown in the figure). The sorting submodule can be configured to sort the scores of the multi-modal corpora in the multi-modal corpus in descending order to obtain a score sequence of the multi-modal corpora. The selection submodule can be configured to select the multi-modal corpora at the first preset positions in the score sequence as initial corpora. The segmentation submodule can be configured to segment the initial corpora, extract and concatenate the segmented entities to obtain the title corresponding to the multi-modal corpus.

[0091] The title generation apparatus provided by the embodiments of the present disclosure first acquires the multi-modal corpus to be recognized by the corpus acquisition unit 501. Then, the type matching unit 502 matches the multi-modal corpus with the historical corpora of different types obtained by pre-clustering to determine the type of the multi-modal corpus. Next, the hotness calculation unit 503 calculates a set of hotness scores of the multi-modal corpus in the type, and each hotness score in the set of hotness scores is used to represent the similarity between each multi-modal corpus and the historical corpus. Then, the relevance calculation unit 504 calculates a set of relevance scores of the multi-modal corpus based on the title of the target site, and each relevance score in the set of relevance scores is used to represent the relevance between each multi-modal corpus and the title of the target site. Finally, the generation unit 505 generates the title corresponding to the multi-modal corpus based on the set of relevance scores, the set of hotness scores, and the multi-modal corpus. Thus, based on the hotness scores and the relevance scores of the multi-modal corpus, the title corresponding to the multi-modal corpus related to the title of the target site and the historical type is generated based on the popularity of the historical corpus and the popularity of the title, which realizes the personalized recommendation of the title and improves the user experience.

[0092] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0093] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0094] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0095] As shown, the device 600 includes a computing unit 601 that can perform various suitable actions and processes in accordance with computer programs stored in a read-only memory (ROM) 602 or loaded into a random access memory (RAM) 603 from a storage unit 608. Various programs and data required by the device 600 for operation can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other by a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604. Figure 6 Various components in the device 600 are connected to the I / O interface 605, including an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through computer networks, such as the Internet, and / or various telecommunication networks.

[0096]

[0097] ​The computing unit 601 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above, such as the title generation method. For example, in some embodiments, the title generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded onto the RAM 603 and executed by the computing unit 601, one or more steps of the title generation method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the title generation method by any other appropriate means, such as by means of firmware.

[0098] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0099] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0100] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0101] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0102] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0103] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0104] It should be understood that the various forms of flow shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure can be achieved, which is not limited herein.

[0105] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A title generation method, the method comprising: obtaining a plurality of multi-modal corpora to be identified, the multi-modal corpora being corpora corresponding to at least one type of modal information; matching the multi-modal corpora with historical corpora of different types obtained by pre-clustering to determine the types of the multi-modal corpora; calculating a set of heat scores of the multi-modal corpora, each heat score in the set of heat scores representing a similarity between a multi-modal corpus and a historical corpus, and calculating a cosine distance value between a vector of each multi-modal corpus and a vector of each historical corpus in the type, and adding a product of the cosine distance value and a normalized heat value of each historical corpus to obtain a heat score of the multi-modal corpus, the normalized heat value representing a popularity of the modal information corresponding to the historical corpus; calculating a set of relevance scores of the multi-modal corpora based on titles of a target site, each relevance score in the set of relevance scores representing a relevance between a multi-modal corpus and a title of the target site; generating a title corresponding to the multi-modal corpora based on the set of relevance scores, the set of heat scores, and the multi-modal corpora, selecting part of the multi-modal corpora as initial corpora based on scores of each multi-modal corpus in the multi-modal corpora, performing word segmentation on the initial corpora, extracting and concatenating segmented entities to obtain the title corresponding to the multi-modal corpora.

2. The method of claim 1, wherein, The matching the multi-modal corpora with historical corpora of different types obtained by pre-clustering to determine the types of the multi-modal corpora comprises: matching the multi-modal corpora with the historical corpora; in response to a successful matching of the multi-modal corpora with a historical corpus of any type, taking the type of the matched historical corpus as the type of the multi-modal corpus.

3. The method of claim 1 or 2, further comprising: The initial historical corpora are clustered using a pre-trained semantic classification model to obtain the historical corpora, and the semantic classification model is trained using the following steps: obtaining an association relationship between sample corpora and sample titles in a preset time period; organizing the sample corpora and the sample titles corresponding to the sample corpora into a bipartite graph based on the association relationship; clustering the bipartite graph to obtain a clustered bipartite graph; training a knowledge-enhanced semantic representation model based on the clustered bipartite graph to obtain the trained semantic classification model.

4. The method of claim 1, wherein, The calculating a set of relevance scores of the multi-modal corpora based on titles of a target site comprises: calculating a cosine distance value between a vector of each multi-modal corpus and a vector of each title of the target site, and adding a product of the cosine distance value and a normalized heat value of each title to obtain a relevance score of the multi-modal corpus, and combining all the relevance scores of the multi-modal corpora to obtain the set of relevance scores of the multi-modal corpora.

5. The method of claim 1, wherein, The generating a title corresponding to the multi-modal corpora based on the set of relevance scores, the set of heat scores, and the multi-modal corpora comprises: add the relevance score of each multi-modal corpus of the multi-modal corpora to the hotness score to obtain a score of each multi-modal corpus; processing the multi-modal corpora based on the score of each multi-modal corpus to generate a title corresponding to the multi-modal corpora.

6. The method of claim 5, wherein, The processing the multi-modal corpora based on the score of each multi-modal corpus to generate a title corresponding to the multi-modal corpora comprises: performing descending order sorting on the score of each multi-modal corpus in the multi-modal corpora to obtain a score sequence of the multi-modal corpora; selecting a multi-modal corpus at a predetermined position in the score sequence as an initial corpus; performing word segmentation on the initial corpus, extracting and splicing the segmented entities to obtain a title corresponding to the multi-modal corpus.

7. A title generation apparatus, the apparatus comprising: a corpus acquisition unit configured to acquire a plurality of multi-modal corpora to be recognized, the multi-modal corpora being corpora corresponding to at least one type of modal information; a type matching unit configured to match the multi-modal corpora with historical corpora of different types obtained by pre-clustering to determine the type of the multi-modal corpora; a hotness calculation unit configured to calculate a set of hotness scores of the multi-modal corpora in the type, each hotness score in the set of hotness scores being used to represent the similarity between each multi-modal corpus and the historical corpora, and to calculate the cosine distance value between the vector of each multi-modal corpus of the multi-modal corpora and the vector of each historical corpus in the historical corpora in the type, and to add the product of the cosine distance value and the normalized hotness value of each historical corpus to obtain the hotness score of the multi-modal corpus, the normalized hotness value being used to represent the popularity of the modal information corresponding to the historical corpus; a relevance calculation unit configured to calculate a set of relevance scores of the multi-modal corpora based on the title of a target site, each relevance score in the set of relevance scores being used to represent the relevance between each multi-modal corpus and the title of the target site; a generation unit configured to generate a title corresponding to the multi-modal corpora based on the set of relevance scores, the set of hotness scores, and the multi-modal corpora, and to select part of the multi-modal corpora as an initial corpus based on the score of each multi-modal corpus in the multi-modal corpora; performing word segmentation on the initial corpus, extracting and splicing the segmented entities to obtain a title corresponding to the multi-modal corpus.

8. The apparatus of claim 7, wherein, The type matching unit comprises: a corpus matching module configured to match the multi-modal corpora with the historical corpora; a type matching module configured to determine the type of the multi-modal corpora as the type of the historical corpus matched successfully in response to the successful matching of the multi-modal corpora with any type of historical corpus.

9. The apparatus of claim 7 or 8, further comprising: a semantic classification model used to cluster initial historical corpora to obtain the historical corpora, the semantic classification model being trained by the following units: an association acquisition unit configured to acquire the association relationship between sample corpora and sample titles in a preset time period; The organization unit is configured to organize the sample corpus and the sample title corresponding to the sample corpus into a bipartite graph based on the association relationship; The clustering unit is configured to cluster the bipartite graph to obtain a clustered bipartite graph; The input unit is configured to train the knowledge-enhanced semantic representation model based on the clustered bipartite graph to obtain a trained semantic classification model.

10. The apparatus of claim 7, wherein, The correlation calculation unit is further configured to calculate a cosine distance value between the vector of each multi-modal corpus of the multi-modal corpus and the vector of each title of the target site, and add the product of the cosine distance value and the normalized heat value of each title to obtain a correlation score of the multi-modal corpus, and combine all correlation scores of the multi-modal corpus to obtain a correlation score set of the multi-modal corpus.

11. The apparatus of claim 7, wherein, The generation unit includes: An addition module configured to add the correlation score of each multi-modal corpus of the multi-modal corpus and the heat score to obtain a score of each multi-modal corpus; A generation module configured to process the multi-modal corpus based on the score of each multi-modal corpus to generate a title corresponding to the multi-modal corpus.

12. The apparatus of claim 11, wherein, The generation module includes: A sorting submodule configured to sort the scores of each multi-modal corpus in the multi-modal corpus in descending order to obtain a score sequence of the multi-modal corpus; A selection submodule configured to select multi-modal corpora in the first set position in the score sequence as initial corpora; A word segmentation submodule configured to perform word segmentation on the initial corpora, extract and splice segmented entities to obtain a title corresponding to the multi-modal corpus.

13. An electronic device, comprising: At least one processor; And The memory is in communication connection with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1-6. The computer instructions are used to enable the computer to execute the method of any one of claims 1-6.

14. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, 15. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-6. ​

Citation Information

Patent Citations

  • News video abstract generation method and device

    CN113660541A