Information processing method and apparatus

By extracting key text information from the descriptive text of multimedia data and selecting a target cover template to synthesize a second cover information, the problem of weak correlation between the cover image and the data content is solved, thereby increasing user interest and click-through rate.

CN113821651BActive Publication Date: 2026-02-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110792430.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-13
Publication Date
2026-02-10
Estimated Expiration
2041-07-13

AI Technical Summary

Technical Problem

The cover image content of multimedia data has a weak connection with the data content, making it difficult to quickly attract user interest.

Method used

By acquiring descriptive text information associated with multimedia data, extracting key text information, and selecting a target cover template based on data attribute information, a second cover information is synthesized to improve relevance.

Benefits of technology

The content of the cover information has been enriched, the correlation between the cover information and multimedia data has been improved, and the user click-through rate and experience have been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113821651B_ABST
    Figure CN113821651B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an information processing method and device. The method comprises: extracting key text information for multimedia data from description text information; obtaining data attribute information of the multimedia data, and selecting a target cover template for the multimedia data from a cover template set according to the data attribute information; obtaining first cover information of the multimedia data, and performing information synthesis on the key text information and the first cover information according to the target cover template to obtain second cover information of the multimedia data. By using the embodiments of the present application, the content of the second cover information of the multimedia data can be enriched, and the relevance between the content of the second cover information and the content of the multimedia data can be improved. The method disclosed in the present application can also be applied to the field of natural language processing technology, for example, the key text information is extracted from the description text information by using the natural language processing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, and in particular to an information processing method and apparatus. Background Technology

[0002] Currently, with the rapid development of computers, the peak daily upload volume of various multimedia data (such as text and video data) from various sources has exceeded millions or even tens of millions. When displaying multimedia data, the cover image has a significant impact on click-through rates and user experience. Currently, cover images for multimedia data are typically uploaded by the creator or uploader of the multimedia data. However, in practice, the inventors have found that the content of cover images uploaded by creators is relatively simple, and the content of the cover image usually has a weak connection with the content of the multimedia data, making it difficult to quickly attract user interest. Therefore, how to enrich the content of multimedia data cover images and thus improve the relevance between the cover image content and the multimedia data content is an urgent problem to be solved. Summary of the Invention

[0003] This application provides an information processing method and apparatus that can enrich the content of the cover information of multimedia data, thereby improving the correlation between the content of the cover information and the content of the multimedia data.

[0004] On one hand, embodiments of this application provide an information processing method, which includes:

[0005] Obtain descriptive text information associated with multimedia data, and extract key text information about the multimedia data from the descriptive text information;

[0006] Obtain the data attribute information of the multimedia data, and select the target cover template for the multimedia data from the cover template set based on the data attribute information;

[0007] The first cover information of the multimedia data is obtained, and the key text information and the first cover information are synthesized according to the target cover template to obtain the second cover information of the multimedia data.

[0008] On one hand, embodiments of this application provide an image data processing apparatus, which includes:

[0009] The acquisition module is used to acquire descriptive text information associated with multimedia data;

[0010] The processing module is used to extract key text information for multimedia data from the descriptive text information;

[0011] The acquisition module also acquires data attribute information of multimedia data;

[0012] The processing module is also used to select a target cover template for multimedia data from the cover template set based on data attribute information;

[0013] The acquisition module is also used to acquire the first cover information of multimedia data;

[0014] The processing module is also used to synthesize key text information and first cover information based on the target cover template to obtain second cover information of multimedia data.

[0015] On one hand, embodiments of this application provide an electronic device, which includes a processor and a memory, wherein the memory is used to store computer program instructions, and the processor is configured to perform the following steps:

[0016] Obtain descriptive text information associated with multimedia data, and extract key text information about the multimedia data from the descriptive text information;

[0017] Obtain the data attribute information of the multimedia data, and select the target cover template for the multimedia data from the cover template set based on the data attribute information;

[0018] The first cover information of the multimedia data is obtained, and the key text information and the first cover information are synthesized according to the target cover template to obtain the second cover information of the multimedia data.

[0019] On one hand, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, perform the following steps:

[0020] Obtain descriptive text information associated with multimedia data, and extract key text information about the multimedia data from the descriptive text information;

[0021] Obtain the data attribute information of the multimedia data, and select the target cover template for the multimedia data from the cover template set based on the data attribute information;

[0022] The first cover information of the multimedia data is obtained, and the key text information and the first cover information are synthesized according to the target cover template to obtain the second cover information of the multimedia data.

[0023] On one hand, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative embodiments described above.

[0024] This application embodiment can extract key text information for multimedia data from the descriptive text information of multimedia data, and select a target cover template for multimedia data from the cover template set according to the data attribute information of multimedia data. Then, based on the target cover template, the key text information and the first cover information are synthesized to obtain the second cover information of multimedia data. This can enrich the content of the second cover information of multimedia data and improve the correlation between the content of the second cover information and the content of multimedia data. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a schematic diagram of the structure of an information processing system provided in an embodiment of this application;

[0027] Figure 2 This is a flowchart illustrating an information processing method provided in an embodiment of this application;

[0028] Figure 3 This is a schematic diagram illustrating the effect of cover information provided in an embodiment of this application;

[0029] Figure 4 This is a flowchart illustrating an information processing method proposed in an embodiment of this application;

[0030] Figure 5 This is a schematic diagram of the structure of a detection model provided in an embodiment of this application;

[0031] Figure 6 This is a schematic diagram illustrating the effect of an initial detection model provided in an embodiment of this application;

[0032] Figure 7 This is a schematic diagram illustrating the effect of a cover template classification method provided in an embodiment of this application;

[0033] Figure 8 This is a schematic diagram illustrating the effect of a second cover information provided in an embodiment of this application;

[0034] Figure 9 This is a flowchart illustrating an information processing method provided in an embodiment of this application;

[0035] Figure 10 This is a schematic diagram of the structure of an information processing system provided in an embodiment of this application;

[0036] Figure 11 This is a schematic diagram of the structure of an information processing device provided in an embodiment of this application;

[0037] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0038] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0039] This application proposes an information processing scheme that can extract key text information from the descriptive text information of multimedia data, select a target cover template according to the data attribute information of the multimedia data, and then perform information fusion between the key text information and the first cover information according to the target cover template to obtain the second cover information of the multimedia data. This can enrich the content of the second cover information of the multimedia data and improve the correlation between the content of the second cover information and the content of the multimedia data.

[0040] In practice, the collection and processing of relevant data in this application should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0041] In this embodiment, multimedia data can be various types of content creation organizations' user-generated content (UGC), professionally generated content (PGC), multi-channel network (MCN), etc. PGC (User Generated Content), UGC (Professionally Generated Content), and MCN (Multi-Channel Network) content are all relevant. PGC (User Generated Content) refers to user-uploaded original content displayed on internet platforms or provided to other users; UGC (Professionally Generated Content) refers to high-quality content produced by traditional broadcasters in a manner almost identical to television programs; MCN (Multi-Channel Network) content combines PGC content, ensuring continuous content output with strong capital support, thereby ultimately achieving stable commercial monetization. For example, this multimedia data can be news articles, user-generated long or short videos, etc., which will not be elaborated further here.

[0042] In one possible implementation, the multimedia data can be displayed in the form of feeds on the client (such as a second client). Feeds are a data format, also known as news sources, and are translated as sources, feeds, information providers, summaries, sources, news subscriptions, or web feeds (English: web feed, news feed, syndicated feed). Feeds disseminate the latest information to users, typically arranged in a timeline format, which is the most original, intuitive, and basic form of display for feeds. For example, if a user subscribes to a data source (such as a friend's account or the account of a public figure), when that data source publishes content, the user can receive updated multimedia data from that data source. If the user subscribes to enough data sources, they can receive continuously updated multimedia data. Understandably, when multimedia data is displayed in the form of feeds on a client (such as a second client), the multimedia data and its cover image (such as a target cover image obtained from the server) can be displayed. These can typically be displayed as single images, small images, large images, or three images, as well as multiple images displayed in a nine-grid or sixteen-grid layout within the feed's focus area. No restrictions are placed here.

[0043] In one possible implementation, the technical solution of this application can be applied to an image data processing system. Please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram of an information processing system provided in an embodiment of this application. The information processing system may include multiple application clients and a server. The application clients are used to upload multimedia data and first cover information. The application clients can also be used to receive multimedia data and second cover information, and then jointly output the multimedia data and second cover information. The server is used to extract key text information based on the descriptive text information of the multimedia data, and to select a target cover template from a set of cover templates. Then, based on the target cover template, it performs information fusion on the first cover information and key text information to obtain the second cover information. This ensures that the generated second cover information contains key text information, improving the correlation between the multimedia data and the second cover information, thereby increasing the user's click-through rate and experience with the multimedia data, ultimately increasing the consumption of information stream content, and thus improving the distribution efficiency of information stream content.

[0044] The technical solution of this application can be applied to servers, which can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0045] In one possible implementation, the process of acquiring the second cover information of multimedia data may generally include the following steps, please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating an information processing method provided in this application embodiment. It can be divided into five steps: Step 1, identifying the data attribute information of multimedia data; Step 2, extracting key text information from the descriptive text information; Step 3, selecting a target cover template; Step 4, obtaining first cover information; and Step 5, generating second cover information. Step 1 can involve manually or mechanically identifying the classification of multimedia data and categorizing its data attribute information into several tones. Step 2 involves extracting key text information from the descriptive text information using Natural Language Processing (NLP) methods. Step 3 involves determining the target cover template for the multimedia data based on its data attribute information and tones. Step 4 involves selecting the first cover information of the multimedia data, then fusing the first cover information and key text information, and improving the clarity and contrast of the cover information to enhance the display effect of the generated second cover information. For example, please refer to... Figure 3 , Figure 3 This is a schematic diagram illustrating the effect of cover information provided in an embodiment of this application. Figure 3 Figure (1) shows a schematic diagram of the effect of a first cover information. The content of this first cover information is relatively monotonous and contains no text information. Figure 3 Figure (2) shows a schematic diagram of the effect of a second cover information, which includes extracted key text information (such as...). Figure 3 As shown in Figure 3201, this greatly improves the correlation between cover information and multimedia data.

[0046] In one possible implementation, the method provided in this application can also be applied to the field of natural language processing (NLP). NLP is an important direction in computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, that is, the language people use in daily life, and thus it is closely related to linguistic research. NLP technologies typically include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies. For example, NLP technology can be used to extract key text information from descriptive text information.

[0047] It is understood that the above scenarios are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0048] Based on the above description, this application proposes an information processing method. Please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a flowchart illustrating an information processing method proposed in an embodiment of this application. The method can be executed by a server. The information processing method may include steps S401-S403.

[0049] S401. Obtain descriptive text information associated with multimedia data, and extract key text information related to multimedia data from the descriptive text information.

[0050] The multimedia data can be video data or image / text data. Video data consists of multiple video frames, while image / text data can be data containing both text and images, or a collection of images only; there are no restrictions here. The descriptive text associated with the multimedia data can be the title text, introduction text, subtitle text, etc., used to describe the multimedia data. The key text information is used to indicate keywords, key phrases, key sentences, main headings, etc., extracted from the descriptive text. For example, it can be obtained by fine-tuning a pre-trained BERT model to obtain a detection model, thereby extracting key text information.

[0051] In one possible implementation, extracting key text information for multimedia data from descriptive text information may further include the following steps: inputting the descriptive text information into a detection model, predicting the keyword probability and part-of-speech information corresponding to each character in the descriptive text information in the detection model; the part-of-speech information corresponding to each character is used to indicate the position of each character in its respective word; identifying characters in the descriptive text information whose corresponding keyword probability is greater than a probability threshold as text keywords; and obtaining key text information containing text keywords based on the part-of-speech information corresponding to each character.

[0052] This detection model is a key text information extraction model trained on sample text information carrying key text labels. Through this model, the probability of a keyword and the part-of-speech tag for each character in the descriptive text can be predicted. Optionally, the model structure can be a sequence labeling algorithm (CRF) layer added after the BERT model. After inputting the descriptive text into the detection model, the BERT model can obtain the character feature vector corresponding to each character in the descriptive text. Then, the character feature vector corresponding to each character is input into the CRF layer to obtain the part-of-speech tag and the probability of a keyword for each character. The BERT model can effectively extract the character feature vector corresponding to each character in the descriptive text. CRF is a sequence labeling algorithm that can accept an input sequence such as X = (x1, x2, ..., x...). n The output target sequence Y = (y1, y2, ..., y3) is calculated. n Let X represent the character feature vector corresponding to each character, and Y represent the part-of-speech information corresponding to each character. Optionally, an attention mechanism can be added to this detection model, which can enable the detection model to determine the association weights between each character and other characters in the word to which it belongs as the weights of the word, and then superimpose the word vector information onto the character vector information to obtain the final character feature vector. This can improve the correlation between characters and words and improve the accuracy of key text information extraction.

[0053] It is understood that the task of extracting key text information in this application embodiment is a supervised part-of-speech tagging task. Using the BERT model and CRF, the part-of-speech information of each character in a sentence or phrase can be quickly determined, thereby determining the probability that each character is a keyword, thus obtaining key text information. Compared to other unsupervised keyword extraction methods, the key information text extracted by this method has higher accuracy. For example, please refer to... Figure 5 , Figure 5This is a schematic diagram of the structure of a detection model provided in an embodiment of this application. The detection model includes a pre-trained BERT model and a CRF layer.

[0054] Keyword probability indicates the probability that a character is a keyword. This keyword probability is any value between 0 and 1, and each character corresponds to a unique keyword probability. The keyword probabilities for each character can be the same or different, depending on the actual detection results. Furthermore, characters in the descriptive text whose corresponding keyword probabilities are greater than a probability threshold can be identified as text keywords. This probability threshold indicates the minimum probability required to classify a character as a text keyword. For example, if the probability threshold is 0.8, and the detection model gives a character a keyword probability of 0.9, then that character is identified as a text keyword. If the detection model gives a character a keyword probability of 0.6, then that character is not identified as a text keyword. This process allows us to obtain all the text keywords in the descriptive text.

[0055] The词性 information corresponding to each character is respectively used to indicate the position of each character in the word it belongs to. For example, each character can be the first value, the middle value or the ending value of the word it belongs to, or it can indicate that the character itself is a word. Optionally, the词性 information of each character can be represented by the BIO ((B-begin, I-inside, O-outside) method. B indicates that the character is the starting position of the word it belongs to, I indicates that the character is the internal position of the word it belongs to, and O indicates that it does not belong to any type. For example, in a text "这件衣服", "衣服" belongs to the same word, "衣" is the starting position of the word, "服" is the internal position of the word, then the词性 information of "衣" is B, and the词性 information of "服" is I. "这" and "件" do not belong to any type, so their词性 information is both O. The词性 information of each character can be represented by the BIOES (B-begin, I-inside, O-outside, E-end, S-single) method. B indicates that the character is the starting position of the word it belongs to, I indicates that the character is the internal position of the word it belongs to, E indicates that the character is the ending position of the word it belongs to, O indicates that the character is a non-entity, and S indicates that the character itself is an entity. For example, in a text "这是俄罗斯", "俄罗斯" belongs to the same word, "俄" is the starting position of the word, "罗" is the internal position of the word, "斯" is the ending position of the word, then the词性 information of "俄" can be represented as B, the词性 information of "罗" can be represented as I, and the词性 information of "斯" can be represented as E. "这" and "是" both represent non-entities, so the词性 information of "这" and "是" can both be represented as O. It can be understood that according to the above representation method, the词性 information of characters can be divided into independent词性 information and non-independent词性 information. The independent词性 information is used to characterize that a character can form an independent word, and the non-independent词性 information is used to characterize that a character cannot form an independent word. The independent word is used to indicate a word that can express a certain meaning without being combined with adjacent characters in a sentence or phrase in the actual context. An independent word can include one character, or two characters or more, without limitation here. For example, in a short sentence "这是俄罗斯", "这" and "是" can both express a certain meaning without being combined with adjacent characters, so "这" and "是" are both independent words. "俄", "罗", and "斯" need to be combined among them to express a certain meaning, so "俄", "罗", and "斯" do not form independent words alone. "俄罗斯" can express a certain meaning only when connected, so the whole "俄罗斯" is an independent word.

[0056] Obtain the key text information containing the text keywords according to the词性 information corresponding to each character, which is used to determine the word to which each text keyword belongs according to the词性 information of the text keyword, and then obtain the key text information according to the word to which each text keyword belongs.

[0057] Optionally, the part-of-speech information of text keywords can be independent or non-independent. Independent part-of-speech information indicates that the text keyword constitutes an independent word; non-independent part-of-speech information indicates that the text keyword does not constitute an independent word. Therefore, obtaining key text information containing text keywords based on the part-of-speech information corresponding to each character can specifically include the following steps: If the part-of-speech information of the text keyword is independent, then the text keyword is used as key text information; if the part-of-speech information of the text keyword is non-independent, then characters associated with the text keyword are obtained from the descriptive text information based on the part-of-speech information corresponding to each character, and the characters associated with the text keyword and the text keyword are combined, and the resulting word is used as key text information. Specifically, if the part-of-speech information of the text keyword is independent, it means that the text keyword itself can constitute an independent word, and therefore the text keyword is used as key text information. For example, in the short sentence "I am Chinese," if the probability of the keyword corresponding to "I" is greater than the probability threshold, and the part-of-speech information of "I" indicates that the character is independent, then "I" is a text keyword, and "I" can be identified as key text information. If the part-of-speech information of a text keyword is not independent, it means that the text keyword itself does not constitute an independent word and needs to be combined with adjacent words to form an independent word. Then, based on the part-of-speech information corresponding to each word, words associated with the text keyword can be obtained from the descriptive text information, and words associated with the text keyword and text keyword can be combined to form a word as key text information. The words associated with text keywords are used to indicate the words in the phrase to which the text keyword belongs, as indicated by the part-of-speech information of each word determined by the detection model in the descriptive text information. For example, in the short sentence "I am Chinese", the detection model can obtain the part-of-speech information of each word. The part-of-speech information represented in the BIO format can be represented as "I / O is / O China / B country / I person / I". It can be clearly seen that according to the part-of-speech information of each word, "China", "country", and "person" belong to the same independent word. When the keyword probability of "China", "country", and "person" is greater than the probability threshold, that is, "China", "country", and "person" are determined as text keywords. "China", "country", and "person" are combined according to the part-of-speech information of each word, and the combined "Chinese person" is determined as the text key information. Optionally, if only the keyword probability of "China" and "country" is greater than the probability threshold, and the keyword probability of "person" is less than the probability threshold, then the words associated with "China" and "country" (i.e., "person") in the descriptive text can be obtained and combined, and the combined "Chinese person" can be determined as the text key information. It is understandable that in a descriptive text, there can be key text information obtained from text keywords that are independent parts of speech, or key text information obtained from words obtained by combining text keywords that are not independent parts of speech with related words.

[0058] For example, please refer to Table 1, which is an extraction illustration table for key text information extraction provided in an embodiment of this application. In Table 1, the first column is used to indicate the key text information extracted based on the title text (i.e., description text information) in the second column.

[0059] Table 1 Key Text Information: Title Text

[0060]

[0061] In one possible implementation, before obtaining key text information describing the text information through the detection model, the detection model can be trained, specifically including the following steps: obtaining sample text information; the sample text information carries key text tags; the key text tags are used to indicate sample key text information in the sample text information; inputting the sample text information into an initial detection model, generating keyword probabilities and part-of-speech information for each sample character in the sample text information; the part-of-speech information for each sample character is used to indicate the position of each sample character in its word; identifying sample characters whose corresponding keyword probabilities are greater than a probability threshold as sample text keywords; obtaining predicted key text information containing sample text keywords based on the part-of-speech information corresponding to each sample character; correcting the model parameters of the initial detection model based on the predicted key text information and the sample key text information indicated by the key text tags, and determining the initial detection model after model parameter correction as the detection model.

[0062] The sample text information is used to train the detection model. The key text tags carried by this sample text information indicate the key text information within it, specifically which words are keywords. Optionally, the sample text information can be obtained from a sample database, which may include multiple sample text information sets, each carrying a corresponding key text tag to indicate the key text information within each set. Further, optionally, the multiple sample text information sets in the sample database can be title text or description text of sample multimedia data (such as user-generated content), and the key text information within these title text or description text can be annotated. For example, please refer to Table 2. Table 2 is a schematic table of sample text information provided in the application embodiment. The first column of data represents the sample key text information marked according to the UGC short text (i.e., sample text information) shown in the second type of data. When marking the sample key text information, the text representing the UGC theme in the UGC short text can be determined as the sample key text information. For example, "shirt", "coffee table" and "The Peculiar Life" in Table 2 are all themes corresponding to the UGC short text. The text that can reflect the user's interests in the UGC short text can also be determined as the sample key text information. For example, "retro stereo" and "record player" in Table 2 belong to interest description.

[0063] Table 2

[0064]

[0065] In one possible implementation, multiple sample text information in the sample database can also be the title text of the viewed resource in the search results obtained from the data retrieval text search. Then, based on the data retrieval text, sample key text information of the title text of the viewed resource can be determined, because users have a strong purpose when searching, and the query phrases they think of are often the core phrases of the topic. This allows for the acquisition of sample text information from multiple perspectives, making the generalization ability of the detection model trained based on this sample text information stronger.

[0066] Specifically, this may include the following steps: obtaining data retrieval text; identifying the retrieval texts in the retrieval text list that have triggered browsing operations as initial sample text information; the retrieval text list containing at least one retrieval text retrieved based on the data retrieval text; determining sample key text information in the initial sample text information based on the data retrieval text; adding key text tags to the initial sample text information based on the sample key text information, and identifying the initial sample text information with added key text tags as sample text information. Here, the data retrieval text is used to indicate the text entered in the search bar of some clients or browsers for retrieval. The retrieval text list is used to indicate at least one retrieval text retrieved based on the data retrieval text. This retrieval text can be the title text, content summary, abstract, or short UGC text corresponding to various retrieved content; there are no restrictions here. The initial sample text information is used to indicate the search text that has triggered a browsing operation among at least one search text retrieved by the data retrieval text. For example, when a user retrieves at least one search text, such as "Search Text A," "Search Text B," and "Search Text C," and after obtaining the at least one search text, the user determines that the desired search result is "Search Text A" based on their actual needs, and clicks on the search text to browse, then it can be determined that "Search Text A" has triggered a browsing operation, and thus "Search Text A" can be identified as the initial sample text information. Furthermore, after obtaining the initial sample text information, the sample key text information can be determined based on the data retrieval text. For example, if the initial sample text information contains all the data retrieval text, then the content in the initial sample text information that is the same as the data retrieval text is identified as the sample key text information; or, if the initial sample text information does not completely contain the words in the data retrieval text, then the words in the initial sample text information that are repeated with the data retrieval text are identified as the sample key text information, and so on, without limitation. Furthermore, after obtaining the key text information of the initial sample text, key text tags can be added to the initial sample text based on the key text information. The initial sample text with added key text tags is then identified as the sample text information. This yields richer sample text information, enabling the detection model trained based on this sample text information to more accurately identify key text information in the text and exhibit stronger generalization ability.

[0067] Optionally, the sample text information can also carry sample attribute information for each sample character. This sample attribute information can be annotated using the BMES four-bit sequence labeling method, where B indicates that the sample character is the beginning position of the word it belongs to, M indicates that the sample character is the middle position of the word it belongs to, E indicates that the sample character is the end position of the word it belongs to, and S indicates that the sample character is a separate word. Thus, after the sample text information is input into the initial detection model, this sample attribute information can be integrated into the input of the BERT model, allowing the initial detection model to learn the relationship between characters and words more quickly.

[0068] The initial detection model's structure corresponds to the detection model's structure. It may include a BERT model and a CRF layer, or a BERT model, an attention mechanism, and a CRF layer; details are omitted here. For example, please refer to... Figure 6 , Figure 6 This is a schematic diagram illustrating the effect of an initial detection model provided in an embodiment of this application. Figure 6 Figure (1) shows the overall structure of the initial detection model, including the BERT model and the CRF layer. Figure 6 (2) can be a more specific initial detection model, including an initial detection model with an attention mechanism, and the sample text information can carry the corresponding sample attribute information annotated by the BMES four-bit sequence labeling method. After the sample text information is input into the BERT model, the embedding layer can incorporate the sample attribute information of each sample character carried in the sample text information. In the BERT model, the attention mechanism can also learn the relationship between each sample character in the sample text information and the word to which the sample character belongs, and thus obtain the character feature vector incorporating the relationship with the word. The part-of-speech information sample character probability corresponding to each sample character in the sample text information is obtained through the CRF layer output, and thus the predicted key text information corresponding to the sample text information can be obtained.

[0069] Inputting sample text information into the initial detection model yields the part-of-speech information and keyword probability for each sample character. The sample character indicates each character in the sample text. The description of the part-of-speech information and keyword probability for each sample character can be found in the description of the part-of-speech information and keyword probability for each character in the sample text, and will not be repeated here. Sample text keywords indicate sample characters in the sample text whose corresponding keyword probability is greater than a probability threshold. It can be understood that the probability threshold set during model training is the same as the probability threshold used by the obtained detection model during key text information extraction. Predicted key text information indicates key text information in the sample text obtained through the initial detection model. The process of obtaining predicted key text information can be found in the description of the key text information acquisition process in the description of the sample text, and will not be repeated here. Furthermore, the model parameters of the initial detection model can be corrected based on the difference between the predicted key text information and the sample key text information corresponding to the sample text information. The initial detection model after model parameter correction is then determined as the detection model.

[0070] Optionally, after obtaining predicted key text information from the sample text information and correcting the model parameters of the initial detection model, the keyword probability and part-of-speech information corresponding to each sample character in the sample text information can be obtained again based on the initial detection model with corrected model parameters. Based on the part-of-speech information corresponding to each sample character, predicted key text information containing the sample text keywords can be obtained. Then, the model parameters of the corrected initial detection model can be further corrected based on the predicted key text information and the sample key text information. This iterative training is performed in this manner until the initial detection model reaches a preset condition. This preset condition indicates the conditions that need to be met for training the initial detection model, such as calculating a loss function using the predicted key text information and the sample key text information. The value of this loss function is less than or equal to a loss threshold, which indicates the minimum loss value of the loss function that must be satisfied before training the initial detection model can be stopped. Further, optionally, multiple sample text information can be obtained from a sample database to train the initial detection model. The training process for each sample text information can refer to the description above. This makes the final detection model more capable of extracting key text information based on the initial detection model, and the accuracy is higher.

[0071] S402. Obtain the data attribute information of the multimedia data, and select the target cover template for the multimedia data from the cover template set based on the data attribute information.

[0072] The data attribute information can be used to indicate the classification / category of multimedia data. For example, the category of multimedia data can be food, current affairs, humor, society, entertainment, etc. Optionally, the attribute information of multimedia data can be the category selected by the user when uploading multimedia data, or it can be the category labeled by the server after receiving the multimedia data through manual or machine recognition. There is no restriction here. The cover template set can include N cover templates, where N is a positive integer. Each cover template can define the style, position, and other information of the key text information displayed in the corresponding generated cover information, such as the color and font of punctuation, numbers, or text in the key text information. The cover template can also define the filter information of the image in the corresponding generated cover information. The cover template can also include some graphics used to enhance the cover effect, which are not elaborated here. The target cover template is used to indicate the cover template of the multimedia data from the cover template set, and the target cover template is the cover template corresponding to the attribute information of the multimedia data.

[0073] In one possible implementation, the cover template set contains N cover templates, where N is a positive integer, and each cover template has a corresponding template tag; the data attribute information contains M data attribute tags for the multimedia data, where M is a positive integer. Selecting a target cover template for the multimedia data from the cover template set based on the data type can specifically include the following steps: The cover template whose template tag is present in at least one of the M data attribute tags is identified as the target cover template. Here, the template tag indicates the category of multimedia data to which the corresponding cover template is applicable. Each cover template can have multiple corresponding template tags; that is, different categories of multimedia data may select the same target cover template. The data attribute tag indicates the tag that indicates the category of the multimedia data. Therefore, by matching the attribute tags of the multimedia data with the template tags of the cover template, the second cover information of the multimedia data is selected, which means obtaining a cover template whose template tag is present in at least one of the M data attribute tags from the cover template set and identifying this cover template as the target cover template. For example, in the set of cover templates, cover template A has template tags of "real-time" and "society", and cover template B has template tags of "life" and "education". If the data attribute tag of the multimedia data is "life", then the template tags corresponding to cover template B contain at least one data attribute tag of the multimedia data. Therefore, cover template B is selected as the target cover template for the multimedia book data.

[0074] Optionally, the M data attribute labels of the multimedia data can be divided into different levels of labels, such as first-level labels and second-level labels. First-level labels indicate categories with a broad scope, second-level labels indicate more detailed classifications within the scope of the first-level label, and so on, including third-level and fourth-level labels. When determining the target cover template, the first-level label is primarily used as the selection criterion. For example, if the first-level label of the multimedia data is "film" and the second-level label is "movie," then the target cover template is selected based on the label "film." Further, optionally, the M data attribute labels of the multimedia data can also include multiple data attribute labels of the same level. However, these same-level data attribute labels have a priority relationship. This priority relationship is based on the correlation between each category and the multimedia data when the multimedia data is divided into multiple categories (i.e., having multiple data attribute labels of the same level). The higher the correlation between the multimedia data and the category, the higher the priority of the corresponding data attribute label; conversely, the lower the correlation, the lower the priority of the corresponding data attribute label. Therefore, when selecting the target cover template, the target cover template is selected based on the data attribute label with the highest priority. Optionally, when the number of embedding positions for key text information in the cover template is limited and insufficient to fully display the key text information extracted from the descriptive text information, key text information belonging to the whitelist of words corresponding to the cover template can be selected first. This whitelist of words indicates the list of words to be selected first when filtering the extracted key text information. For example, if the data attribute tag of the multimedia data is "funny", and the key text information "raccoon", "grapes", and "its reaction is enough to make me laugh for a year" are extracted from the descriptive text information "The raccoon is sitting on the sofa eating grapes, and after finding that there are no grapes in the bowl, its reaction is enough to make me laugh for a year", then the cover template corresponding to the template tag "funny" can be determined based on the data attribute tag "funny". Since the key text information that can be displayed in this cover template is limited, it is necessary to filter the above key text information. If the whitelist of words corresponding to this cover template contains the word "laugh", then the key text information "its reaction is enough to make me laugh for a year" can be displayed first. This can improve the relevance between the key text information and the multimedia data and improve the effect of the cover.

[0075] In one possible implementation, the data attribute tags of each multimedia data item can be categorized, such as into several tones: a slightly serious tone, a neutral and general tone, and a slightly relaxed tone. Correspondingly, the cover templates in the cover template set can also be categorized into slightly serious tone, neutral and general tone, and slightly relaxed tone, and the template tags of the cover templates can be categorized under the corresponding tone. This allows for determining the tone to which the data attribute tags of the multimedia data belong, and then selecting the target cover template corresponding to that data attribute tag from the cover templates under the corresponding tone based on the template tag. Please see [link to relevant documentation]. Figure 7 , Figure 7 This is a schematic diagram illustrating the effect of a cover template classification method provided in an embodiment of this application, such as... Figure 7 As shown, the template tags of the cover templates in the cover template collection can be categorized into three tones: template tags such as "society" and "history" can be categorized into serious tones; template tags such as "sports," "documentary," "workplace," and "anime" are neutral and general tones; and template tags such as "funny" can be categorized into lighthearted tones. Therefore, after obtaining the data attribute tags of the multimedia data, the target cover template can be selected from the cover templates under the corresponding tone in the cover template collection based on the tone corresponding to the data attribute tag. Optionally, the same data attribute tag can be categorized into two tones simultaneously. For example, the data attribute tag "society" can be categorized into a more serious tone or a neutral / general tone. Similarly, the template tag "society" can be categorized into either a more serious tone or a neutral / general tone. The same template tag can correspond to different tones, resulting in different cover templates. A cover template categorized as more serious might have a more formal style for the key text information, while another categorized as more serious might have a less formal style for the key text information. Therefore, after obtaining the data attribute tags of the multimedia data, the tone category of the data attribute tag is determined based on the tone of the multimedia data. Then, the target cover template corresponding to that data attribute tag is selected from the cover templates corresponding to that tone category in the cover template set. This allows for the determination of the target cover template corresponding to the template tag under the corresponding tone based on the tone of the multimedia data's data attribute tag, making the generated cover information more closely match the content of the multimedia data and improving accuracy.

[0076] S403. Obtain the first cover information of the multimedia data, and synthesize the key text information and the first cover information according to the target cover template to obtain the second cover information of the multimedia data.

[0077] The first cover information indicates the cover image information of the multimedia data that has not been synthesized with the key text information. This first cover information can be the cover image of the multimedia data uploaded by the user, or cover information obtained from the multimedia data (such as video frames in video data, image data in text and image data), or a cover image processed from the user-uploaded cover image and the cover image obtained from the multimedia data (such as a cover image after cropping, rotating, or adding filters). There are no restrictions here. Synthesizing the key text information and the first cover information according to the target cover template involves importing the first cover information into the target cover template at its designated display position, importing the key text information into the target cover template at its designated display position, determining the style of the key text information according to the style defined in the target cover template, and then adjusting the target cover template after importing the key text information and the first cover information using filters and other information defined in the cover template. Finally, a second cover information is obtained, which indicates the cover information of the multimedia data that has added key text information based on the first cover information.

[0078] In one possible implementation, obtaining the first cover information of multimedia data may further include the following steps: obtaining an initial cover image of the multimedia data; and performing image enhancement processing on the initial cover image to obtain the first cover information. The initial cover image can be a cover image of multimedia data uploaded by the user, or it can be cover information obtained from the multimedia data (such as video frames in video data, image data in text and image data), without limitation. The methods for performing image enhancement processing on the initial cover image can include various methods, such as sharpening the edges of objects in the initial cover image, enhancing the colors in the initial cover image, adding filters to the initial cover image, adjusting the contrast and clarity of the initial cover image, etc., without limitation. This can improve the effect of the obtained first cover information, thereby enhancing the effect of the corresponding second cover information.

[0079] In one possible implementation, after obtaining the second cover information of the multimedia data, if a data push instruction for the multimedia data is received, the second cover information of the multimedia data is pushed to the application client, so that the application client can associate and output the multimedia data and the second cover information. The data push instruction is used to instruct the multimedia data and the second cover information to be pushed to the application client. This data push instruction can be generated based on a data acquisition request sent by the application client, or it can be an instruction generated by the server based on a multimedia user recommendation mechanism to push multimedia data to the application client; no limitation is made here. The application client can be any application client that receives the multimedia data and the second cover information. The application client can associate and output / display the multimedia data and the second cover information; for example, please refer to [reference needed]. Figure 8 , Figure 8 This is a schematic diagram illustrating the effect of a second cover information provided in an embodiment of this application, such as... Figure 8 As shown in (1), the multimedia data is a video data, and the video's cover image is the second cover information generated based on the first cover information and key text information of the multimedia data. Figure 8 As shown in Figure 801, the key text information extracted from the descriptive text information allows for the rapid identification of the main content of the news item as "animal warfare" from the cover image of the multimedia data. This enables users of the application client to quickly determine that the multimedia data describes a battle between animals. Figure 8 As shown in (2), the multimedia data is a video data, and the video's cover image is the second cover information generated based on the first cover information and key text information of the multimedia data. Figure 8 As shown in Figure 802, the key text information extracted from the descriptive text information allows for the rapid identification of a video describing a "fishing crash course" from the cover image of the multimedia data. This quickly attracts the interest of users on the application client, thereby increasing the click-through rate of the multimedia data. Figure 8 As shown in (3), the multimedia data is a video data, and the cover image of the video is the second cover information generated based on the first cover information and key text information of the multimedia data. Figure 8 As shown in Figure 803, the key text information extracted from the descriptive text information can be quickly determined from the cover image of the multimedia data to describe "small animals eating". This can quickly arouse the interest of the corresponding users of the application client, increase the attractiveness of the multimedia data to users, and thus increase the click-through rate of the multimedia data.

[0080] This application embodiment can extract key text information specific to multimedia data from the descriptive text information of multimedia data, and select a target cover template for the multimedia data from a set of cover templates based on the data attribute information of the multimedia data. Then, it synthesizes the key text information and the first cover information based on the target cover template to obtain the second cover information of the multimedia data. This can enrich the content of the second cover information of the multimedia data, thereby improving the correlation between the content of the second cover information and the content of the multimedia data.

[0081] Please see Figure 9 , Figure 9 This is a flowchart illustrating an information processing method provided in an embodiment of this application. The method can be executed by a server. The information processing method may include steps S901-S907.

[0082] S901. Obtain descriptive text information associated with multimedia data, and extract key text information related to multimedia data from the descriptive text information.

[0083] The description of step S901 can be found in the description of step S401, and will not be repeated here.

[0084] S902. Obtain the data attribute information of the multimedia data, and obtain the user attribute information corresponding to the K user groups of the multimedia data respectively; K is a positive integer.

[0085] The description of the data attributes of the multimedia data can be found in step S402, and will not be repeated here. The user group refers to the group of users who can receive the multimedia data. Each user group has similar characteristics, which can be described by the user attribute information corresponding to that user group. This user attribute information can be corresponding user tags. Optionally, the user attribute information can be user tags determined by acquiring a large amount of user behavior or user information. For example, after acquiring a user's user behavior, it is found that the user frequently browses humorous content and fashion and beauty content, indicating that the user likes light and lively content. Therefore, the user can be tagged as someone who likes light and lively content. When generating the second cover information for this type of user, some lively and lighthearted cover templates can be selected to generate the corresponding second cover information. Similarly, after acquiring a user's user behavior, it is found that the user browses current events and social information, indicating that the user likes serious and formal content. Therefore, the user can be tagged as someone who likes serious and formal content. When generating the second cover information for this type of user, some serious and formal cover templates can be selected to generate the corresponding second cover information. This allows for a more targeted selection of cover templates for users.

[0086] S903. Based on the data attribute information and the user attribute information corresponding to each user group, select the target cover template for each user group from the cover template set.

[0087] In this set of cover templates, each cover template can also include an applicable group tag describing the user group to which the cover template is suitable. Therefore, there can be multiple cover templates for the same template tag, and these multiple cover templates for the same template tag can have different applicable group tags. For example, for the neutral and general data attribute tag "sports," the cover templates could include one suitable for users who prefer a lively and relaxed style, with more lively and relaxed keyword text styles, added graphics, and image filters; or it could include one suitable for users who prefer a more serious and formal style, with more serious and formal keyword text styles, added graphics, and image filters. Thus, the same template tag can have both types of cover templates to meet the needs of different user groups. Furthermore, after determining the data attribute information of the multimedia data, a cover template containing at least one data attribute tag from that data attribute information can be determined. And based on the user attribute information corresponding to each user group, the cover template containing that at least one data attribute tag corresponding to each user attribute information can be determined as the target cover template. For example, if the data attribute tag of multimedia data is "film and television", the user attribute information of the user group for this multimedia data can be user attribute information indicating whether they like lively and relaxed or serious and formal. In the cover template set, the cover templates with the template tag "film and television" can include cover templates suitable for both lively and relaxed and serious and formal users. Therefore, these two cover templates can be determined as the target cover templates for users who like lively and relaxed and the target cover templates for users who like serious and formal users, respectively.

[0088] S904. Based on the target cover template corresponding to each user group, synthesize the key text information and the first cover information to obtain the second cover information of multimedia data for each user group.

[0089] In step S904, the description of information synthesis of key text information and first cover information can be found in step S403, and will not be repeated here. It is understood that by synthesizing information for the target cover template of each user group, the second cover information corresponding to each user group can be obtained, thus yielding second cover information for different user groups of multimedia data.

[0090] Optionally, after obtaining the second cover information for each user group, the second cover information for each user group can be stored in the storage area so that the second cover information of multimedia data can be quickly retrieved from the storage area according to the user attribute information, thereby improving the data response speed.

[0091] S905. When a data push command for an application client is detected, obtain the user attribute information of the user to whom the application client belongs.

[0092] The data push instruction specifies the user attribute information of the user to whom the data is pushed to the application client. This user attribute information indicates the user attribute information of the user logged in to the application client.

[0093] S906. Determine the user groups corresponding to the user attribute information of the users belonging to the application client in the K user groups as the target user groups.

[0094] The target user group is used to indicate the user group corresponding to the application client to which the data push instruction is sent, thereby determining the second cover information corresponding to the target user group. For example, if the user attribute information of the client indicated by the data push instruction is that the user attribute information is that they like lively and relaxed activities, then the user group corresponding to the user attribute information of the client can be determined as the target user group. It can be understood that by generating corresponding second cover information for each user group, the second cover information can be made more targeted to each user group, thereby bridging the gap between users, between users and content, and between content, providing better results for multimedia data recommendation and search.

[0095] S907. Push the multimedia data and the second cover information corresponding to the target user group to the application client so that the application client can output the multimedia data and the second cover information corresponding to the target user group in association.

[0096] Specifically, the multimedia data and the second cover information corresponding to the target user group are pushed to the aforementioned application client. The description of how the application client associates and outputs the multimedia data and the second cover information corresponding to the target user group can be found in step S403, and will not be repeated here. This makes the second cover information of the multimedia data displayed on the application client more suitable for the users of that application client, thereby attracting user attention and increasing the click-through rate of the multimedia data.

[0097] In a practical application scenario, the embodiments of this application can be applied to an information processing system. Please refer to... Figure 10 , Figure 10This is a schematic diagram illustrating the effect of an information processing system provided in an embodiment of this application. The system mainly includes an intelligent enhancement service, a detection model, a template database, a cover image acquisition service, and a scheduling center service. The intelligent enhancement service receives scheduling from the scheduling center service, completes information enhancement of the first cover information of the multimedia data (i.e., information fusion of the first cover information and key text information), and finally generates the enhanced cover information (i.e., the second cover information), which may specifically include, for example... Figure 2 The multiple steps shown are not elaborated here. The detection model is used to receive the scheduling of the intelligent enhancement service. The supervised training method constructed above is used to build a detection model that can fully learn the semantic information of sentences. This detection model is used to extract key text information from the descriptive text information of multimedia data. The template database includes the above-mentioned set of cover templates. Each cover template in the template database can be labeled with a corresponding template to facilitate the determination of a suitable target cover template, that is, to communicate with the intelligent enhancement service. This allows the extracted key text information and the key phrases of the first cover information to be fused according to the style and strategy indicated by the cover template. The cover image acquisition service is a pre-service before the cover image enhancement process. It can include the selection and cropping of cover images, etc. The selection of cover images mainly involves filtering and screening cover images according to basic quality characteristics such as clarity, aesthetics, unsuitable images, mosaics, etc., removing some low-quality images that are not suitable for the cover. Then, based on the selection of a suitable cover image, the selected cover image (that is, the initial cover image) is subjected to image enhancement processing to obtain the first cover information. This dispatch center service is used to coordinate various services, such as calling the intelligent enhancement service to enhance the cover image of multimedia data, that is, to fuse the first cover information with key text information, etc. There are no restrictions here. In other words, the dispatch center service can call the intelligent enhancement service to generate second cover information for multimedia data (such as...). Figure 10 As shown in Figure 1001), the intelligent enhancement service, after receiving the scheduling from the scheduling center service, can call the detection model to extract key text information (such as...) from the descriptive text information of the multimedia data. Figure 10 As shown in Figure 1002), and can select a target cover template for multimedia data from the template database (such as...). Figure 10 As shown in Figure 1003), and the cover image acquisition service is called to determine the first cover information of the multimedia data (such as...). Figure 10 As shown in Figure 1004, the intelligent enhancement service then integrates the key text information and the first cover information obtained based on the target cover template to obtain the second cover information after the multimedia data has been enhanced. The second cover information may include key text information.

[0098] Optionally, the image data processing system may also include modules such as a content database, upstream and downstream content interface servers, a manual review system, a file download service, a deduplication service, a content distribution export service, a content storage service, a first client (also known as a content producer), a second client (also known as a content consumer), etc., without limitation. The content database can store metadata about multimedia data published by producers, such as file size, cover image link, bitrate, file format, title, publication time, author, video file size, video format, whether it is original or a first release, and also includes the classification of multimedia data during the manual review process, etc., which will not be elaborated here. This manual review system serves as the carrier of human service capabilities. Typically a web-based system, it uses machine-filtered results as input for manual confirmation and verification. The verification results are recorded in the content information metadata database. Simultaneously, the manual verification results can be used to evaluate the actual effectiveness of machine tagging and filtering models online. It can also report detailed review processes to a statistics server, including the source of the manual review task, review results, and start and end times. Furthermore, it is used to review and filter content that machines cannot determine, including tagging and secondary verification of short videos and micro-videos. It also serves as a reviewer for reports and negative feedback from content consumers, and so on. The upstream and downstream content interface server manages the upstream and downstream flow of multimedia data. For example, it communicates directly with the content producer (i.e., the first client), storing content submitted by the first client into the content database. This typically includes information such as the content title, publisher, summary, original cover image, publication time, and file size. It can also synchronize content submitted by multimedia data publishers (including content provided by external channels) to the dispatch center server for subsequent multimedia data processing and flow. This file download service is used to download data from content storage services, controlling the download speed and progress. This file download service typically consists of a group of parallel servers, comprised of related task scheduling and distribution clusters, such as downloading and retrieving multimedia data from content storage services. The deduplication service can be used to deduplicate multimedia data titles, cover images, content text, and video and audio fingerprints. For example, it vectorizes the titles and text of image and text content, uses simmhash (a hash algorithm) and BERT to determine the text vector and image vector, and then deduplicates them. For video content, it extracts video and audio fingerprints to construct vectors, and then calculates the distance between vectors, such as Euclidean distance, to determine if there is a duplicate. The purpose of deduplication is mainly to reduce the amount of content review and ensure that the same multimedia data exists only once in the recommendation distribution pool, thereby protecting the user experience.This content distribution exit service is used to indicate the exit point for multimedia data output in machine and manual processing links. Through this content distribution exit, multimedia data is distributed to the content consumer end (i.e., the second client). The distribution method can be recommendation algorithm distribution or manual operation; there are no restrictions here. This content storage service is used to store multimedia-related data. It typically consists of a widely distributed set of storage servers, located far from the C-side user, usually with peripheral CDN acceleration servers for distributed caching. Uplink and downlink content interface servers store video and image content uploaded by content producers. After obtaining content index information, the second client can also directly access the content storage server to download the corresponding content (i.e., the content storage server stores the data source for external services), and can also provide a file download service with raw multimedia data for related processing (i.e., the content storage server stores the data source for internal services). Optionally, this information processing system may also include other service modules and implement corresponding functions, which will not be elaborated here.

[0099] This application embodiment can extract key text information specific to multimedia data from the descriptive text information of multimedia data, and select a target cover template corresponding to each user group from the cover template set based on the data attribute information of the multimedia data and the user attribute information of each user group. Then, based on the target cover template, the key text information and the first cover information are synthesized to obtain the second cover information of the multimedia data for each user group. This can enrich the content of the second cover information of the multimedia data, thereby improving the relevance between the content of the second cover information and the content of the multimedia data.

[0100] Based on the description of the above information processing method embodiments, this application also discloses an information processing apparatus, which can be configured in the aforementioned server. For example, the apparatus can be a computer program (including program code) running on the server. The apparatus can execute... Figure 4 The method shown. Please refer to [link / reference]. Figure 11 The device can operate the following modules:

[0101] The acquisition module 1101 is used to acquire descriptive text information associated with multimedia data;

[0102] Processing module 1102 is used to extract key text information for multimedia data from descriptive text information;

[0103] The acquisition module 1101 also acquires the data attribute information of the multimedia data;

[0104] The processing module 1102 is also used to select a target cover template for multimedia data from the cover template set based on the data attribute information;

[0105] The acquisition module 1101 is also used to acquire the first cover information of the multimedia data;

[0106] The processing module 1102 is also used to synthesize key text information and first cover information according to the target cover template to obtain second cover information of multimedia data.

[0107] In one embodiment, the processing module 1102 is specifically used for:

[0108] The descriptive text information is input into the detection model, which predicts the keyword probability and part-of-speech information for each character in the descriptive text information. The part-of-speech information for each character is used to indicate the position of each character in the word it belongs to.

[0109] Characters whose probability of corresponding keywords in the descriptive text information is greater than the probability threshold are identified as text keywords;

[0110] Extract key text information containing text keywords based on the part-of-speech information corresponding to each character.

[0111] In one implementation, the part-of-speech information of the text keywords is either independent part-of-speech information or non-independent part-of-speech information; independent part-of-speech information indicates that the text keywords constitute independent words; non-independent part-of-speech information indicates that the text keywords do not constitute independent words; the processing module 1102 is specifically used for:

[0112] If the part-of-speech information of the text keywords is independent part-of-speech information, then the text keywords will be treated as key text information;

[0113] If the part-of-speech information of the text keyword is not independent part-of-speech information, then the words associated with the text keyword are obtained from the descriptive text information according to the part-of-speech information corresponding to each character, and the words associated with the text keyword and the text keyword are combined, and the combined words are used as key text information.

[0114] In one embodiment, the acquisition module 1101 is further configured to:

[0115] Obtain sample text information; the sample text information carries key text tags; the key text tags are used to indicate the key text information in the sample text information.

[0116] Processing module 1102 is also used for:

[0117] The sample text information is input into the initial detection model, which generates the keyword probability and part-of-speech information for each sample character in the sample text. The part-of-speech information for each sample character is used to indicate the position of each sample character in its word.

[0118] The sample words whose corresponding keyword probability is greater than the probability threshold are identified as sample text keywords.

[0119] Based on the part-of-speech information corresponding to each sample character, obtain the predicted key text information containing the keywords of the sample text;

[0120] The model parameters of the initial detection model are corrected based on the sample key text information indicated by the predicted key text information and key text labels, and the initial detection model after model parameter correction is determined as the detection model.

[0121] In one implementation, the processing module 1102 specifically performs the following:

[0122] Retrieve data retrieval text;

[0123] The search texts in the search text list that have been triggered to perform browsing operations are identified as the initial sample text information; the search text list contains at least one search text retrieved based on the data search text.

[0124] Based on the data retrieval text, determine the key text information of the initial sample text information;

[0125] Key text tags are added to the initial sample text information based on the key text information of the sample, and the initial sample text information with added key text tags is determined as the sample text information.

[0126] In one implementation, the processing module 1102 specifically performs the following:

[0127] Obtain the initial cover image for multimedia data;

[0128] Image enhancement processing is performed on the initial cover image to obtain the first cover information.

[0129] In one implementation, the cover template set contains N cover templates, where N is a positive integer, and each cover template has a corresponding template tag; the data attribute information contains M data attribute tags for the multimedia data, where M is a positive integer; the processing module 1102 is specifically used for:

[0130] The cover template whose template tag is present in at least one of the M data attribute tags among the N cover templates is identified as the target cover template.

[0131] In one embodiment, the processing module 1102 is specifically used for:

[0132] Obtain user attribute information for K user groups in the multimedia data; K is a positive integer.

[0133] Based on the data attribute information and the user attribute information corresponding to each user group, select the target cover template for each user group from the cover template set;

[0134] Based on the target cover template corresponding to each user group, the key text information and the first cover information are synthesized to obtain the second cover information for each user group in the multimedia data.

[0135] In one embodiment, the processing module 1102 is further configured to:

[0136] The user groups corresponding to the user attribute information of the users belonging to the application client in the K user groups are identified as the target user groups;

[0137] The multimedia data and the second cover information corresponding to the target user group are pushed to the application client so that the application client can associate and output the multimedia data and the second cover information corresponding to the target user group.

[0138] In the various embodiments of this application, the functional modules can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules, and this application does not impose any limitations on this.

[0139] This application embodiment can extract key text information specific to multimedia data from the descriptive text information of multimedia data, and select a target cover template for the multimedia data from a set of cover templates based on the data attribute information of the multimedia data. Then, it synthesizes the key text information and the first cover information based on the target cover template to obtain the second cover information of the multimedia data. This can enrich the content of the second cover information of the multimedia data, thereby improving the correlation between the content of the second cover information and the content of the multimedia data.

[0140] Please see again Figure 12 , Figure 12 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. The electronic device includes a processor 1201 and a memory 1202. Optionally, the electronic device may also include a network interface 1203 or a power supply module, etc. The processor 1201, memory 1202, and network interface 1203 can exchange data. The network interface 1203 is controlled by the processor to send and receive messages. The memory 1202 stores computer programs, including program instructions. The processor 1201 executes the program instructions stored in the memory 1202. The processor 1201 is configured to invoke the program instructions to execute the aforementioned method.

[0141] The memory 1202 may include volatile memory, such as random-access memory (RAM); the memory 1202 may also include non-volatile memory, such as flash memory, solid-state drive (SSD), etc.; the memory 1202 may also include a combination of the above types of memory.

[0142] The processor 1201 may be a central processing unit (CPU). In one embodiment, the processor 1201 may also be a graphics processing unit (GPU). The processor 1201 may also be a combination of a CPU and a GPU.

[0143] In one embodiment, memory 1202 is used to store program instructions. Processor 1201 can invoke the program instructions to perform the following steps:

[0144] Obtain descriptive text information associated with multimedia data, and extract key text information about the multimedia data from the descriptive text information;

[0145] Obtain the data attribute information of the multimedia data, and select the target cover template for the multimedia data from the cover template set based on the data attribute information;

[0146] The first cover information of the multimedia data is obtained, and the key text information and the first cover information are synthesized according to the target cover template to obtain the second cover information of the multimedia data.

[0147] In one implementation, processor 1201 specifically executes:

[0148] The descriptive text information is input into the detection model, which predicts the keyword probability and part-of-speech information for each character in the descriptive text information. The part-of-speech information for each character is used to indicate the position of each character in the word it belongs to.

[0149] Characters whose probability of corresponding keywords in the descriptive text information is greater than the probability threshold are identified as text keywords;

[0150] Extract key text information containing text keywords based on the part-of-speech information corresponding to each character.

[0151] In one implementation, the part-of-speech information of the text keywords is either independent part-of-speech information or non-independent part-of-speech information; independent part-of-speech information indicates that the text keywords constitute independent words; non-independent part-of-speech information indicates that the text keywords do not constitute independent words; processor 1201 specifically executes:

[0152] If the part-of-speech information of the text keywords is independent part-of-speech information, then the text keywords will be treated as key text information;

[0153] If the part-of-speech information of the text keyword is not independent part-of-speech information, then the words associated with the text keyword are obtained from the descriptive text information according to the part-of-speech information corresponding to each character, and the words associated with the text keyword and the text keyword are combined, and the combined words are used as key text information.

[0154] In one implementation, processor 1201 also performs:

[0155] Obtain sample text information; the sample text information carries key text tags; the key text tags are used to indicate the key text information in the sample text information.

[0156] The sample text information is input into the initial detection model, which generates the keyword probability and part-of-speech information for each sample character in the sample text. The part-of-speech information for each sample character is used to indicate the position of each sample character in its word.

[0157] The sample words whose corresponding keyword probability is greater than the probability threshold are identified as sample text keywords.

[0158] Based on the part-of-speech information corresponding to each sample character, obtain the predicted key text information containing the keywords of the sample text;

[0159] The model parameters of the initial detection model are corrected based on the sample key text information indicated by the predicted key text information and key text labels, and the initial detection model after model parameter correction is determined as the detection model.

[0160] In one implementation, processor 1201 specifically performs:

[0161] Retrieve data retrieval text;

[0162] The search texts in the search text list that have been triggered to perform browsing operations are identified as the initial sample text information; the search text list contains at least one search text retrieved based on the data search text.

[0163] Based on the data retrieval text, determine the key text information of the initial sample text information;

[0164] Key text tags are added to the initial sample text information based on the key text information of the sample, and the initial sample text information with added key text tags is determined as the sample text information.

[0165] In one implementation, processor 1201 specifically performs:

[0166] Obtain the initial cover image for multimedia data;

[0167] Image enhancement processing is performed on the initial cover image to obtain the first cover information.

[0168] In one implementation, the cover template set contains N cover templates, where N is a positive integer, and each cover template has a corresponding template tag; the data attribute information contains M data attribute tags for the multimedia data, where M is a positive integer; the processor 1201 specifically executes:

[0169] The cover template whose template tag is present in at least one of the M data attribute tags among the N cover templates is identified as the target cover template.

[0170] In one implementation, processor 1201 specifically performs:

[0171] Obtain user attribute information for K user groups in the multimedia data; K is a positive integer.

[0172] Based on the data attribute information and the user attribute information corresponding to each user group, select the target cover template for each user group from the cover template set;

[0173] Based on the target cover template, key text information and first cover information are synthesized to obtain second cover information for multimedia data, including:

[0174] Based on the target cover template corresponding to each user group, the key text information and the first cover information are synthesized to obtain the second cover information for each user group in the multimedia data.

[0175] In one implementation, processor 1201 also performs:

[0176] When a data push command is detected for the application client, the user attribute information of the user to which the application client belongs is obtained;

[0177] The user groups corresponding to the user attribute information of the users belonging to the application client in the K user groups are identified as the target user groups;

[0178] The multimedia data and the second cover information corresponding to the target user group are pushed to the application client so that the application client can associate and output the multimedia data and the second cover information corresponding to the target user group.

[0179] In specific implementations, the device, processor 1201, memory 1202, etc., described in the embodiments of this application can execute the implementation methods described in the above method embodiments, or they can execute the implementation methods described in the embodiments of this application, which will not be repeated here.

[0180] This application also provides a computer (readable) storage medium storing a computer program. The computer program includes program instructions, which, when executed by a processor, can perform some or all of the steps executed in the above method embodiments. Optionally, the computer storage medium can be volatile or non-volatile.

[0181] This application also provides a computer program product or computer program, which includes program instructions that can be stored in a computer-readable storage medium. A processor of a computer device reads the program instructions from the computer-readable storage medium and executes the program instructions, causing the computer to perform some or all of the steps described in the above method, which will not be elaborated further here.

[0182] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer storage medium, which can be a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the methods described above. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0183] The above-disclosed embodiments are merely some of the embodiments of this application, and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments, and equivalent changes made in accordance with the claims of this application, still fall within the scope of this application.

Claims

1. An information processing method, characterized in that, The method includes: Obtain descriptive text information associated with multimedia data, and extract key text information related to the multimedia data from the descriptive text information; Obtain the data attribute information of the multimedia data, and the user attribute information corresponding to K user groups of the multimedia data; K is a positive integer; the data attribute information includes data attribute tags of the multimedia data, and the data attribute tags are used to indicate the category of the multimedia data; Based on the tone corresponding to the multimedia data, determine the tone category to which the data attribute tag of the multimedia data belongs; the same data attribute tag can be classified into two tones simultaneously; Based on the data attribute information, the tone category of the data attribute tags, and the user attribute information corresponding to each user group, target cover templates for each user group are selected from the cover template set. Each cover template has at least one corresponding template tag, which indicates the category of multimedia data to which the corresponding cover template is applicable. Each cover template has a corresponding tone category; the same template tag corresponds to different tones, resulting in different cover templates. The user attribute information indicates the tone category of the user's preferred content. Tone categories include serious tone, neutral and general tone, and relaxed tone. The tone category of the target cover template corresponds to the tone category of the data attribute tags and the tone category of the corresponding user group. The first cover information of the multimedia data is obtained, and the key text information and the first cover information are synthesized according to the target cover template corresponding to each user group to obtain the second cover information of the multimedia data for each user group.

2. The method according to claim 1, characterized in that, The step of extracting key text information related to the multimedia data from the descriptive text information includes: The descriptive text information is input into a detection model, which predicts the keyword probability corresponding to each character in the descriptive text information and the part-of-speech information corresponding to each character; the part-of-speech information corresponding to each character is used to indicate the position of each character in the word to which it belongs. Characters in the descriptive text information whose probability of corresponding keywords is greater than a probability threshold are identified as text keywords; The key text information containing the text keywords is obtained based on the part-of-speech information corresponding to each character.

3. The method according to claim 2, characterized in that, The part-of-speech information of the text keywords can be independent part-of-speech information or non-independent part-of-speech information; the independent part-of-speech information indicates that the text keywords constitute independent words; the non-independent part-of-speech information indicates that the text keywords do not constitute independent words. The step of obtaining the key text information containing the text keywords based on the part-of-speech information corresponding to each character includes: If the part-of-speech information of the text keyword is the independent part-of-speech information, then the text keyword is used as the key text information; If the part-of-speech information of the text keyword is the non-independent part-of-speech information, then according to the part-of-speech information corresponding to each character, the characters associated with the text keyword are obtained from the descriptive text information, and the characters associated with the text keyword and the text keyword are combined, and the combined words are used as the key text information.

4. The method according to claim 2, characterized in that, The method further includes: Obtain sample text information; the sample text information carries key text tags; the key text tags are used to indicate key sample text information in the sample text information; The sample text information is input into an initial detection model, which generates the keyword probability corresponding to each sample character in the sample text information and the part-of-speech information corresponding to each sample character. The part-of-speech information corresponding to each sample character is used to indicate the position of each sample character in its respective word. The sample characters whose keyword probability in the sample text information is greater than the probability threshold are identified as sample text keywords; Based on the part-of-speech information corresponding to each sample character, predictive key text information containing the keywords of the sample text is obtained; The model parameters of the initial detection model are corrected based on the predicted key text information and the sample key text information indicated by the key text labels, and the initial detection model after model parameter correction is determined as the detection model.

5. The method according to claim 4, characterized in that, The acquisition of sample text information includes: Retrieve data retrieval text; The search texts in the search text list that have been triggered by browsing operations are identified as initial sample text information; the search text list contains at least one search text retrieved based on the data search text; Based on the data retrieval text, determine the key text information of the sample in the initial sample text information; Based on the key text information of the sample, the key text tag is added to the initial sample text information, and the initial sample text information with the key text tag is determined as the sample text information.

6. The method according to claim 1, characterized in that, The acquisition of the first cover information of the multimedia data includes: Obtain the initial cover image of the multimedia data; The initial cover image is subjected to image enhancement processing to obtain the first cover information.

7. The method according to claim 1, characterized in that, The cover template set contains N cover templates, where N is a positive integer, and each cover template has a corresponding template tag; the data attribute information contains M data attribute tags of the multimedia data, where M is a positive integer. The step of selecting the multimedia data from the cover template set for each user group includes: Based on the template tags corresponding to the N cover templates, the cover templates that have at least one of the M data attribute tags, and the user attribute information corresponding to each user group, the target cover templates for each user group of the multimedia data are determined.

8. The method according to claim 1, characterized in that, The method further includes: When a data push command is detected for an application client, the user attribute information of the user to which the application client belongs is obtained; The user groups corresponding to the user attribute information of the users to which the application client belongs in the K user groups are determined as the target user groups; The multimedia data and the second cover information corresponding to the target user group are pushed to the application client, so that the application client can associate and output the multimedia data and the second cover information corresponding to the target user group.

9. An information processing device, characterized in that, include: The acquisition module is used to acquire descriptive text information associated with multimedia data; The processing module is used to extract key text information related to the multimedia data from the descriptive text information; The acquisition module is further configured to acquire data attribute information of the multimedia data, and user attribute information corresponding to K user groups of the multimedia data; K is a positive integer; the data attribute information includes data attribute tags of the multimedia data, and the data attribute tags are used to indicate the category of the multimedia data; The processing module is further configured to determine the tone category to which the data attribute tags of the multimedia data belong based on the tone corresponding to the multimedia data; and select target cover templates for each user group from the cover template set based on the data attribute information, the tone category to which the data attribute tags belong, and the user attribute information corresponding to each user group. Each cover template has at least one corresponding template tag. The template tag is used to indicate the category of multimedia data applicable to the corresponding cover template. Each cover template has a corresponding tone category. The same template tag will result in different cover templates for different tones. The user attribute information is used to indicate the tone category of the user's preferred content; the tone categories include serious tone, neutral and general tone, and relaxed tone; The tone category of the target cover template corresponds to the tone category of the data attribute tag and the tone category of the corresponding user group; the same data attribute tag can be classified into two tones at the same time; The acquisition module is also used to acquire the first cover information of the multimedia data; The processing module is further configured to synthesize the key text information and the first cover information according to the target cover template corresponding to each user group, so as to obtain the second cover information of the multimedia data for each user group.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory for storing computer program instructions, and the processor is configured to perform the method as described in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, are used to perform the method as described in any one of claims 1-8.

12. A computer program product, characterized in that, The computer program product includes computer instructions that are executed to perform the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Video recommendation method and device, storage medium, terminal and server

    CN111246255A

  • Guidance language generation method and device, electronic device and computer storage medium

    CN112100357A

  • Information flow processing method and device and electronic equipment

    CN112100501A

  • Content cover generator and method based on web front end

    CN112257406A