Methods for handling response content and interactive methods for media content.
By performing style recognition and distribution analysis on the interactive content awaiting a response in media content, the system automatically generates response content, solving the problem of low efficiency in traditional responses and achieving a highly efficient response process.
Patent Information
- Application Number
- CN202210904173.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-07-28
AI Technical Summary
In traditional methods, users browsing media content need to actively input responses, resulting in low response efficiency.
By acquiring the interactive content to be responded to from media content, style identification is performed based on the description content, interactive content and its publishing object information, style distribution is generated, style category vector is determined, and corresponding response content is automatically generated.
It enables automatic generation of reply content without user input, thus improving reply efficiency.
Smart Images

Figure CN115186085B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for processing response content, as well as an interactive method, apparatus, computer device, storage medium, and computer program product for interactive media content. Background Technology
[0002] With the development of computer technology, more and more interactive content is being published in media, such as comments and bullet comments. Viewers of the media content can reply to published interactive content to engage in discussion and interaction.
[0003] In traditional technology, the way to reply to published interactive content is for the person browsing the media content to enter a reply to the published interactive content.
[0004] However, traditional methods suffer from low response efficiency because the recipients of the media content need to actively input their responses. Summary of the Invention
[0005] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for processing response content that can improve response efficiency in response to the above-mentioned technical problems.
[0006] Firstly, this application provides a method for processing response content. The method includes:
[0007] Get the pending interactive content related to media content;
[0008] Based on the description of the media content, the interactive content to be replied to, and the information of the recipients of the interactive content to be replied to, style identification is performed on the interactive content to be replied to, and the style distribution of the interactive content to be replied to is obtained.
[0009] Determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to;
[0010] Based on the description content, the interactive content to be replied to, and the style category vector, generate reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
[0011] Secondly, this application also provides a response content processing apparatus. The apparatus includes:
[0012] The interactive content acquisition module is used to acquire interactive content that needs to be responded to in relation to media content;
[0013] The style recognition module is used to identify the style of the interactive content to be replied to based on the description of the media content, the interactive content to be replied to, and the publishing object information of the interactive content to be replied to, so as to obtain the style distribution of the interactive content to be replied to.
[0014] The style category vector determination module is used to determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to;
[0015] The response content generation module is used to generate response content corresponding to at least one style category in the style distribution of the interaction content to be responded to, based on the description content, the interaction content to be responded to, and the style category vector.
[0016] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0017] Get the pending interactive content related to media content;
[0018] Based on the description of the media content, the interactive content to be replied to, and the information of the recipients of the interactive content to be replied to, style identification is performed on the interactive content to be replied to, and the style distribution of the interactive content to be replied to is obtained.
[0019] Determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to;
[0020] Based on the description content, the interactive content to be replied to, and the style category vector, generate reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
[0021] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0022] Get the pending interactive content related to media content;
[0023] Based on the description of the media content, the interactive content to be replied to, and the information of the recipients of the interactive content to be replied to, style identification is performed on the interactive content to be replied to, and the style distribution of the interactive content to be replied to is obtained.
[0024] Determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to;
[0025] Based on the description content, the interactive content to be replied to, and the style category vector, generate reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
[0026] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0027] Get the pending interactive content related to media content;
[0028] Based on the description of the media content, the interactive content to be replied to, and the information of the recipients of the interactive content to be replied to, style identification is performed on the interactive content to be replied to, and the style distribution of the interactive content to be replied to is obtained.
[0029] Determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to;
[0030] Based on the description content, the interactive content to be replied to, and the style category vector, generate reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
[0031] Sixthly, this application provides an interactive method for media content. The method includes:
[0032] Interactive content that displays media content;
[0033] In response to any interactive event, determine the interactive content to be replied to;
[0034] Corresponding to the interactive content to be replied to, display at least one style category of alternative reply content corresponding to the interactive content to be replied to;
[0035] In response to the selection event of any candidate reply content, the candidate reply content selected by the selection operation is displayed as the interactive content to be replied to.
[0036] Seventhly, this application also provides an interactive device for media content interaction. The device includes:
[0037] The interactive content display module is used to display interactive content from media content.
[0038] The response module is used to respond to trigger events for any interactive content and determine the interactive content to be replied to;
[0039] The alternative response content display module is used to display alternative response content corresponding to at least one style category of the interactive content to be responded to.
[0040] The reply content display module is used to respond to the selection event of any candidate reply content and display the selected candidate reply content as the interactive content to be replied to.
[0041] Eighthly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0042] Interactive content that displays media content;
[0043] In response to any interactive event, determine the interactive content to be replied to;
[0044] Corresponding to the interactive content to be replied to, display at least one style category of alternative reply content corresponding to the interactive content to be replied to;
[0045] In response to the selection event of any candidate reply content, the candidate reply content selected by the selection operation is displayed as the interactive content to be replied to.
[0046] Ninthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0047] Interactive content that displays media content;
[0048] In response to any interactive event, determine the interactive content to be replied to;
[0049] Corresponding to the interactive content to be replied to, display at least one style category of alternative reply content corresponding to the interactive content to be replied to;
[0050] In response to the selection event of any candidate reply content, the candidate reply content selected by the selection operation is displayed as the interactive content to be replied to.
[0051] Tenthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0052] Interactive content that displays media content;
[0053] In response to any interactive event, determine the interactive content to be replied to;
[0054] Corresponding to the interactive content to be replied to, display at least one style category of alternative reply content corresponding to the interactive content to be replied to;
[0055] In response to the selection event of any candidate reply content, the candidate reply content selected by the selection operation is displayed as the interactive content to be replied to.
[0056] The aforementioned response content processing method, apparatus, computer equipment, storage medium, and computer program product acquire interactive content to be responded to for media content. Based on the description of the media content, the interactive content to be responded to, and the publishing object information of the interactive content to be responded to, style recognition is performed on the interactive content to be responded to, thereby obtaining the style distribution of the interactive content to be responded to. This allows the determination of the style category vector corresponding to at least one style category in the style distribution of the interactive content to be responded to. Based on the description, the interactive content to be responded to, and the style category vector, response content corresponding to at least one style category in the style distribution of the interactive content to be responded to is generated. By automatically generating response content corresponding to at least one style category as the interactive content selection, the need for inputting response content is eliminated, thus improving response efficiency.
[0057] The aforementioned interactive methods, apparatuses, computer devices, storage media, and computer program products for interactive media content, by displaying interactive media content, responding to a trigger event for any interactive content, determining the interactive content to be replied to, and displaying alternative reply content corresponding to at least one style category of the interactive content to be replied to, can provide alternative reply content corresponding to at least one style category as interactive content selection. Therefore, in response to a selection event for any alternative reply content, the alternative reply content selected by the selection operation can be displayed as the interactive content to reply to the interactive content to be replied to, eliminating the need for inputting reply content and improving reply efficiency. Attached Figure Description
[0058] Figure 1 This is an application environment diagram of the response content processing method in one embodiment;
[0059] Figure 2 This is a flowchart illustrating a response content processing method in one embodiment;
[0060] Figure 3 This is a schematic diagram of an interactive content style recognition model in one embodiment;
[0061] Figure 4 This is a schematic diagram of an interactive content style recognition model in another embodiment;
[0062] Figure 5 This is a schematic diagram of a response content generation model in one embodiment;
[0063] Figure 6 This is a schematic diagram of an interaction rate prediction model in one embodiment;
[0064] Figure 7 This is a schematic diagram illustrating the process of generating dynamic style video comment replies in one embodiment;
[0065] Figure 8 This is a schematic diagram of a video comment style recognition model in one embodiment;
[0066] Figure 9 This is a schematic diagram of a dynamic style comment reply generation model in one embodiment;
[0067] Figure 10 This is a schematic diagram of a different style of comment reply priority model in one embodiment;
[0068] Figure 11 This is a flowchart illustrating a method for interacting with media content in one embodiment.
[0069] Figure 12 This is a schematic diagram illustrating interactive content displaying media content in one embodiment;
[0070] Figure 13 This is a schematic diagram illustrating interactive content displaying media content in another embodiment;
[0071] Figure 14 This is a schematic diagram showing alternative response content corresponding to at least one style category in one embodiment;
[0072] Figure 15 This is a schematic diagram showing alternative response content corresponding to at least one style category in another embodiment;
[0073] Figure 16 This is a schematic diagram showing alternative response content corresponding to at least one style category in yet another embodiment;
[0074] Figure 17 This is a structural block diagram of a response content processing device in one embodiment;
[0075] Figure 18 This is a structural block diagram of an interactive device for media content interaction in one embodiment;
[0076] Figure 19 This is an internal structural diagram of a computer device in one embodiment;
[0077] Figure 20 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation
[0078] The solutions provided in this application relate to the field of artificial intelligence (AI) technology. AI is the theory, methods, technology, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine capable of reacting in a manner similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0079] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0080] This application primarily concerns Natural Language Processing (NLP) technology. NLP is an important field within computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it is closely related to linguistic research. NLP technologies typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0081] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0082] The response content processing method provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on the cloud or other servers. Server 104 obtains the interactive content to be replied to from terminal 102. Based on the description of the media content, the interactive content to be replied to, and the publishing object information of the interactive content to be replied to, it performs style recognition on the interactive content to be replied to, obtains the style distribution of the interactive content to be replied to, determines the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to, and generates reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to based on the description, the interactive content to be replied to, and the style category vector. Terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, portable wearable devices, and aircraft. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented using a standalone server, a server cluster consisting of multiple servers, or a cloud server, or it can be a node on a blockchain.
[0083] In one embodiment, such as Figure 2 As shown, a method for processing reply content is provided. This method can be executed by the terminal or the server alone, or by the terminal and the server working together. In this embodiment, the method is described using the server as an example, and includes the following steps:
[0084] Step 202: Obtain the interactive content to be responded to for the media content.
[0085] Media content refers to media data published on public platforms or applications. For example, media content can specifically refer to video data published on public platforms or applications. Similarly, media content can specifically refer to audio data published on public platforms or applications. A public platform is a publicly accessible and browsable platform. For example, a public platform can specifically refer to a video website. Similarly, a public platform can specifically refer to an audio website. An application is a program that can be used to publish media content. For example, an application can specifically refer to an audio application used to publish media content. Similarly, an application can specifically refer to a video application used to publish media content.
[0086] Here, "interactive content awaiting a response" refers to interactive content that requires a reply. Interactive content refers to content posted by those viewing media content that expresses their views on the media content. For example, interactive content could specifically refer to comments posted by those viewing media content. Another example is "bullet comments" posted by those viewing media content.
[0087] Specifically, when processing response content, the server retrieves the interactive content to be responded to for the media content. In a specific application, the server responds to the selection of any interactive content on the media content and retrieves the interactive content to be responded to for that media content. In a specific application, when the terminal displays interactive content on the media content, in response to a trigger event for any interactive content, after determining the interactive content to be responded to, it sends the interactive content to be responded to to the server, so that the server can retrieve the interactive content to be responded to for the media content.
[0088] Step 204: Based on the description of the media content, the interactive content to be replied to, and the information of the recipient of the interactive content to be replied to, perform style identification on the interactive content to be replied to and obtain the style distribution of the interactive content to be replied to.
[0089] The descriptive content refers to the content used to describe the media content. For example, the descriptive content can specifically refer to content extracted from the media content. For instance, descriptive content could be subtitles extracted from the media content. Or, it could be content extracted from interactive content based on the media content. For example, descriptive content could be keywords extracted from interactive content based on the media content; interactive content could include bullet comments, comments, etc. The publishing target information refers to the information of the object that published the comment to be replied to, and can be used to determine the style distribution of the objects that published the comment to be replied to. For instance, publishing target information could specifically refer to the historical interactive content published by the object that published the comment to be replied to.
[0090] In this context, "style" refers to a comprehensive set of characteristics exhibited in literary creation. In this embodiment, it primarily refers to language style. Style recognition involves analyzing the style category of the interactive content to be responded to and identifying that category. Style category refers to the type of language style, which can be configured according to the actual application scenario. For example, specific style categories could include philosophical, humorous, satirical, or poetic styles.
[0091] Style distribution describes the distribution of style categories. It includes at least the set of style categories corresponding to the style identification object. For example, when the style distribution includes the set of style categories corresponding to the style identification object, the style distribution could take the form of [humorous style, philosophical style, poetic style]. In specific applications, the style distribution can also include the probability of each style category within the style category set. For example, when the style distribution includes the style category set and the probability of each style category within the style category set, the style distribution could take the form of [humorous style 0.55, philosophical style 0.3, poetic style 0.15].
[0092] Specifically, the server performs joint analysis based on the media content description, the interactive content to be replied to, and the recipient information of the interactive content to be replied to, in order to identify the style of the interactive content to be replied to and obtain its style distribution. In specific applications, the server can perform style identification based on the media content description, the interactive content to be replied to, and the recipient information of the interactive content to be replied to separately, and obtain the style distribution of the interactive content to be replied to based on the style distribution obtained from each style identification. The server can also perform style identification jointly on the media content description and the interactive content to be replied to, and perform style identification based on the recipient information of the interactive content to be replied to, and obtain the style distribution of the interactive content to be replied to based on the style distribution obtained from each style identification.
[0093] In a specific application, when obtaining the style distribution of the interactive content to be replied to based on the style distribution obtained from each style recognition, the server can perform a weighted average of the style category probabilities of each style category in the style distribution obtained from each style recognition to obtain the style distribution of the interactive content to be replied to. The weighting coefficients can be configured according to the actual application scenario when performing the weighted average.
[0094] In practical applications, before performing style recognition, the server needs to obtain the descriptive content of the media content. In one application, the server can extract key content from the media content using methods such as optical character recognition (OCR) and automatic speech recognition, and use this extracted key content as the descriptive content. In another application, the server can extract keywords from the interactive content of the media content, obtain the corresponding keywords, and use these keywords as the descriptive content. In yet another application, the server can also use both the key content extracted from the media content and the corresponding keywords from the interactive content as the descriptive content.
[0095] When the media content is video, the server can obtain the corresponding video frames and acquire video subtitles through optical character recognition (OCR), using these subtitles as key content extracted from the media content. Additionally, the server can obtain the corresponding dialogue text through automatic speech recognition (ACR), using this dialogue text as key content extracted from the media content. In a specific application, the server can also simultaneously use both video subtitles and dialogue text as key content extracted from the media content.
[0096] In this embodiment, the method of keyword extraction is not limited. For example, keyword extraction can be performed based on the TF-IDF (Term Frequency / Inverse Document Frequency) algorithm. Alternatively, keyword extraction can be performed based on the Textrank algorithm.
[0097] Step 206: Determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to.
[0098] Among them, the style category vector is a vector used to describe the style category, and different style categories correspond to different style category vectors.
[0099] Specifically, after obtaining the style distribution of the interactive content to be replied to, the server determines the style category vector corresponding to at least one style category in the style distribution. In specific applications, the server can determine the style category vector corresponding to at least one style category in the style distribution by querying a preset style category vector table, which pre-configures the correspondence between style categories and style category vectors. Alternatively, the server can directly encode at least one style category in the style distribution of the interactive content to be replied to, thereby obtaining the style category vector corresponding to at least one style category.
[0100] In a specific application, the server can use a pre-trained language representation model to encode at least one style category from the style distribution of the interactive content to be replied to, thereby obtaining a style category vector corresponding to at least one style category. The pre-trained language representation model can be selected and pre-trained according to the actual application scenario. For example, the pre-trained language representation model can specifically be a BERT model (Bidirectional Encoder Representation from Transformers).
[0101] Step 208: Based on the description content, the interactive content to be replied to, and the style category vector, generate reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
[0102] Specifically, the server segments and vectorizes the description and the interactive content to be replied to, obtaining vectorized representations of each word in the description and the interactive content to be replied to. Then, based on the vectorized representations of each word and the style category vector, it generates reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to. In specific applications, the server first predicts reply words based on the vectorized representations of each word and the style category vector, obtaining at least one predicted reply word corresponding to at least one style category in the style distribution of the interactive content to be replied to. Then, based on at least one predicted reply word corresponding to at least one style category, it generates reply content corresponding to at least one style category.
[0103] In a specific application, when predicting response words based on the vectorized representations and style category vectors of each word, the server uses a copy mechanism to choose between generating the predicted response word from a pre-set vocabulary or directly copying a word from the description content and the content to be replied to. The copy mechanism selects between generating and directly copying based on maximizing probability; that is, the server calculates the probability of each word being a predicted response word and the probability of each candidate word in the pre-set vocabulary being a predicted response word, and then chooses between generating or directly copying based on these two probabilities.
[0104] Furthermore, the copying mechanism involves a simple constraint rule: if a word does not appear in the input (i.e., the probability of each word being a predicted response word is less than the first threshold), then it will definitely not be copied. If a word only appears in the input but is not in the vocabulary (i.e., the probability of each candidate word in the preset vocabulary being a predicted response word is less than the second threshold), then it will definitely be copied. In this embodiment, the input is the vectorized representation of each word. The first and second thresholds can be configured according to the actual application scenario; they can be the same or different. For example, both the first and second thresholds can be 0.
[0105] In practical applications, the style distribution of the interactive content to be replied to includes a set of style categories and the corresponding style category probability for each style category in the style category set. When generating reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to, the server will filter out the style categories for which reply content needs to be generated based on the corresponding style category probabilities. That is, for style categories whose style category probabilities meet the category probability threshold, the corresponding reply content will be generated. The category probability threshold can be configured according to the actual application scenario. In this way, computational resources can be saved when generating reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
[0106] The above-described response processing method acquires the interactive content to be responded to for media content. Based on the description of the media content, the interactive content to be responded to, and the information of the recipient of the interactive content, it performs style recognition on the interactive content to be responded to, thereby obtaining the style distribution of the interactive content to be responded to. This allows it to determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be responded to. Based on the description, the interactive content to be responded to, and the style category vector, it generates response content corresponding to at least one style category in the style distribution of the interactive content to be responded to. By automatically generating response content corresponding to at least one style category as the interactive content selection, it eliminates the need to input response content, thus improving response efficiency.
[0107] In one embodiment, based on the description of the media content, the interactive content to be replied to, and the recipient information of the interactive content to be replied to, style identification is performed on the interactive content to be replied to to obtain the style distribution of the interactive content to be replied to, including:
[0108] Style identification is performed based on the descriptive content of the media content and the interactive content to be replied to, resulting in the first style distribution;
[0109] Style identification is performed based on the information of the publishing object of the interactive content to be replied to, resulting in a second style distribution;
[0110] Based on the first style distribution and the second style distribution, the style distribution of the interactive content to be responded to is obtained.
[0111] The first style distribution describes the distribution of style categories of the interactive content to be replied to, based on the description content. The second style distribution describes the distribution of style categories of the recipients of the interactive content to be replied to. The style distribution of the interactive content to be replied to describes the distribution of style categories of the interactive content to be replied to, taking into account the description content, the interactive content to be replied to, and the recipient information simultaneously.
[0112] Specifically, the server performs style identification based on the description of the media content and the interactive content to be replied to, obtaining a first style distribution. It then performs style identification based on the publishing object information of the interactive content to be replied to, obtaining a second style distribution that describes the distribution of style categories of the publishing object of the interactive content to be replied to. By comprehensively considering the first and second style distributions, the style distribution of the interactive content to be replied to is obtained.
[0113] In practical applications, the first style distribution includes a set of style categories, and the second style distribution also includes a set of style categories. The server uses each style category in the first and second style distributions as the corresponding style category of the interactive content to be replied to, thus obtaining the style distribution of the interactive content to be replied to. At this time, the style distribution of the interactive content to be replied to includes each style category in the first and second style distributions. For example, if the first style distribution is [philosophical, humorous, parody], and the second style distribution is [humorous, parody, poetic], then the style distribution of the interactive content to be replied to will be [philosophical, humorous, parody, poetic].
[0114] In practical applications, the first style distribution includes a set of style categories and the corresponding style category probability for each style category in the style category set. The second style distribution includes a set of style categories and the corresponding style category probability for each style category in the style category set. The server performs a weighted average of the style category probabilities included in the first and second style distributions to obtain the style distribution of the interactive content to be replied to. The style distribution of the interactive content to be replied to includes a set of style categories and the corresponding style category probability for each style category in the style category set. For example, if the first style distribution is [philosophical 0.6, humorous 0.2, parody 0.2] and the second style distribution is [humorous 0.4, parody 0.4, poetic 0.2], then the style distribution of the interactive content to be replied to is [philosophical 0.3, humorous 0.3, parody 0.3, poetic 0.1].
[0115] In this embodiment, by simultaneously considering the style distribution of the descriptive content of the media content, the interactive content to be replied to, and the publishing object information of the interactive content to be replied to, the style distribution of the interactive content to be replied to can be determined, which can improve the accuracy of style recognition and achieve accurate identification of the style distribution of the interactive content to be replied to.
[0116] In one embodiment, style identification is performed based on the descriptive content of the media content and the interactive content to be responded to, resulting in a first style distribution, including:
[0117] The descriptive content of the media content and the interactive content to be responded to are encoded to obtain the corresponding target vector representations of the descriptive content and the interactive content to be responded to;
[0118] The first style distribution is obtained based on the target vector representation.
[0119] In this context, the target vector representation refers to the vector used to characterize the descriptive content and the interactive content to be responded to. For example, the target vector representation can specifically be the word vectors of each word in the descriptive content and the interactive content to be responded to.
[0120] Specifically, the server encodes the description of the media content and the interactive content to be replied to, obtaining the corresponding target vector representations of the description and the interactive content to be replied to. The word vectors of each word in the description and the interactive content to be replied to in the target vector representation are fused, and the first style distribution is obtained based on the first fused vector representation.
[0121] In practical applications, when encoding the description of media content and the interactive content to be replied to, the server will segment the description and the interactive content to be replied to into words to obtain a word set. Then, each word in the word set will be encoded to obtain the corresponding vector representation of each word. The corresponding vector representation of each word will be used as the corresponding target vector representation.
[0122] In a specific application, when encoding each word in a word set to obtain its corresponding vector representation, the server first serializes each word in the word set to obtain its corresponding serialized representation, and then encodes the serialized representation to obtain its corresponding vector representation.
[0123] In practical applications, the server performs pooling operations on the word vectors of the descriptive content in the target vector representation and the word vectors of each word in the interactive content to be replied to. The pooled vectors are then fused, and style classification is performed based on the first fused vector representation to obtain a first style distribution. The pooling operation can specifically be max pooling, mean pooling, etc. This embodiment does not specifically limit the pooling operation. Max pooling refers to taking the point with the largest value in the local receptive field, while mean pooling refers to averaging all values in the local receptive field.
[0124] In a specific application, the server can perform style recognition based on a pre-trained interactive content style recognition model to obtain a first style distribution. This is achieved by inputting the description of the media content and the interactive content to be replied to into the pre-trained model. This pre-trained model can be trained under supervised conditions on a dataset of interactive content labeled with style categories. For example, the format of the labeled interactive content dataset could be something like this: [Media content 1 Interactive content 1 Interactive content 1 Style Interactive content reply 1 Interactive content reply 1 Style.......].
[0125] In a specific application, this embodiment does not limit the specific interactive content style recognition model, as long as it can achieve interactive content style recognition. For example, Figure 3 As shown, the pre-trained interactive content style recognition model can be a BERT-based model. During style recognition, the BERT model (including a Transformer Encoder, specifically a 12-layer encoder) first encodes the CLS flag, the description of the media content, and the interactive content to be replied to, obtaining the vector representations of the CLS flag, the description, and the target vector representations of the interactive content to be replied to. Then, a fully connected layer fuses the word vectors of the CLS flag vector, the description, and the interactive content to be replied to. The first fused vector representation is then passed through a classification output layer to obtain the first style distribution.
[0126] Specifically, when encoding the CLS flag, the description of the media content, and the interactive content to be replied to using the BERT model, the description of the media content and the interactive content to be replied to are first segmented into words to obtain the corresponding word sets of the description and interactive content to be replied to. Each word in the corresponding word set is then serialized to obtain the corresponding serialized representation of each word, and then the corresponding serialized representation of each word is encoded.
[0127] Specifically, when fusing the word vectors of the CLS flag corresponding vector representation, the descriptive content in the target vector representation, and the word vectors of each word in the interactive content to be replied to, it can be done by max pooling the word vectors of the CLS flag corresponding vector representation, the descriptive content in the target vector representation, and the word vectors of each word in the interactive content to be replied to; that is, pooling first and then fusing. For example, Figure 3 As shown, the CLS flag is placed first, and the corresponding vector representation of the CLS flag obtained by the BERT model can be used for subsequent classification tasks.
[0128] In this embodiment, by encoding the descriptive content of the media content and the interactive content to be replied to, the target vector representations of the descriptive content and the interactive content to be replied to can be obtained, and thus the first style distribution can be obtained based on the target vector representations.
[0129] In one embodiment, the publishing object information includes at least one historical interaction content of the publishing object. Style identification is performed based on the publishing object information of the interaction content to be responded to, resulting in a second style distribution including:
[0130] Perform style identification on at least one historical interaction content to obtain the historical style distribution corresponding to at least one historical interaction content;
[0131] The second style distribution is obtained by weighted averaging the style category probabilities of each historical style category in the historical style distribution.
[0132] Historical interaction content refers to interactive content that the publisher has previously posted. For example, historical interaction content could specifically be comments posted by the publisher in the past. Another example is bullet comments posted by the publisher in the past. Historical style distribution describes the distribution of style categories within historical interaction content.
[0133] Specifically, the server encodes at least one historical interaction to obtain a vector representation of that interaction. Based on this vector representation, it then obtains a historical style distribution corresponding to that historical interaction. Finally, it performs a weighted average of the probabilities of each style category within this distribution to obtain a second style distribution. For example, if the first style distribution is [philosophical 0.6, humorous 0.2, parody 0.2] and the second style distribution is [humorous 0.4, parody 0.4, poetic 0.2], then the style distribution of the interaction to be replied to will be [philosophical 0.3, humorous 0.3, parody 0.3, poetic 0.1].
[0134] In practical applications, the server segments at least one piece of historical interaction content into words, obtaining a corresponding word set. Then, it encodes each word in the word set to obtain a vector representation for each word. This vector representation is then used as the vector representation for the at least one piece of historical interaction content. In another specific application, the server first serializes each word in the word set to obtain a serialized representation for each word. Then, it encodes this serialized representation to obtain a vector representation for each word.
[0135] In practical applications, the server fuses the vector representations of each word in the vector representation of at least one historical interaction content, and performs style classification based on the second fused vector representation to obtain the historical style distribution corresponding to at least one historical interaction content. In a specific application, the server first performs a pooling operation on the vector representations of each word in the vector representation, and then fuses the pooled vectors. The pooling operation can specifically be max pooling, mean pooling, etc. This embodiment does not specifically limit the pooling operation. Max pooling refers to taking the point with the largest value in the local receptive field, while mean pooling refers to averaging all values in the local receptive field.
[0136] In a specific application, the server can perform style recognition based on a pre-trained interactive content style recognition model to obtain the historical style distribution corresponding to at least one historical interactive content. That is, by inputting at least one historical interactive content into the pre-trained interactive content style recognition model, the server obtains the historical style distribution corresponding to at least one historical interactive content. This pre-trained interactive content style recognition model can be obtained through supervised training on an interactive content dataset labeled with style categories. For example, the format of the interactive content dataset labeled with style categories can be [Media Content 1 Interactive Content 1 Interactive Content 1 Style Interactive Content Reply 1 Interactive Content Reply 1 Style.......]
[0137] In a specific application, this embodiment does not limit the specific interactive content style recognition model, as long as it can achieve interactive content style recognition. For example, Figure 4 As shown, the pre-trained interactive content style recognition model can be a BERT-based model. During style recognition, the BERT model (including a Transformer Encoder, specifically a 12-layer encoder) first encodes the CLS flag and historical interactive content, obtaining vector representations of the CLS flag and historical interactive content. Then, a fully connected layer fuses these vector representations. Finally, the fused vector representation is passed through a classification output layer to obtain the historical style distribution of the historical interactive content.
[0138] In the BERT model, when encoding the CLS flag and historical interaction content, the historical interaction content is first segmented into words to obtain a corresponding word set. Then, each word in the word set is serialized to obtain its serialized representation, which is then encoded. When fusing the vector representations of the CLS flag and the historical interaction content, a max-pooling method can be used, i.e., pooling followed by fusion. For example... Figure 4 As shown, the CLS flag is placed first, and the corresponding vector representation of the CLS flag obtained by the BERT model can be used for subsequent classification tasks.
[0139] In this embodiment, by performing style recognition on at least one historical interactive content, a historical style distribution corresponding to at least one historical interactive content can be obtained. Then, by weighting the style category probabilities of each historical style category in the historical style distribution, a style distribution of the publishing object can be constructed, thereby obtaining the second style distribution.
[0140] In one embodiment, the response content processing method further includes:
[0141] Obtain similar interactive content to the interactive content to be responded to from the interactive content of media content;
[0142] Based on the description content, the interaction content to be responded to, and the style category vector, generate response content corresponding to at least one style category in the style distribution of the interaction content to be responded to:
[0143] Based on the description content, the interactive content to be replied to, similar interactive content, and style category vector, generate reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
[0144] Similar interactive content refers to interactive content in media content that is similar to interactive content.
[0145] Specifically, the server, based on the interactive content of the media content, the style distribution of the interactive content, the interactive content to be replied to, and the style distribution of the interactive content to be replied to, obtains similar interactive content to the interactive content to be replied to from the interactive content of the media content. Then, based on the description content, the interactive content to be replied to, and the style category vector, it performs response word prediction to obtain at least one predicted response word corresponding to at least one style category in the style distribution of the interactive content to be replied to. Based on the at least one predicted response word corresponding to at least one style category, it generates response content corresponding to at least one style category. In specific applications, the server first performs style recognition on the interactive content of the media content to obtain the style distribution of the interactive content. The method of style recognition on interactive content is similar to the method of style recognition on at least one historical interactive content described above, and will not be repeated here in this embodiment.
[0146] In this embodiment, by obtaining similar interactive content to the interactive content to be replied to from the interactive content of the media content, the text of the generated reply content can be enriched and improved. By generating reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to based on the description content, the interactive content to be replied to, the similar interactive content, and the style category vector, the generation quality can be improved.
[0147] In one embodiment, obtaining similar interactive content to the interactive content to be responded to from the interactive content of media content includes:
[0148] The style similarity between the interactive content to be responded to and the interactive content of the media is obtained separately, and the content similarity between the interactive content to be responded to and each interactive content is obtained separately;
[0149] Based on the style and content similarity of each interactive content, similar interactive content to the interactive content to be responded to is selected from the interactive content of the media content.
[0150] Style similarity describes the degree of similarity in style between the interactive content to be responded to and the interactive content of the media content. For example, style similarity can specifically refer to the similarity between the style distribution of the interactive content to be responded to and the style distribution of the interactive content of the media content. Content similarity describes the degree of similarity between the interactive content to be responded to and the interactive content of the media content.
[0151] Specifically, the server obtains the style similarity between the interactive content to be replied to and each interactive content of the media content, based on the style distribution of the interactive content to be replied to and the corresponding style distribution of each interactive content of the media content. It also obtains the content similarity between the interactive content to be replied to and each interactive content. The style similarity and content similarity of each interactive content are then weighted and summed to obtain the similarity between each interactive content and the interactive content to be replied to. Based on the similarity between each interactive content and the interactive content to be replied to, interactive content from the interactive content of the media content with a similarity greater than a similarity threshold is selected as similar interactive content to the interactive content to be replied to. The weighting coefficients assigned to style similarity and content similarity during the weighted summation, as well as the similarity threshold, can be configured according to the actual application scenario.
[0152] In practical applications, style distribution includes a set of style categories and the corresponding style category probability for each style category within that set. The server can calculate the style similarity between the style distribution of the interactive content to be responded to and the style distribution of each interactive content within the media content. In a specific application, the server can calculate the cosine similarity between the style distribution of the interactive content to be responded to and the style distribution of each interactive content within the media content, and then obtain the style similarity for each interactive content based on the cosine similarity. For example, the style similarity can specifically be (1 - cosine similarity).
[0153] In practical applications, when obtaining the content similarity between the interactive content to be replied to and each interactive content, the server can first calculate the edit distance between the interactive content to be replied to and each interactive content. Then, based on the edit distance, the content length of the interactive content to be replied to, and each interactive content, the server can obtain the corresponding content similarity of each interactive content. Edit distance, also known as Lewinstein distance, is a quantitative measure of the difference between two strings (e.g., English words). It measures the minimum number of processing steps required to transform one string into another. In this embodiment, the primary focus is on determining the minimum number of processing steps required to transform interactive content into interactive content to be replied to.
[0154] In a specific application, the server determines the maximum content length for each interactive content based on the content to be replied to and the content length of each interactive content. For each interactive content, it calculates the ratio of the edit distance between the content to be replied to and the interactive content to the maximum content length of the interactive content. Based on this ratio, the content similarity of the interactive content is obtained. For example, the content similarity can be (1 - edit distance / maximum content length).
[0155] In this embodiment, by obtaining the style similarity and content similarity between the interactive content to be replied to and the interactive content of the media content, it is possible to select similar interactive content to be replied to from the interactive content of the media content, considering both style and content.
[0156] In one embodiment, generating response content corresponding to at least one style category in the style distribution of the interaction content to be responded to, based on the description content, the interaction content to be responded to, similar interaction content, and style category vectors, includes:
[0157] The description content, the interactive content to be replied to, and similar interactive content are segmented into words to obtain a set of segmented words;
[0158] The words in the word segmentation set are vectorized to obtain the vectorized representation of each word;
[0159] Based on the vectorized representation of each word and the style category vector, the response word is predicted to obtain at least one predicted response word corresponding to at least one style category in the style distribution of the interactive content to be responded to.
[0160] Generate response content corresponding to at least one style category based on at least one predicted response word corresponding to at least one style category.
[0161] Specifically, the server segments the description content, the interactive content to be replied to, and similar interactive content into words, obtaining a set of segmented words. Then, it vectorizes each word in the word set through encoding to obtain the vectorized representation of each word. Based on the vectorized representation of each word and the style category vector, it predicts the reply word. That is, based on the copy mechanism, it chooses to generate the reply word from the preset word list or directly copy a word from each word in the segmented word set as the predicted reply word. It obtains at least one predicted reply word corresponding to at least one style category in the style distribution of the interactive content to be replied to. It combines at least one predicted reply word corresponding to at least one style category to generate reply content corresponding to at least one style category.
[0162] In practical applications, when generating each predicted response word, the server determines the copy probability of each word in the word segmentation set. Based on the copy probability, it determines whether to generate a response word from a preset word list or directly copy a word from each word in the word segmentation set as the predicted response word. When the copy probability of each word in the word segmentation set is less than the copy probability threshold, it determines to generate a response word from the preset word list. When there is at least one target word in the word segmentation set with a copy probability greater than or equal to the copy probability threshold, it determines to directly copy a word from each word in the word segmentation set as the predicted response word. The copied word is the word with the highest copy probability among at least one target word.
[0163] In a specific application, the server can generate response content corresponding to at least one style category in the style distribution of the interactive content to be responded to, based on a pre-trained response content generation model. This pre-trained response content generation model can be obtained through supervised training on a dataset of interactive content labeled with style categories. During training, the style category vector is represented by the vector corresponding to the labeled style category. For example, the format of the interactive content dataset labeled with style categories can be [Media Content 1 Interactive Content 1 Interactive Content 1 Style Interactive Content Reply 1 Interactive Content Reply 1 Style.......].
[0164] In a specific application, this embodiment does not limit the specific response content generation model, as long as it can achieve response content generation. For example, Figure 5 As shown, the pre-trained response content generation model can be a model based on Transformer Encoder and Transformer Decoder. When generating response content, the description, the interaction to be responded to, and similar interaction content are first segmented into words to obtain a word set. Then, each word in the word set is vectorized through encoding to obtain a vectorized representation of each word. Next, response word prediction is performed based on the vectorized representation of each word (i.e., copied from the source text) and the style category vector (i.e., style representation). Specifically, based on a copy mechanism (i.e., an encoder-decoder attention mechanism), the model chooses whether to generate a response word from a pre-defined vocabulary or directly copy a word from each word in the word set as the predicted response word. This obtains at least one predicted response word corresponding to at least one style category in the style distribution of the interaction to be responded to. Combining these at least one predicted response word corresponding to at least one style category generates response content corresponding to at least one style category.
[0165] In the process of predicting response words based on the vectorized representations and style category vectors of each word, the style category vector and start symbol corresponding to the style category for which the response content needs to be generated are first processed. <s>The process involves decoding the style category vector and the corresponding word vector, predicting response word 1 based on the decoded vector and the vectorized representation of each word, and then decoding the style category vector and response word 1 again. This process continues until the prediction stops. The resulting n predicted response words are then combined to generate a response for at least one style category. The minimum value of n is 1, and the maximum value can be configured in the prediction stopping condition according to the actual application scenario.
[0166] In this embodiment, by segmenting and vectorizing the descriptive content, the interactive content to be replied to, and the similar interactive content, vectorized representations of each word in the descriptive content, the interactive content to be replied to, and the similar interactive content can be obtained. Then, based on the vectorized representations of each word and the style category vector, the reply word can be accurately predicted, improving the generation effect of proper words related to media content. At least one predicted reply word corresponding to at least one style category in the style distribution of the interactive content to be replied to can be obtained, so as to generate reply content corresponding to at least one style category based on at least one predicted reply word corresponding to at least one style category.
[0167] In one embodiment, response word prediction is performed based on the vectorized representation and style category vector of each word to obtain at least one predicted response word corresponding to at least one style category in the style distribution of the interactive content to be responded to, including:
[0168] Based on the vectorized representation of each word and the style category vector, the response word is predicted to obtain the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to.
[0169] When the prediction of the response word meets the prediction stopping condition, the current response word is taken as at least one predicted response word corresponding to at least one style category;
[0170] When the prediction of the response word meets the conditions for continuing prediction, the prediction of the response word continues based on the vectorized representation of each word, the style category vector, and the current response word, so as to obtain at least one predicted response word corresponding to at least one style category.
[0171] The prediction stopping condition refers to the condition for stopping response word prediction, which can be configured according to the actual application scenario. For example, the prediction stopping condition can be that the number of predicted response words reaches a response word count threshold. This threshold can be configured according to the actual application scenario. Another example is that the copy probability of each word in the description content, the interaction content to be replied to, and similar interaction content, as well as the word similarity of each candidate word in the preset word list, are all less than a preset response word probability threshold. This threshold can be configured according to the actual application scenario. Yet another example is that the number of predicted response words reaches the response word count threshold, or the copy probability of each word in the description content, the interaction content to be replied to, and similar interaction content, as well as the word similarity of each candidate word in the preset word list, are all less than the preset response word probability threshold. The continue prediction condition refers to the condition for continuing response word prediction. When the response word prediction no longer meets the prediction stopping condition, the continue prediction condition is considered met.
[0172] Specifically, the server predicts response words based on the vectorized representations of each word and the style category vectors. First, it obtains the current response word corresponding to at least one style category in the style distribution of the interactive content to be replied to. It then determines whether the current response word prediction meets the prediction stopping condition. If the response word prediction meets the prediction stopping condition, the prediction stops, and the current response word is taken as at least one predicted response word corresponding to at least one style category. If the response word prediction does not meet the prediction stopping condition, it is considered that the prediction meets the condition to continue. Based on the vectorized representations of each word, the style category vectors, and the current response word, the server continues to predict response words to obtain at least one predicted response word corresponding to at least one style category. At this time, the number of predicted response words is at least two.
[0173] In this embodiment, by predicting the response word based on the vectorized representation of each word and the style category vector, it is possible to obtain the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to. By using prediction stop conditions and prediction continue conditions, it is possible to determine whether to stop the prediction of the response word and obtain at least one predicted response word corresponding to at least one style category.
[0174] In one embodiment, response word prediction is performed based on the vectorized representation of each word and the style category vector to obtain the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to, including:
[0175] The similarity between the vectorized representation and style category vector of each word is calculated to obtain the copy probability of each word.
[0176] When there is at least one target word with a copy probability greater than or equal to the copy probability threshold, obtain the current reply word corresponding to at least one style category in the style distribution of the interactive content to be replied to based on at least one target word;
[0177] When the copy probability is less than the copy probability threshold, obtain at least one current response word corresponding to a style category from the preset word list.
[0178] The copy probability is used to characterize the probability that each word will be directly copied as the current response word. The preset vocabulary refers to a vocabulary that is pre-configured according to the actual application scenario, including the vectorized representation of at least one candidate word.
[0179] Specifically, the server calculates the similarity between the vectorized representation and style category vector of each word to obtain the copy probability of each word. It then compares the copy probability of each word with a copy probability threshold. If there is at least one target word with a copy probability greater than or equal to the copy probability threshold, it means that a word can be directly copied from at least one target word as the current response word. The copied word is the one with the highest copy probability among the at least one target word. If the copy probabilities are all less than the copy probability threshold, it means that a response word needs to be generated from a preset word list. The server will obtain at least one current response word corresponding to a style category from the preset word list.
[0180] In practical applications, when the number of target words is 1, the target word can be directly copied as the current reply word. When the number of target words is at least 2, it is necessary to compare the copy probabilities of the target words and select the target word with the highest copy probability as the current reply word.
[0181] In this embodiment, by calculating the similarity between the vectorized representation of each word and the style category vector, the copy probability of each word can be obtained. Thus, by comparing the copy probability and the copy probability threshold, it can be determined whether to generate a reply word from a preset word list or to directly copy a word from each word in the word segmentation set as the current reply word. The generation of the current reply word is realized based on the copy mechanism.
[0182] In one embodiment, the preset vocabulary includes vectorized representations of at least one candidate word. When the copy probability is less than a copy probability threshold, the current response word corresponding to at least one style category is obtained from the preset vocabulary, including:
[0183] When the copy probability is less than the copy probability threshold, the similarity is calculated for the vectorized representation and style category vector of each candidate word in the preset vocabulary to obtain the word similarity of each candidate word.
[0184] Based on word similarity, obtain the current response word corresponding to at least one style category.
[0185] Specifically, when the copy probability is less than the copy probability threshold, the server will calculate the similarity between the vectorized representation and style category vector of each candidate word in the preset word list, obtain the word similarity of each candidate word, compare the word similarity of each candidate word, and select the candidate word with the largest word similarity as the current response word corresponding to at least one style category.
[0186] In this embodiment, by calculating the similarity between the vectorized representation of each candidate word in the preset word list and the style category vector, the word similarity of each candidate word can be obtained. Then, based on the word similarity, the current reply word can be determined, and at least one style category corresponding to the current reply word can be obtained.
[0187] In one embodiment, when the predicted response word meets the conditions for continuing prediction, the prediction of the response word continues based on the vectorized representation of each word, the style category vector, and the current response word, to obtain at least one predicted response word corresponding to at least one style category, including:
[0188] When the prediction of the response word meets the conditions for continuing the prediction, the next response word is predicted based on the vectorized representation of each word, the style category vector, and the current response word.
[0189] The next reply word is used as the new current reply word. The process then jumps to the step of predicting the next reply word based on the vectorized representation of each word, the style category vector, and the current reply word.
[0190] Until the prediction of the response word meets the prediction stopping condition, based on the current response word obtained from each prediction of the response word, at least one predicted response word corresponding to at least one style category is obtained.
[0191] Specifically, when the response word prediction meets the conditions for continuing prediction, it means that in addition to generating the current response word, it is necessary to continue predicting response words. The server will predict response words based on the vectorized representation of each word, the style category vector, and the current response word to obtain the next response word corresponding to the current response word. The next response word is used as the new current response word, and the process jumps to the step of predicting response words based on the vectorized representation of each word, the style category vector, and the current response word to obtain the next response word corresponding to the current response word. This process continues to generate new current response words until the response word prediction meets the prediction stopping condition. Based on the current response words obtained from each response word prediction, at least one predicted response word corresponding to at least one style category is obtained.
[0192] In practical applications, when the next response word corresponding to the current response word is obtained each time, the server first decodes the style category vector and the current response word to obtain the decoded vector. Then, based on the decoded vector and the vectorized representation of each word, the server predicts the response word. That is, based on the copy mechanism, the server chooses to generate the next response word corresponding to the current response word from the preset word list or to directly copy a word from each word in the word segmentation set as the next response word corresponding to the current response word.
[0193] In a specific application, the server calculates the similarity between the vectorized representation and the decoded vector of each word, obtaining the copy probability of each word and its corresponding decoded vector. It then compares the copy probabilities of each word with the corresponding copy probability threshold. If there is at least one target word with a copy probability greater than or equal to the copy probability threshold, it means that a word can be directly copied from at least one target word as the next response word. The copied word is the one with the highest copy probability among the at least one target word. If all copy probabilities are less than the copy probability threshold, it means that the next response word needs to be generated from a preset vocabulary list, and the server will obtain the next response word from the preset vocabulary list.
[0194] In a specific application, when the copy probability is less than the copy probability threshold, the server will calculate the similarity between the vectorized representation and style category vector of each candidate word in the preset vocabulary, obtain the word similarity between each candidate word and the decoded vector, compare the word similarity between each candidate word and the decoded vector, and select the candidate word with the highest word similarity as the next response word.
[0195] In one embodiment, after generating response content corresponding to at least one style category in the style distribution of the interaction content to be responded to based on the description content, the interaction content to be responded to, and the style category vector, the method further includes:
[0196] Obtain the style distribution of the reply content and the style distribution of the reply content's publisher;
[0197] Based on the style distribution of the corresponding reply content and the style distribution of the object that published the reply content, the style similarity of the corresponding reply content is obtained;
[0198] Interaction prediction is performed based on the response content, the corresponding style distribution of the response content, the description content, the interaction content to be responded to, and the style distribution of the interaction content to be responded to, to obtain the estimated interaction rate of the response content.
[0199] The responses are sorted based on their stylistic similarity and estimated interaction rate to obtain the ranking results.
[0200] The style distribution of the response content describes the distribution of style categories among the response content. The style distribution of the target audience for the response content describes the distribution of style categories among the target audience for the response content. Interaction prediction refers to predicting whether the response content will be interacted with; the estimated interaction rate is the predicted probability that the response content will be interacted with.
[0201] Specifically, the server performs style recognition on the reply content to obtain the corresponding style distribution, and retrieves at least one published interaction from the corresponding reply content's publisher. Based on this published interaction, it obtains the style distribution of the reply content's publisher. After obtaining the style distributions of the reply content and the publisher, the server calculates the style distribution similarity between them, and obtains the style similarity of the reply content. Based on the style distribution similarity, it performs interaction prediction based on the reply content, its corresponding style distribution, description content, the interaction to be replied to, and the style distribution of the interaction to be replied to, obtaining the estimated interaction rate. Finally, based on the style similarity and the estimated interaction rate, the reply content is sorted to obtain the ranking result.
[0202] In practical applications, the method for style identification of reply content is similar to that for style identification of at least one historical interaction, and will not be elaborated further in this embodiment. When obtaining the style distribution of the reply content's publishing object, the server performs style identification on at least one published interaction, obtaining a style distribution corresponding to that at least one published interaction. A weighted average of the style category probabilities of each style category in the style distribution corresponding to the at least one published interaction is then performed to obtain the style distribution of the reply content's publishing object. The method for style identification of at least one published interaction is similar to that for style identification of at least one historical interaction, and will not be elaborated further in this embodiment.
[0203] In practical applications, the style distribution similarity between the style distribution of the response content and the style distribution of the object that published the response content can be specifically cosine similarity. The server can first calculate the cosine similarity of the style distribution, and then obtain the style similarity of the response content based on the cosine similarity. For example, the style similarity can be (1 - cosine similarity).
[0204] In this embodiment, the style similarity of the reply content can be obtained based on the style distribution of the reply content and the style distribution of the object that published the reply content. Interaction prediction can be achieved based on the reply content, the style distribution of the reply content, the description content, the interactive content to be replied to, and the style distribution of the interactive content to be replied to, so as to obtain the estimated interaction rate of the reply content. Then, the reply content can be sorted according to the style similarity of the reply content and the estimated interaction rate to obtain the reply content sorting result. In order to sort and display the reply content according to the reply content sorting result, the adoption rate and interaction rate of the reply content can be improved.
[0205] In one embodiment, interaction prediction is performed based on the response content, the corresponding style distribution of the response content, the descriptive content, the interaction content to be responded to, and the style distribution of the interaction content to be responded to, to obtain the estimated interaction rate corresponding to the response content, including:
[0206] The reply content and its corresponding style distribution are encoded to obtain the first vector representation of the reply content, and the style distribution of the interaction content to be replied to and the interaction content to be replied to are encoded to obtain the second vector representation of the interaction content to be replied to.
[0207] The description content is encoded to obtain a corresponding vector representation of the description content;
[0208] The vector representation, first vector representation, and second vector representation of the description content are fused to obtain the estimated interaction rate of the response content.
[0209] Specifically, the server encodes the reply content and its corresponding style distribution to obtain a first vector representation of the reply content, and encodes the style distribution of the interaction content to be replied to to obtain a second vector representation of the interaction content to be replied to. The server also encodes the description content to obtain a vector representation of the description content. The server then merges the vector representation of the description content, the first vector representation, and the second vector representation, and obtains the estimated interaction rate of the reply content based on the third merged vector representation.
[0210] In specific applications, this embodiment does not limit the method of fusing the vector representations of the response content, the interactive content to be responded to, and the descriptive content. For example, the vector representations of the descriptive content, the first vector representation, and the second vector representation can be fused based on an attention mechanism. The attention mechanism is an allocation mechanism whose core idea is to highlight certain important features of an object and reallocate resources (i.e., weights) according to the importance of the attention object. The core idea is to find the correlation between existing data and then highlight certain important features. In this embodiment, the correlation between the response content, the interactive content to be responded to, and the descriptive content is found based on the first vector representation, the second vector representation, and the vector representation of the descriptive content. This is achieved by fusion to highlight certain important features in their vector representations, so that the estimated interaction rate of the response content can be obtained based on these important features.
[0211] In a specific application, the server can obtain the estimated interaction rate of the reply content based on a pre-trained interaction rate prediction model. This pre-trained interaction rate prediction model can be trained on data such as published reply content and published interaction content. Data with interaction counts meeting a certain threshold can be selected as positive samples, and data with interaction counts not meeting the threshold can be selected as negative samples for training. The threshold can be configured according to the actual application scenario, and the number of interactions can include the number of likes, shares, replies, etc.
[0212] In a specific application, this embodiment does not limit the specific interaction rate prediction model, as long as it can achieve interaction prediction. For example, Figure 6 As shown, the interaction rate prediction model can be a BERT-based model. When predicting interaction, the BERT model encodes the response content and its corresponding style distribution, the interaction content to be responded to, its style distribution, and the description content, respectively, resulting in a first vector representation of the response content, a second vector representation of the interaction content to be responded to, and a vector representation of the description content. The vector representations of the description content, the first vector representation, and the second vector representation are then fused, and the predicted interaction rate for the response content is obtained based on the third fused vector representation.
[0213] In this embodiment, interaction prediction can be performed by comprehensively considering the response content, the corresponding style distribution of the response content, the descriptive content, the interactive content to be responded to, and the style distribution of the interactive content to be responded to, thereby determining the estimated interaction rate of the response content.
[0214] In one embodiment, the reply content processing method of this application is illustrated by its application to generating replies to dynamic style video comments. Here, replying to a video comment refers to the process by which an object viewing a video, when browsing comments from other objects below the video, considers the style of their reply when intending to discuss and interact with these comments, such as a poetic style or a humorous style. For example, the object viewing the video can specifically be a user viewing the video.
[0215] The inventors believe that using the content of comments to be replied to to generate replies results in monotonous reply styles, failing to meet the diverse style needs of the vast audience watching videos on video platforms. Furthermore, it leads to a lack of variety in the style of interactions within the video platform community, reducing overall platform activity. This application's reply content processing method, applied to the generation of dynamic style video comment replies, identifies the style distribution of comments to be replied to through joint deep modeling analysis of video content, comments to be replied to, and the commenting audience. It then generates multiple styles of reply content for the current video viewer, allowing them to directly select a suitable style, reducing input costs and improving interaction efficiency. The reply styles are also more aligned with the viewer's intent. Furthermore, by combining the relevance of the generated replies to the current viewer and interaction prediction, the method optimizes different reply styles, further enhancing viewer satisfaction and efficiency with the selective reply approach. By combining the style preferences of comments to be replied to, the commenting audience, and the current replyer, and predicting future interactions, the method can generate diverse styles of comments for different viewers at different times, increasing the dynamic diversity of comment replies and enhancing the interactive atmosphere of the video platform.
[0216] Specifically, a flowchart illustrating the process of generating dynamic style video comment replies is shown below. Figure 7 As shown. In response to the operation of replying to comments under a video, the server will perform style recognition on the comments to be replied to, style recognition on the object that posted the comments to be replied to, and extract key text information. Based on the style distribution of the comments to be replied to, the style distribution of the object that posted the comments to be replied to, and the extracted key text information, the server will generate video comment replies of different styles, sort the video comment replies of different styles, and return the multi-style video comment replies for the audience to use.
[0217] First, it is necessary to identify the style distribution of the video (media content) comments to be replied to (i.e., the interactive content to be replied to), that is, to identify the style distribution of the comments to be replied to.
[0218] The server performs style recognition based on the video content (i.e., the descriptive content) and the comment text of the video comments to obtain the style distribution of the comments to be replied to. In specific applications, the server encodes the video content and the comment text of the video comments to obtain the corresponding target vector representations of the video content and the comment text. The word vectors of each word in the descriptive content and the interactive content to be replied to in the target vector representation are fused. Based on the first fused vector representation, the initial style distribution of the comments to be replied to is obtained (i.e., the first style distribution).
[0219] In a specific application, the server can, for example, Figure 8 The video comment style recognition model shown identifies the comment style distribution, i.e., the probability distribution of comment style types, for comments to be replied to. After inputting the video content and comment text into the model, it encodes the CLS flag, video content, and comment text using a BERT model, obtaining vector representations of the CLS flag and target vector representations of the video content and comment text. A fully connected network then fuses the word vectors of the descriptive content in the CLS flag vector representation, the target vector representation, and the word vectors in the interactive content to be replied to. Based on the first fused vector representation, style classification is performed to obtain the comment style distribution. The video comment style recognition model can be trained under supervision on a dataset of interactive content labeled with style categories. For example, the format of the interactive content dataset labeled with style categories can be [Media Content 1 Interactive Content 1 Interactive Content 1 Style Interactive Content Reply 1 Interactive Content Reply 1 Style.......].
[0220] Secondly, it is necessary to identify the style of the person who posted the comment to be replied to.
[0221] The server retrieves at least one historical comment (i.e., historical interaction content) from the target user, performs style identification on this comment, and obtains a historical style distribution corresponding to it. It then performs a weighted average of the style category probabilities for each historical style category in the historical style distribution to obtain the style distribution of the target user for the comment to be replied to (i.e., the second style distribution). The inventors believe that considering the style of the target user for the comment to be replied to improves the accuracy of comment style identification. Therefore, the server comprehensively considers both the initially obtained style distribution of the target comment and the style distribution of the target user for the comment to be replied to, to determine the style distribution of the target comment (i.e., the style distribution of the interaction content to be replied to).
[0222] In a specific application, the server can, for example, Figure 8 The video comment style recognition model shown identifies the historical style distribution of previously posted comments. After inputting previously posted comments into the model, it encodes the CLS flag and the historical comments using a BERT model, obtaining vector representations of both the CLS flag and the historical comments. These vector representations are then fused using a fully connected network. Based on this second fused vector representation, style classification is performed to obtain the historical style distribution of the previously posted comments.
[0223] The video comment style recognition model can be obtained through supervised training on an interactive content dataset labeled with style categories. For example, the format of the interactive content dataset labeled with style categories can be [Media Content 1 Interactive Content 1 Interactive Content 1 Style Interactive Content Reply 1 Interactive Content Reply 1 Style.......]
[0224] Secondly, it is necessary to extract key text information.
[0225] To enrich and improve the text used to generate replies and enhance the quality of the generated content, the server performs style recognition on other comments (excluding the comment to be replied to) and replies (i.e., interactive content) of the video, using other comments and replies with similar styles and content as key text information. At this point, the server calculates the style similarity between other comments and replies and the comment to be replied to, as well as the content similarity between other comments and replies and the comment to be replied to. Based on the corresponding style and content similarity of other comments and replies, the server determines the similarity between other comments and replies and the comment to be replied to, and uses other comments and replies whose similarity meets the similarity threshold as key text information.
[0226] The method for style identification of other comments and replies to a video is similar to that for style identification of at least one historically posted comment, and will not be elaborated further in this embodiment. In a specific application, the similarity between other comments and replies and the comment to be replied to = ws1 * style similarity + ws2 * content similarity, where ws1 and ws2 are weight coefficients configured according to the actual application scenario. Style similarity = the similarity of the style distribution vectors of the two comments, specifically (1 - cosine similarity). Content similarity = 1 - edit distance between the two comments / max (length of the two comments).
[0227] Secondly, it is necessary to generate video comment replies in different styles.
[0228] The server segments the video content, the comments to be replied to, and other similar comments into words, resulting in a word set. Each word in the word set is then vectorized to obtain its vector representation. Based on the vector representation of each word and the style category vector corresponding to at least one style category in the style distribution of the comments to be replied to, the server predicts the reply word, obtaining at least one predicted reply word corresponding to at least one style category in the style distribution of the comments to be replied to. Based on the at least one predicted reply word corresponding to at least one style category, the server generates a video comment reply (i.e., reply content) corresponding to at least one style category.
[0229] When generating predicted response words, the server uses a copy mechanism to choose between generating predicted response words from a preset vocabulary or directly copying a word from the description content and the words in the interaction content to be replied to. Specifically, the server predicts response words based on the vectorized representations and style category vectors of each word, obtaining the current response word corresponding to at least one style category in the style distribution of the comment to be replied to. When the response word prediction meets the prediction stopping condition, the current response word is used as at least one predicted response word corresponding to at least one style category. When the response word prediction meets the prediction continuing condition, the server predicts response words based on the vectorized representations of each word, the style category vector, and the current response word, obtaining the next response word corresponding to the current response word. The next response word is used as the new current response word, and the process jumps to the step of predicting response words based on the vectorized representations of each word, the style category vector, and the current response word to obtain the next response word corresponding to the current response word. This process continues until the response word prediction meets the prediction stopping condition. Based on the current response word obtained from each response word prediction, at least one predicted response word corresponding to at least one style category is obtained.
[0230] When generating the current reply word, the server calculates the similarity between the vectorized representation and style category vector of each word to obtain the copy probability of each word. When there is at least one target word with a copy probability greater than or equal to the copy probability threshold, the server obtains the current reply word corresponding to at least one style category in the style distribution of the comment to be replied to based on the at least one target word. When the copy probabilities are all less than the copy probability threshold, the server calculates the similarity between the vectorized representation and style category vector of each candidate word in the preset word list to obtain the word similarity of each candidate word. Based on the word similarity, the server obtains the current reply word corresponding to at least one style category.
[0231] When generating video comment responses corresponding to at least one style category, the server can filter out the style categories for which video comment responses need to be generated based on the style category probabilities of each style category. Specifically, for style categories whose probabilities meet a category probability threshold, the server generates the corresponding video comment responses. This category probability threshold can be configured according to the actual application scenario. This method saves computational resources when generating video comment responses corresponding to at least one style category in the style distribution of the interactive content to be responded to.
[0232] In a specific application, the server can, for example, Figure 9 The dynamic style comment reply generation model shown generates video comment replies corresponding to at least one style category. When generating video comment replies, the model first performs word segmentation and vectorization on the video content, other similar comment texts, and the current comment text to be replied to, obtaining vectorized representations of each word in the video content, other similar comment texts, and the current comment text to be replied to (i.e., copied from the source text). Then, based on the vectorized representations of each word and the style category vector, it predicts the reply word. Specifically, based on a copy mechanism (i.e., an encoder-decoder attention mechanism), it chooses whether to generate a reply word from a preset vocabulary or directly copy a word from each word as the predicted reply word. This obtains at least one predicted reply word corresponding to at least one style category in the style distribution of the comment to be replied to. Combining these at least one predicted reply word corresponding to at least one style category generates a video comment reply corresponding to at least one style category.
[0233] In the process of predicting response words based on the vectorized representations and style category vectors of each word, the style category vector and start symbol corresponding to the style category for which the video comment response needs to be generated are first processed. <s>The process involves decoding the style category vector and the corresponding word vector, predicting the first response word based on the decoded vector and the vectorized representation of each word. Then, the style category vector and the first response word are decoded again, and the second response word is predicted based on the decoded vector and the vectorized representation of each word. This process continues until a prediction stopping condition is met. The resulting n predicted response words are then combined to generate a video comment response corresponding to at least one style category. The minimum value of n is 1, and the maximum value can be configured in the prediction stopping condition according to the actual application scenario.
[0234] The dynamic style comment reply generation model can be obtained through supervised training on an interactive content dataset labeled with style categories. During training, the style category vectors are represented by the corresponding vectors of the labeled style categories. For example, the format of the interactive content dataset labeled with style categories can be [media content 1 interactive content 1 interactive content 1 style interactive content reply 1 interactive content reply 1 style.......].
[0235] Finally, the comments on videos of different styles need to be sorted and optimized.
[0236] The above steps generate various styles of replies for the current comments to be replied to. This step calculates the style matching of the video comment replies for each style category with the style of the audience currently viewing the video, and estimates the interaction rate of the generated reply content. This can increase the adoption rate of dynamic style replies by the audience viewing the video, while also increasing the interaction rate of the comment replies.
[0237] The server obtains the style distribution of video comment replies and the style distribution of the currently viewed video (i.e., the object that posted the video comment reply). Based on the style distribution of the video comment replies and the style distribution of the currently viewed video, it calculates the style similarity between the video comment replies for each style category and the currently viewed video, thus obtaining the style similarity of the video comment replies. Based on the video comment replies, the style distribution of the video comment replies, the video content, the comments to be replied to, and the style distribution of the comments to be replied to, it performs interaction prediction to obtain the estimated interaction rate of the video comment replies. Based on the style similarity of the video comment replies and the estimated interaction rate, it sorts the video comment replies to obtain the video comment reply ranking result.
[0238] The style similarity of video comment replies can be calculated as the cosine similarity between the style distribution of the video comment replies and the style distribution of the currently viewed video. During interaction prediction, the server encodes the video comment replies and their corresponding style distributions to obtain a first vector representation of the video comment replies. It also encodes the style distributions of the comments to be replied to and the comments to be replied to, obtaining a second vector representation of the comments to be replied to. The video content is then encoded to obtain its corresponding vector representation. Finally, the vector representations of the descriptive content, the first vector representation, and the second vector representation are fused together. Based on this third fused vector representation, the estimated interaction rate of the video comment replies is obtained.
[0239] The ranking of video comment replies is determined by their style similarity and estimated interaction rate. These scores are then used to rank the replies, resulting in a ranking result. The ranking score is calculated as: Ranking Score = w1 * Style Similarity + w2 * Estimated Interaction Rate, where w1 and w2 are weights and can be configured according to the specific application scenario.
[0240] In a specific application, the server can, for example, Figure 10 The different style-priority models shown generate the estimated interaction rate for each video comment reply. During interaction prediction, the BERT model encodes the video comment reply, its corresponding style score, the comment to be replied to, the style distribution of the comment to be replied to, and the video content, respectively. This yields a style-specific reply representation for the video comment reply, a text representation of the comment to be replied to, and a video content representation. These three representations are then fused, and based on the third fused vector representation, the probability of interaction for that style reply (i.e., the estimated interaction rate) is generated.
[0241] The different style comment reply priority model can be trained on data such as published reply content and published interaction content. Data with interaction counts meeting a certain threshold can be selected as positive samples, and data with interaction counts not meeting the threshold can be selected as negative samples for training. The threshold can be configured according to the actual application scenario, and the number of interaction counts can include the number of likes, reposts, replies, etc.
[0242] In one embodiment, such as Figure 11 As shown, an interactive method for media content is provided, which can be applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps:
[0243] Step 1102: Display the interactive content of the media content.
[0244] Specifically, the terminal displays interactive content containing media content. In practical applications, the interactive content can be comments posted by users viewing the media content. For example, when the interactive content is a comment, a diagram illustrating the interactive content displayed on the terminal could be as follows: Figure 12 As shown, while displaying the media content, interactive content (i.e., comment content 1, comment content 2, comment content 3, etc.) is also displayed. In specific applications, the interactive content can be bullet comments posted by the viewers of the media content. For example, when the interactive content is bullet comments, the diagram illustrating how the terminal displays the interactive content of the media content can be as follows: Figure 13 As shown, while the media content is displayed in full screen on the terminal, interactive content (i.e., bullet comments 1, 2, 3, 4, etc.) is displayed on top of the media content. When the interactive content is a bullet comment, its position and transparency can be set by the user viewing the media content.
[0245] Step 1104: In response to the trigger event of any interactive content, determine the interactive content to be replied to.
[0246] Specifically, the terminal will respond to the trigger event for any interactive content and determine the interactive content to be replied to. In a practical application, when an object browsing media content wants to reply to any interactive content, it will select any displayed interactive content, and the terminal will respond to the trigger event for any interactive content and determine the interactive content to be replied to.
[0247] In practical applications, the trigger event for any interactive content can be a selection event for that interactive content. Objects browsing media content can select any interactive content by clicking on its corresponding display area. The clicking method can be a single click, double click, etc., and this embodiment does not limit the clicking method here.
[0248] Step 1106: For the interactive content to be replied to, display alternative reply content corresponding to at least one style category of the interactive content to be replied to.
[0249] Among them, at least one style-corresponding alternative response content refers to the response content that the object browsing the media content can choose from. The object browsing the media content can select any alternative response content from at least one style-corresponding alternative response content of the displayed interactive content to be responded to as the interactive content to be responded to.
[0250] Specifically, the terminal, corresponding to the interactive content to be replied to, will display alternative response content corresponding to at least one style category of the interactive content to be replied to. In specific applications, the terminal, corresponding to the interactive content to be replied to, will obtain alternative response content corresponding to at least one style category of the interactive content to be replied to, and display alternative response content corresponding to at least one style category.
[0251] In a specific application, at least one style category's corresponding alternative response content can be displayed above other interactive content besides the content to be responded to. For example, an illustration of displaying at least one style category's corresponding alternative response content could be as follows: Figure 14 As shown, at least one style category of alternative response content is displayed below the interactive content to be responded to and above other interactive content (i.e., alternative response content 1, alternative response content 2, and alternative response content 3 in the figure), meaning that other interactive content is not displayed at this time.
[0252] In a specific application, alternative responses corresponding to at least one style category can be displayed on top of the media content. For example, an illustration of displaying alternative responses corresponding to at least one style category could be as follows: Figure 15 As shown, at least one style category's corresponding alternative response content is displayed above the media content, partially obscuring it (i.e., alternative response content 1, alternative response content 2, and alternative response content 3 in the figure). Furthermore, to avoid affecting the display of the media content, the viewer can adjust the display transparency of the alternative response content corresponding to at least one style category.
[0253] In a specific application, while displaying alternative responses corresponding to at least one style category, the terminal also displays the style category corresponding to the alternative responses, allowing users browsing media content to quickly select based on the style category. For example, a diagram illustrating the display of alternative responses corresponding to at least one style category can be shown below. Figure 16 As shown, while displaying alternative response content corresponding to at least one style category, the style category corresponding to the alternative response content is also displayed.
[0254] Step 1108: In response to the selection event of any candidate reply content, the candidate reply content selected by the selection operation is displayed as the interactive content to be replied to.
[0255] Specifically, users browsing media content can view alternative response content for at least one style category through the terminal. When they want to select any alternative response content as the interactive content to be replied to, the terminal will respond to the selection event of any alternative response content and display the selected alternative response content as the interactive content to be replied to.
[0256] In practical applications, users browsing media content can select any alternative response by clicking on the corresponding display area. The clicking method can be a single click, double click, etc., but this embodiment does not limit the clicking method.
[0257] The aforementioned interactive method for media content displays interactive content, responds to a trigger event for any interactive content, determines the interactive content to be replied to, and displays alternative reply content corresponding to at least one style category of the interactive content to be replied to. It can provide alternative reply content corresponding to at least one style category as interactive content selection, thereby responding to a selection event for any alternative reply content, displaying the alternative reply content selected by the selection operation as the interactive content to reply to the interactive content to be replied to, so that there is no need to input reply content to reply, which can improve reply efficiency.
[0258] In one embodiment, corresponding to the interactive content to be responded to, the alternative response content displayed for at least one style category of the interactive content to be responded to includes:
[0259] Corresponding to the interactive content to be replied to, alternative response content is displayed in descending order of the estimated interaction rate of the alternative response content, with at least one style category corresponding to the interactive content to be replied to.
[0260] Specifically, corresponding to the interactive content to be responded to, the terminal obtains the estimated interaction rates of the candidate response content in descending order, and displays candidate response content corresponding to at least one style category of the interactive content to be responded to, according to the descending order of the estimated interaction rates. In specific applications, the estimated interaction rate can be obtained using the method described in the above embodiments for obtaining the estimated interaction rate of the response content.
[0261] In practical applications, while displaying alternative response content corresponding to at least one style category of the interactive content to be responded to, the terminal will also display the estimated interaction rate of the alternative response content, so that the audience browsing the media content can select the alternative response content based on the estimated interaction rate.
[0262] In practical applications, while displaying alternative response content corresponding to at least one style category of the interactive content to be responded to, the terminal will also display the estimated interaction rate and style category of the alternative response content, so that the audience browsing the media content can select the alternative response content based on the estimated interaction rate and style category.
[0263] In this embodiment, by displaying alternative response content corresponding to at least one style category of the interactive content to be responded to in descending order of the estimated interaction rate of the alternative response content, it is possible to achieve rapid selection based on the estimated interaction rate.
[0264] In one embodiment, alternative response content corresponding to at least one style category is obtained through the above-described response content processing method.
[0265] Specifically, at least one alternative response content corresponding to a style category can be obtained through the above response content processing method. The terminal will perform style recognition on the interactive content to be responded to based on the description content of the media content, the interactive content to be responded to, and the publishing object information of the interactive content to be responded to, to obtain the style distribution of the interactive content to be responded to, to determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be responded to, and to generate alternative response content corresponding to at least one style category in the style distribution of the interactive content to be responded to based on the description content, the interactive content to be responded to, and the style category vector.
[0266] In this embodiment, the above-described response content processing method can be used to obtain alternative response content corresponding to at least one style category.
[0267] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0268] Based on the same inventive concept, this application also provides a response content processing device for implementing the above-mentioned response content processing method and a media content interactive content interaction device for implementing the media content interactive content interaction method. The solution provided by this device is similar to the solution described in the above-described method. Therefore, the specific limitations of the one or more response content processing device embodiments provided below can be found in the limitations of the response content processing method above, and the specific limitations of the one or more media content interactive content device embodiments provided below can be found in the limitations of the media content interactive content method above, and will not be repeated here.
[0269] In one embodiment, such as Figure 17 As shown, a response content processing device is provided, including: an interactive content acquisition module 1702, a style recognition module 1704, a style category vector determination module 1706, and a response content generation module 1708, wherein:
[0270] Interactive content acquisition module 1702 is used to acquire interactive content to be responded to in relation to media content;
[0271] The style recognition module 1704 is used to perform style recognition on the interactive content to be replied to based on the description content of the media content, the interactive content to be replied to, and the publishing object information of the interactive content to be replied to, so as to obtain the style distribution of the interactive content to be replied to.
[0272] Style category vector determination module 1706 is used to determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to;
[0273] The response content generation module 1708 is used to generate response content corresponding to at least one style category in the style distribution of the interaction content to be responded to, based on the description content, the interaction content to be responded to, and the style category vector.
[0274] The aforementioned response content processing device acquires the interactive content to be responded to for media content. Based on the description of the media content, the interactive content to be responded to, and the publishing information of the interactive content to be responded to, it performs style recognition on the interactive content to be responded to, thereby obtaining the style distribution of the interactive content to be responded to. This allows it to determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be responded to. Based on the description, the interactive content to be responded to, and the style category vector, it generates response content corresponding to at least one style category in the style distribution of the interactive content to be responded to. By automatically generating response content corresponding to at least one style category as the interactive content selection, it eliminates the need to input response content for replying, thereby improving response efficiency.
[0275] In one embodiment, the style recognition module is further configured to perform style recognition based on the description content of the media content and the interactive content to be replied to, to obtain a first style distribution; perform style recognition based on the publishing object information of the interactive content to be replied to, to obtain a second style distribution; and obtain the style distribution of the interactive content to be replied to based on the first style distribution and the second style distribution.
[0276] In one embodiment, the style recognition module is further used to encode the descriptive content of the media content and the interactive content to be replied to, to obtain the target vector representations of the descriptive content and the interactive content to be replied to, and to obtain the first style distribution based on the target vector representations.
[0277] In one embodiment, the publishing object information includes at least one historical interaction content of the publishing object, and the style recognition module is further used to perform style recognition on at least one historical interaction content to obtain a historical style distribution corresponding to at least one historical interaction content, and to perform a weighted average of the style category probabilities of each historical style category in the historical style distribution to obtain a second style distribution.
[0278] In one embodiment, the response content generation module is further configured to obtain similar interactive content to the interactive content to be responded to from the interactive content of the media content, and generate response content corresponding to at least one style category in the style distribution of the interactive content to be responded to based on the description content, the interactive content to be responded to, the similar interactive content, and the style category vector.
[0279] In one embodiment, the response content generation module is further configured to obtain the style similarity between the interactive content to be responded to and each interactive content of the media content, and to obtain the content similarity between the interactive content to be responded to and each interactive content, and to filter out similar interactive content to the interactive content to be responded to from the interactive content of the media content based on the corresponding style similarity and content similarity of each interactive content.
[0280] In one embodiment, the response content generation module is further configured to segment the description content, the interactive content to be responded to, and similar interactive content into words to obtain a segmented word set, vectorize each word in the segmented word set to obtain a vectorized representation of each word, predict response words based on the vectorized representation of each word and the style category vector, obtain at least one predicted response word corresponding to at least one style category in the style distribution of the interactive content to be responded to, and generate response content corresponding to at least one style category based on at least one predicted response word corresponding to at least one style category.
[0281] In one embodiment, the response content generation module is further configured to predict response words based on the vectorized representation of each word and the style category vector, to obtain the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to. When the response word prediction meets the prediction stopping condition, the current response word is used as at least one predicted response word corresponding to at least one style category. When the response word prediction meets the prediction continuing condition, the response word prediction is continued based on the vectorized representation of each word, the style category vector and the current response word, to obtain at least one predicted response word corresponding to at least one style category.
[0282] In one embodiment, the response content generation module is further configured to perform similarity calculation on the vectorized representation and style category vector of each word to obtain the copy probability of each word. When there is at least one target word with a copy probability greater than or equal to the copy probability threshold, the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to is obtained based on at least one target word. When the copy probabilities are all less than the copy probability threshold, the current response word corresponding to at least one style category is obtained from the preset word list.
[0283] In one embodiment, the preset vocabulary includes vectorized representations of at least one candidate word. The response content generation module is further configured to calculate the similarity between the vectorized representations and style category vectors of each candidate word in the preset vocabulary when the copy probability is less than the copy probability threshold, obtain the word similarity of each candidate word, and obtain the current response word corresponding to at least one style category based on the word similarity.
[0284] In one embodiment, the response content generation module is further configured to, when the response word prediction meets the conditions for continuing prediction, perform response word prediction based on the vectorized representation of each word, the style category vector, and the current response word, obtain the next response word corresponding to the current response word, use the next response word as the new current response word, and jump to the step of performing response word prediction based on the vectorized representation of each word, the style category vector, and the current response word to obtain the next response word corresponding to the current response word, until the response word prediction meets the prediction stopping condition, and obtain at least one predicted response word corresponding to at least one style category based on the current response word obtained from each response word prediction.
[0285] In one embodiment, the response content processing device further includes a sorting module. The sorting module is used to obtain the style distribution of the response content and the style distribution of the response content publishing object. Based on the style distribution of the response content and the style distribution of the response content publishing object, it obtains the style similarity of the response content. Based on the response content, the style distribution of the response content, the description content, the interactive content to be responded to, and the style distribution of the interactive content to be responded to, it performs interaction prediction to obtain the estimated interaction rate of the response content. Based on the style similarity of the response content and the estimated interaction rate, it sorts the response content to obtain the response content sorting result.
[0286] In one embodiment, the sorting module is further configured to encode the reply content and the style distribution corresponding to the reply content to obtain a first vector representation of the reply content, and to encode the style distribution of the interaction content to be replied to and the interaction content to be replied to to obtain a second vector representation of the interaction content to be replied to, and to encode the description content to obtain a vector representation of the description content, and to fuse the vector representation of the description content, the first vector representation, and the second vector representation to obtain the estimated interaction rate corresponding to the reply content.
[0287] In one embodiment, such as Figure 18 As shown, an interactive device for media content interaction is provided, including: an interactive content display module 1802, a response module 1804, an alternative response content display module 1806, and a response interactive content display module 1808, wherein:
[0288] Interactive content display module 1802 is used to display interactive content of media content;
[0289] The response module 1804 is used to respond to a trigger event of any interactive content and determine the interactive content to be replied to.
[0290] The alternative reply content display module 1806 is used to display alternative reply content corresponding to at least one style category of the interactive content to be replied to.
[0291] The reply content display module 1808 is used to respond to the selection event of any candidate reply content and display the candidate reply content selected by the selection operation as the interactive content to be replied to.
[0292] The aforementioned interactive device for media content displays interactive media content. In response to a trigger event for any interactive content, it determines the interactive content to be replied to. Corresponding to the interactive content to be replied to, it displays alternative reply content of at least one style category. It can provide alternative reply content of at least one style category as interactive content selection. In response to a selection event for any alternative reply content, it displays the alternative reply content selected by the selection operation as the interactive content to reply to the interactive content to be replied to, so that there is no need to input reply content to reply, thereby improving reply efficiency.
[0293] In one embodiment, the alternative response content display module is further configured to display alternative response content corresponding to at least one style category of the interactive content to be responded to, in descending order of the estimated interaction rate of the alternative response content.
[0294] In one embodiment, alternative response content corresponding to at least one style category is obtained through the above-described response content processing method.
[0295] Each module in the aforementioned response content processing device and media content interactive device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0296] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 19 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data such as preset dictionaries. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a response content processing method.
[0297] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 20 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an interactive method for media content. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0298] Those skilled in the art will understand that Figure 19 and Figure 20 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0299] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0300] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0301] In one embodiment, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above-described method embodiments.
[0302] It should be noted that the information (including but not limited to information about the target audience) and data (including but not limited to data used for analysis, data stored, and data displayed) involved in this application are all information and data authorized by the target audience or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0303] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0304] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0305] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.< / s> < / s>
Claims
1. A method for processing reply content, characterized in that, The method includes: Get the pending interactive content related to media content; The description of the media content and the interactive content to be replied to are encoded to obtain the target vector representations of the description and the interactive content to be replied to. Based on the target vector representation, a first style distribution is obtained; the first style distribution is used to describe the distribution of style categories of the interactive content to be replied to based on the description content. Style identification is performed based on the publishing object information of the interactive content to be replied to, resulting in a second style distribution; the second style distribution is used to describe the distribution of style categories of the publishing object of the interactive content to be replied to. Based on the first style distribution and the second style distribution, the style distribution of the interactive content to be responded to is obtained; Determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to; Based on the description content, the interactive content to be replied to, and the style category vector, generate reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
2. The method according to claim 1, characterized in that, The information of the publishing object includes at least one historical interaction of the publishing object; The process of style identification based on the publishing object information of the interactive content to be replied to, to obtain the second style distribution, includes: Perform style identification on the at least one historical interaction content to obtain the historical style distribution corresponding to the at least one historical interaction content; The second style distribution is obtained by weighted averaging the style category probabilities of each historical style category in the historical style distribution.
3. The method according to claim 1, characterized in that, The method further includes: From the interactive content of the media content, obtain similar interactive content to the interactive content to be responded to; Based on the description content, the interactive content to be replied to, and the style category vector, the system generates reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to. Based on the description content, the interactive content to be replied to, the similar interactive content, and the style category vector, generate reply content corresponding to at least one style category in the style distribution of the interactive content to be replied to.
4. The method according to claim 3, characterized in that, The step of obtaining similar interactive content to the interactive content to be responded to from the interactive content of the media content includes: The style similarity between the interactive content to be replied to and each interactive content of the media content is obtained, and the content similarity between the interactive content to be replied to and each interactive content is obtained. Based on the style similarity and content similarity of each interactive content, similar interactive content to the interactive content to be replied to is selected from the interactive content of the media content.
5. The method according to claim 3, characterized in that, The step of generating response content corresponding to at least one style category in the style distribution of the interaction content to be responded to, based on the description content, the interaction content to be responded to, the similar interaction content, and the style category vector, includes: The description content, the interactive content to be replied to, and the similar interactive content are segmented into words to obtain a set of segmented words. The words in the segmented word set are vectorized to obtain the vectorized representation of each word; Based on the vectorized representation of each word and the style category vector, the response word is predicted to obtain at least one predicted response word corresponding to at least one style category in the style distribution of the interactive content to be responded to. Based on at least one predicted response word corresponding to the at least one style category, generate response content corresponding to the at least one style category.
6. The method according to claim 5, characterized in that, The step of predicting response words based on the vectorized representations of each word and the style category vector, to obtain at least one predicted response word corresponding to at least one style category in the style distribution of the interactive content to be responded to, includes: Based on the vectorized representation of each word and the style category vector, the response word is predicted to obtain the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to; When the prediction of the response word meets the prediction stopping condition, the current response word is taken as at least one predicted response word corresponding to the at least one style category; When the prediction of the response word meets the conditions for continuing prediction, the prediction of the response word continues based on the vectorized representation of each word, the style category vector, and the current response word, so as to obtain at least one predicted response word corresponding to the at least one style category.
7. The method according to claim 6, characterized in that, The step of predicting response words based on the vectorized representations of each word and the style category vectors to obtain the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to includes: The similarity between the vectorized representation of each word and the style category vector is calculated to obtain the copy probability of each word. When there is at least one target word with a copy probability greater than or equal to the copy probability threshold, the current reply word corresponding to at least one style category in the style distribution of the interactive content to be replied to is obtained based on the at least one target word; When the copy probability is less than the copy probability threshold, the current response word corresponding to the at least one style category is obtained from the preset word list.
8. The method according to claim 7, characterized in that, The preset vocabulary includes a vectorized representation of at least one candidate word; When the copy probability is less than the copy probability threshold, obtaining the current response word corresponding to the at least one style category from the preset word list includes: When the copy probability is less than the copy probability threshold, the similarity is calculated between the vectorized representation of each candidate word in the preset word list and the style category vector to obtain the word similarity of each candidate word. Based on the word similarity, the current response word corresponding to the at least one style category is obtained.
9. The method according to claim 6, characterized in that, When the predicted response word meets the conditions for continuing prediction, the process of continuing to predict response words based on the vectorized representations of each word, the style category vector, and the current response word, to obtain at least one predicted response word corresponding to the at least one style category, includes: When the prediction of the response word meets the conditions for continuing prediction, the next response word is predicted based on the vectorized representation of each word, the style category vector, and the current response word. The next reply word is used as the new current reply word, and the process jumps to the step of predicting the next reply word based on the vectorized representation of each word, the style category vector, and the current reply word to obtain the next reply word corresponding to the current reply word. Until the prediction of the response word meets the prediction stopping condition, at least one predicted response word corresponding to the at least one style category is obtained based on the current response word obtained from each prediction of the response word.
10. The method according to any one of claims 1-9, characterized in that, After generating response content corresponding to at least one style category in the style distribution of the interaction content to be responded to based on the description content, the interaction content to be responded to, and the style category vector, the method further includes: Obtain the style distribution corresponding to the reply content and the style distribution of the reply content's publishing object; Based on the style distribution of the reply content and the style distribution of the object that published the reply content, the style similarity of the reply content is obtained. Interaction prediction is performed based on the response content, the corresponding style distribution of the response content, the description content, the interaction content to be responded to, and the style distribution of the interaction content to be responded to, to obtain the estimated interaction rate corresponding to the response content; The responses are sorted based on their style similarity and estimated interaction rate to obtain a ranking result.
11. The method according to claim 10, characterized in that, The step of predicting the interaction rate based on the response content, the corresponding style distribution of the response content, the description content, the interaction content to be responded to, and the style distribution of the interaction content to be responded to, to obtain the estimated interaction rate corresponding to the response content, includes: The reply content and its corresponding style distribution are encoded to obtain a first vector representation of the reply content, and the interaction content to be replied to and its style distribution are encoded to obtain a second vector representation of the interaction content to be replied to. The description content is encoded to obtain a vector representation of the description content; The vector representations of the description content, the first vector representation, and the second vector representation are fused to obtain the estimated interaction rate of the response content.
12. A method for interactive media content, characterized in that, The method includes: Interactive content that displays media content; In response to a trigger event for any of the said interactive content, determine the interactive content to be replied to; Corresponding to the interactive content to be replied to, alternative reply content of at least one style category of the interactive content to be replied to is displayed; the alternative reply content of at least one style category is obtained by the reply content processing method described in any one of claims 1-11 above; In response to a selection event for any of the candidate response content, the candidate response content selected by the selection event is displayed as the interactive content to be replied to.
13. The method according to claim 12, characterized in that, The provision of alternative response content corresponding to at least one style category of the interactive content to be responded to, including: Corresponding to the interactive content to be replied to, alternative response content corresponding to at least one style category of the interactive content to be replied to is displayed in descending order of the estimated interaction rate of the alternative response content.
14. A response content processing device, characterized in that, The device includes: The interactive content acquisition module is used to acquire interactive content that needs to be responded to in relation to media content; A style recognition module is used to encode the description content of the media content and the interactive content to be replied to, to obtain target vector representations of the description content and the interactive content to be replied to, and to obtain a first style distribution based on the target vector representations; the first style distribution is used to describe the distribution of style categories of the interactive content to be replied to based on the description content, and to perform style recognition based on the publishing object information of the interactive content to be replied to, to obtain a second style distribution; the second style distribution is used to describe the distribution of style categories of the publishing object of the interactive content to be replied to, and to obtain the style distribution of the interactive content to be replied to based on the first style distribution and the second style distribution; The style category vector determination module is used to determine the style category vector corresponding to at least one style category in the style distribution of the interactive content to be replied to; The response content generation module is used to generate response content corresponding to at least one style category in the style distribution of the interaction content to be responded to, based on the description content, the interaction content to be responded to, and the style category vector.
15. The apparatus according to claim 14, characterized in that, The information of the publishing object includes at least one historical interaction content of the publishing object; the style recognition module is further used to perform style recognition on the at least one historical interaction content to obtain a historical style distribution corresponding to the at least one historical interaction content, and to perform a weighted average of the style category probabilities of each historical style category in the historical style distribution to obtain a second style distribution.
16. The apparatus according to claim 14, characterized in that, The response content generation module is further configured to obtain similar interactive content of the interactive content to be responded to from the interactive content of the media content, and generate response content corresponding to at least one style category in the style distribution of the interactive content to be responded to based on the description content, the interactive content to be responded to, the similar interactive content, and the style category vector.
17. The apparatus according to claim 16, characterized in that, The response content generation module is further configured to obtain the style similarity between the interactive content to be responded to and each interactive content of the media content, and to obtain the content similarity between the interactive content to be responded to and each interactive content. Based on the corresponding style similarity and content similarity of each interactive content, similar interactive content to the interactive content to be responded to is selected from the interactive content of the media content.
18. The apparatus according to claim 16, characterized in that, The response content generation module is further configured to segment the description content, the interactive content to be responded to, and the similar interactive content into words to obtain a segmented word set; to vectorize each word in the segmented word set to obtain a vectorized representation of each word; to predict response words based on the vectorized representations of each word and the style category vector; to obtain at least one predicted response word corresponding to at least one style category in the style distribution of the interactive content to be responded to; and to generate response content corresponding to the at least one style category based on the at least one predicted response word corresponding to the at least one style category.
19. The apparatus according to claim 18, characterized in that, The response content generation module is further configured to predict response words based on the vectorized representations of each word and the style category vector, to obtain the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to. When the response word prediction meets the prediction stopping condition, the current response word is used as at least one predicted response word corresponding to the at least one style category. When the response word prediction meets the prediction continuing condition, the module continues to predict response words based on the vectorized representations of each word, the style category vector, and the current response word, to obtain at least one predicted response word corresponding to the at least one style category.
20. The apparatus according to claim 19, characterized in that, The response content generation module is further used to perform similarity calculation on the vectorized representation of each word and the style category vector to obtain the copy probability of each word. When there is at least one target word with a copy probability greater than or equal to the copy probability threshold, the current response word corresponding to at least one style category in the style distribution of the interactive content to be responded to is obtained based on the at least one target word. When the copy probabilities are all less than the copy probability threshold, the current response word corresponding to the at least one style category is obtained from the preset word list.
21. The apparatus according to claim 20, characterized in that, The preset word list includes vectorized representations of at least one candidate word. The response content generation module is further configured to calculate the similarity between the vectorized representations of each candidate word in the preset word list and the style category vector when the copy probability is less than the copy probability threshold, to obtain the word similarity of each candidate word, and to obtain the current response word corresponding to the at least one style category based on the word similarity.
22. The apparatus according to claim 19, characterized in that, The reply content generation module is further configured to, when the reply word prediction meets the conditions for continuing prediction, perform reply word prediction based on the vectorized representation of each word, the style category vector, and the current reply word to obtain the next reply word corresponding to the current reply word, use the next reply word as the new current reply word, and jump to the step of performing reply word prediction based on the vectorized representation of each word, the style category vector, and the current reply word to obtain the next reply word corresponding to the current reply word, until the reply word prediction meets the prediction stopping condition, and obtain at least one predicted reply word corresponding to the at least one style category based on the current reply word obtained from each reply word prediction.
23. The apparatus according to any one of claims 14-22, characterized in that, The device further includes a sorting module, which is used to obtain the style distribution corresponding to the reply content and the style distribution of the reply content's publishing object; based on the style distribution corresponding to the reply content and the style distribution of the reply content's publishing object, obtain the style similarity corresponding to the reply content; perform interaction prediction based on the reply content, the style distribution corresponding to the reply content, the description content, the interaction content to be replied to, and the style distribution of the interaction content to be replied to, obtain the estimated interaction rate corresponding to the reply content; and sort the reply content according to the style similarity corresponding to the reply content and the estimated interaction rate to obtain the reply content sorting result.
24. The apparatus according to claim 23, characterized in that, The sorting module is further configured to encode the reply content and its corresponding style distribution to obtain a first vector representation of the reply content, encode the interaction content to be replied to and its style distribution to obtain a second vector representation of the interaction content to be replied to, encode the description content to obtain a vector representation of the description content, and fuse the vector representation of the description content, the first vector representation, and the second vector representation to obtain the estimated interaction rate of the reply content.
25. An interactive device for media content interaction, characterized in that, The device includes: The interactive content display module is used to display interactive content from media content. The response module is used to determine the interactive content to be replied to in response to any of the interactive content trigger events; The alternative response content display module is used to display alternative response content corresponding to at least one style category of the interactive content to be responded to, corresponding to the interactive content to be responded to; the alternative response content corresponding to the at least one style category is obtained by the response content processing method described in any one of claims 1-11. The reply content display module is used to respond to a selection event of any of the candidate reply content, and display the candidate reply content selected by the selection event as the interactive content to reply to the interactive content to be replied to.
26. The apparatus according to claim 25, characterized in that, The alternative response content display module is also used to display alternative response content corresponding to at least one style category of the interactive content to be responded to, in descending order of the estimated interaction rate of the alternative response content.
27. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 13.
28. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.
29. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Information recommendation method and device and electronic equipment
CN111984767A
Reply information processing method and system
CN112434140A
Recommendation information generation method and device, storage medium and computing equipment
CN114595671A