A comment generation method and device, electronic equipment and storage medium

By segmenting and extracting features from video content on social media platforms, and combining this with user descriptions, a trained model is used to generate personalized comments. This solves the problem of inaccurate comments in existing systems and improves comment acceptance rates and user interaction efficiency.

CN114970494BActive Publication Date: 2025-12-30TENCENT TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110211022.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-25
Publication Date
2025-12-30
Estimated Expiration
2041-04-09

AI Technical Summary

Technical Problem

The lack of personalization in video comments generated by existing social media platforms results in a low rate of comment adoption.

Method used

By segmenting and extracting features from the target multimedia content, and combining the descriptive information of the target object with a pre-set set of candidate object descriptive features, a trained comment generation model is used to predict personalized comments.

Benefits of technology

It improved the accuracy and acceptance rate of comments, and enhanced the auxiliary effect of the automatic comment function on user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114970494B_ABST
    Figure CN114970494B_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field, in particular to a comment generation method and device, electronic equipment and storage medium, to improve the comment generation efficiency and accuracy. Wherein, the method comprises: in response to a comment request triggered by a target object for target multimedia content, performing word segmentation processing on first description information of the target multimedia content to obtain each word segmentation in the first description information; and obtaining a target description feature of the target object based on second description information of the target object and a preset candidate object description feature set; and predicting comment information to be published by the target object for the target multimedia content based on the target description feature and each word segmentation. The application simultaneously learns the personalized representation of the object when generating the comment, and simultaneously uses the target multimedia content and the personalized description information of the object when providing the automatic comment function for the user, so that the comment is quickly and accurately generated, and the personalized demand of the automatic comment for the current object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly to a comment generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of internet technology, various social application platforms have emerged, including various instant messaging applications and content sharing platforms. These existing social application platforms all have social functions. Users can publish articles, videos, pictures and other multimedia content on the social application platforms. Platform contacts can forward, like, and comment on each other's posts, interact and discuss the published content, and post comments.

[0003] Taking video comments as an example, when generating video comments for users in related technologies, it is generally achieved by copying comments from similar videos or generating comments based on video content. The generated comments are basically the same for all users, resulting in inaccurate comments and a low acceptance rate. Summary of the Invention

[0004] This application provides a comment generation method, apparatus, electronic device, and storage medium to improve the efficiency and accuracy of comment generation.

[0005] This application provides a comment generation method, including:

[0006] In response to a comment request triggered by a target object regarding target multimedia content, the first description information of the target multimedia content is segmented to obtain each segmented word in the first description information; and

[0007] Based on the second description information of the target object and a preset set of candidate object description features, target description features for the target object are obtained;

[0008] Based on the target description features and the word segmentation, predict the comment information that the target object will publish in response to the target multimedia content.

[0009] This application provides a comment generation device, comprising:

[0010] A word segmentation processing unit is configured to, in response to a comment request triggered by a target object for target multimedia content, perform word segmentation processing on the first description information of the target multimedia content to obtain each word in the first description information; and

[0011] The feature extraction unit is used to obtain target description features for the target object based on the second description information of the target object and a preset set of candidate object description features;

[0012] The prediction unit is used to predict the comment information that the target object will publish in relation to the target multimedia content based on the target description features and the various word segments.

[0013] Optionally, the word segmentation processing unit is specifically used for:

[0014] The first description information of the target multimedia content is input into the trained comment generation model;

[0015] The first description information is segmented and encoded based on the encoding part of the trained comment generation model to obtain each segment of the first description information and the word vector of each segment.

[0016] Optionally, the prediction unit is specifically used for:

[0017] The word vectors of each segmented word and the target description features are input into the decoding part of the trained comment generation model. Based on the decoding part, decoding processing is performed to obtain the comment information output by the trained comment generation model.

[0018] The trained comment generation model is trained on a training sample dataset. Each training sample in the training sample dataset includes an object profile of the sample object, descriptive information of the sample multimedia content, and real comments published by the sample object on the sample multimedia content.

[0019] Optionally, if the comment information includes multiple comment terms, the prediction unit is specifically used for:

[0020] The system uses an iterative loop to generate each comment term from the comment information sequentially. During each iteration, the following operations are performed:

[0021] The previously output comment word is input into the decoding part, wherein the first input into the decoding part is a pre-set start marker word;

[0022] Based on the previously output comment words, the word vectors of each word segment, and the target description features, the current output comment words are generated through decoding.

[0023] Optionally, the device further includes:

[0024] The model training unit is used to train the trained comment generation model in the following ways:

[0025] Select training samples from the training sample dataset;

[0026] The untrained comment generation model is iteratively trained using the training samples to obtain the trained comment generation model. Each training iteration includes the following operations:

[0027] The object profiles of the sample objects in the training samples and the descriptive information of the sample multimedia content are input into the untrained comment generation model to obtain the predicted comments output by the untrained comment generation model.

[0028] Based on the error between the predicted comments and the real comments in the corresponding training samples, the parameters of the untrained comment generation model are adjusted.

[0029] Optionally, the model training unit is specifically used for:

[0030] Alternatively, select training samples from the training sample dataset whose number of comments on the sample objects reaches a third preset threshold, or select training samples whose number of comments on the multimedia content of the sample objects reaches a fourth preset threshold.

[0031] An electronic device provided in this application includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor performs the steps of any of the above-described comment generation methods.

[0032] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of any of the above-described comment generation methods.

[0033] This application provides a computer-readable storage medium including program code. When the program product is run on an electronic device, the program code is used to cause the electronic device to perform the steps of any of the above-described comment generation methods.

[0034] The beneficial effects of this application are as follows:

[0035] This application provides a comment generation method, apparatus, electronic device, and storage medium. Because this application learns the personalized representation of the object while generating comments, when providing automated commenting functions to users, it uses both the target multimedia content and the personalized description information of the object to generate comments quickly and accurately, improving the personalized needs of the current object for automated comments, further increasing the object adoption rate of automated comments, and enhancing the auxiliary effect of automated commenting functions on object interaction.

[0036] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0037] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0038] Figure 1 This is an optional schematic diagram of an application scenario in an embodiment of this application;

[0039] Figure 2 This is a flowchart illustrating a comment generation method according to an embodiment of this application;

[0040] Figure 3 This is a flowchart of a method for generating personalized automatic comments in an embodiment of this application;

[0041] Figure 4 This is a flowchart illustrating the construction process of a user-personalized comment generation model in an embodiment of this application.

[0042] Figure 5 This is a schematic diagram of the structure of a comment generation model in an embodiment of this application;

[0043] Figure 6 This is a flowchart illustrating an optional comment generation model training method in an embodiment of this application.

[0044] Figure 7 This application provides a method for obtaining user vectors for users with a small number of comments.

[0045] Figure 8 This is a schematic diagram of the composition structure of a comment generation device according to an embodiment of this application;

[0046] Figure 9 This is a schematic diagram of the composition structure of an electronic device according to an embodiment of this application;

[0047] Figure 10 This is a schematic diagram of the composition structure of another electronic device using an embodiment of this application. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0049] The following describes some of the concepts involved in the embodiments of this application.

[0050] Multimedia content: A human-computer interactive information exchange and dissemination medium that combines two or more media. Media includes text, images, sound, video, etc. In the embodiments of this application, multimedia content can be articles, news, videos, music, etc.

[0051] Description Information: In this embodiment, the description information is divided into description information of multimedia content and description information of objects. The description information of multimedia content mainly describes the attributes of the multimedia content. Taking video as an example, the description information of video mainly refers to the video text content, including title text, dialogue text recognized by Automatic Speech Recognition (ASR), and subtitle text recognized by Optical Character Recognition (OCR). The description information of objects includes one or more of the object profile and the object identifier. The object mainly refers to something used for [specific purposes], and the object profile can be a user profile. In this embodiment, both the first and second description information refer to description information. The first and second are used for distinction; specifically, the first description information is the description information of the multimedia content, and the second description information is the description information of the object.

[0052] Candidate object descriptive feature set: This is a feature set pre-constructed in this embodiment of the application. This feature set contains descriptive features of many candidate objects. These descriptive features can also be in the form of deep vector representations and can be obtained based on machine learning. Here, candidate objects refer to sample objects.

[0053] User personas, also known as user roles, are an effective tool for outlining target users and connecting user needs with design direction. They primarily include basic information such as the user's age, gender, occupation, and behavioral preferences. User personas are widely used across various fields. In practice, they typically connect user attributes, behaviors, and expectations using the simplest and most relatable language. As virtual representatives of actual users, user personas are constructed based on users' multimedia activity data and preferences over a period of time. The resulting user personas need to represent the main audience and target group of the product. In this embodiment, the user personas of both candidate and target objects can be represented as a set of information tag weights. These information tags primarily describe the user's interests and hobbies, and can also be called interest tags. This set contains at least one interest tag and the corresponding weight for each interest tag.

[0054] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0055] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0056] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0057] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0058] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0059] The solutions provided in the embodiments of this application involve technologies such as artificial intelligence, natural language processing, and machine learning. The method for training a comment generation model proposed in the embodiments of this application can be divided into two parts: a training part and an application part. The training part involves the technical field of machine learning. In the training part, the comment generation model is trained using machine learning technology. The training samples provided in the embodiments of this application, which include sample objects and sample multimedia content, are used to train the comment generation model. After the training samples are processed by the comment generation model, the predicted comments output by the model are obtained. Combined with the real comments labeled in the training samples, the model parameters are continuously adjusted through optimization algorithms. The application part is used to generate comments for the target object on the target multimedia content to be published using the comment generation model trained in the training part. Furthermore, it should be noted that the comment generation model in the embodiments of this application can be trained online or offline, without specific limitations.

[0060] The design concept of the embodiments of this application is briefly introduced below:

[0061] With the rapid development of information technology and the internet, multimedia content that allows readers or viewers to post comments, such as online news, audio and video, short videos, e-books, online articles, and forum posts, is becoming increasingly popular and a major way for people to obtain information in their daily lives. People can access and browse various multimedia content presented in the form of pictures, text, or videos through major online portals, large news websites, or short video applications (APPs).

[0062] Taking video as an example of multimedia content, when generating video comments for users in related technologies, it is generally achieved by copying comments from similar videos or generating comments based on video content. The generated comments are basically the same for all users, resulting in inaccurate comments and a low acceptance rate.

[0063] In view of this, embodiments of this application propose a comment generation method, apparatus, electronic device, and storage medium. Since this application learns the user's personalized representation while generating comments, when providing automated comment functions to users, it uses target multimedia content and the user's personalized description information to generate comments quickly and accurately, improving the personalized needs of the current user for automated comments, further increasing the user adoption rate of automated comments, and enhancing the auxiliary effect of automated comment functions on user interaction.

[0064] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and descriptive features in the embodiments of this application can be combined with each other without conflict.

[0065] like Figure 1 The diagram illustrates an application scenario of an embodiment of this application. The application scenario diagram includes two terminal devices 110 and one server 120. Terminal devices 110 are equipped with a client, which can be accessed by logging into the client through the terminal device 110. The client involved in this embodiment can be software, a webpage, a mini-program, etc., and the server 120 is the backend server corresponding to the software, webpage, mini-program, etc., without limiting the specific type of client.

[0066] It should be noted that, Figure 1 The examples shown are merely illustrative; in reality, the number of terminal devices and servers is unlimited and is not specifically limited in the embodiments of this application.

[0067] The terminal device 110 and the server 120 can communicate via a communication network. In one optional embodiment, the communication network is a wired network or a wireless network. The terminal device 110 and the server 120 can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0068] In this embodiment, the terminal device 110 is an electronic device used by a user. This electronic device can be a personal computer, mobile phone, tablet computer, laptop, e-book reader, smart TV, smart home device, or other computer device with certain computing capabilities that runs instant messaging software and websites or social networking software and websites. The server 120 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0069] The comment generation model can be deployed on server 120 for training. Server 120 can store a large number of training samples for training the comment generation model. Optionally, after training the comment generation model based on the training method in this embodiment, the trained comment generation model can be directly deployed on server 120 or terminal device 110. In this embodiment, the comment generation model can be deployed on server 120. In this embodiment, the comment generation model is mainly used to automatically generate personalized comments for the target object when the target object intends to comment on the target multimedia content.

[0070] In one possible application scenario, the training samples in this application can be stored using cloud storage technology. Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to aggregate a large number of storage devices of various types (storage devices are also called storage nodes) in a network to work together through application software or application interfaces to jointly provide data storage and business access functions.

[0071] In one possible application scenario, to reduce communication latency, servers 120 can be deployed in various regions, or for load balancing, different servers 120 can serve the regions corresponding to each terminal backup 10. Multiple servers 120 can share data through blockchain, forming a data-sharing system. For example, terminal device 110 located at location a communicates with server 120, while terminal device 110 located at location b communicates with other servers 120.

[0072] Each server 120 in the data sharing system has a corresponding node identifier. Each server 120 can also store the node identifiers of other servers 120 in the data sharing system, so that the generated blocks can be broadcast to the other servers 120 in the data sharing system based on their node identifiers. Each server 120 can maintain a node identifier list as shown in the table below, storing the server 120 name and node identifier in the list. The node identifier can be an IP (Internet Protocol) address or any other information that can be used to identify the node. Table 1 only uses IP addresses as an example.

[0073] Table 1

[0074] Server Name Node identifier Node 1 119.115.151.174 Node 2 118.116.189.145 … … Node N 119.124.789.258

[0075] The following describes the comment generation method provided by the exemplary embodiments of this application in conjunction with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this respect.

[0076] It should be noted that the comment generation method in this application embodiment can be executed by the server or the terminal device alone, or by the server and the terminal device together.

[0077] See Figure 2 The diagram shown is a flowchart illustrating an implementation of a comment generation method provided in this application, specifically executed by a terminal device. The specific implementation process of this method is as follows:

[0078] S21: In response to a comment request triggered by a target object regarding target multimedia content, the terminal device performs word segmentation processing on the first description information of the target multimedia content to obtain each word in the first description information; and

[0079] S22: The terminal device obtains target description features for the target object based on the second description information of the target object and a preset set of candidate object description features;

[0080] Multimedia content refers to online information, audio and video, short videos, e-books, online articles, and forum posts that allow readers or viewers to post comments. This application primarily uses video as an example for detailed explanation.

[0081] In this embodiment, a comment request is triggered when the target object intends to comment on the target multimedia content. For example, a user clicks the control to send comments while watching a video, or clicks the comment control while watching a short video. At this time, a personalized comment for the target multimedia content can be generated based on the automatic comment generation method in this embodiment. Automatic comment generation improves the speed at which the target object inputs comments, reduces the cost of inputting comments, and enhances the interaction efficiency and experience for the target object.

[0082] Specifically, the first step is to extract features from the target object and segment the target multimedia content into words. Further, a comment is generated in step S23.

[0083] S23: The terminal device predicts the comment information that the target object will publish for the target multimedia content based on the target description features and each word segmentation.

[0084] Because this application embodiment learns the user's personalized representation while generating comments, when providing the user with automated commenting functionality, it uses both target multimedia content and the user's personalized description information to quickly and accurately generate comments, improving the personalized needs of the current user for automated comments, further increasing the user adoption rate of automated comments, and enhancing the auxiliary effect of automated commenting functionality on user interaction.

[0085] It should be noted that, in this embodiment, a candidate object description feature set needs to be pre-constructed. This feature set contains description features of many candidate objects, and these description features can also be in the form of deep vector representations, which can be obtained based on machine learning. Here, candidate objects refer to sample objects.

[0086] In one alternative implementation, the candidate objects in this application embodiment are all those that have posted comments a certain threshold number of times, i.e., high-frequency comment objects.

[0087] In one optional implementation, step S22 can be divided into the following two cases:

[0088] Scenario 1: The target object is one of the candidate objects.

[0089] In this case, the second description information of the target object includes the object identifier of the target object, i.e., the user ID. At this time, the description feature corresponding to the ID can be queried from the candidate object description feature set based on the user ID, which is the target description feature of the target object.

[0090] Scenario 2: The target object is not one of the candidate objects.

[0091] When the target object posts few comments or has not posted any comments, the corresponding number of comments will be less than the set threshold, that is, the target object is a low-frequency comment object. At this time, some similar objects can be selected from the preset candidate object description feature set, and the target description features of the target object can be further determined based on the description features of these similar objects.

[0092] Specifically, when searching for similar objects, the similarity between each candidate object and the target object is determined based on the object profile in the second descriptive information, and then the selection is based on the similarity. Finally, the descriptive features of at least one candidate object with a similarity reaching a first preset threshold are obtained from the candidate object descriptive feature set, and the target descriptive features of the target object are determined based on the obtained descriptive features of at least one candidate object.

[0093] In this embodiment of the application, considering that the number of candidate objects whose similarity reaches the first preset threshold is uncertain, when determining the target description features of the target object based on the description features of at least one candidate object, it can also be divided into the following two methods:

[0094] Method 1: If the number of at least one candidate object found is less than the set number, the corresponding similarity is used as a weight to perform a weighted summation of the descriptive features of the at least one candidate object found, and the target descriptive features are obtained.

[0095] For example, if the quantity is set to 5, and assuming there are 4 candidate objects found, with corresponding descriptive feature vectors D1, D2, D3, and D4, and the corresponding weights representing the similarity between the candidate objects and the target object, assuming they are s1, s2, s3, and s4 respectively, then the target descriptive feature of the target object is D = s1·D1 + s2·D2 + s3·D3 + s4·D4.

[0096] Method 2: If at least one candidate object is found, the candidate objects that meet the set number are selected based on the corresponding similarity. The corresponding similarity is used as a weight to perform a weighted summation of the descriptive features of each selected candidate object to obtain the target descriptive features.

[0097] For example, if the number is set to 5, and there are 8 candidate objects found, then we can sort them by similarity and select the top 5 candidate objects with the highest similarity. Let's assume that the vectors of the corresponding descriptive features are D1, D2, D3, D4 and D5, and the corresponding weights are the similarity between the candidate objects and the target object, let's assume they are s1, s2, s3, s4 and s5, then the target descriptive feature of the target object is D = s1·D1 + s2·D2 + s3·D3 + s4·D4 + s5·D5.

[0098] In this embodiment, high-frequency commenters can directly query their user vector representation based on their user ID, while low-frequency commenters need to find similar high-frequency commenters and then calculate their vector representation based on the similar high-frequency user representations.

[0099] It should be noted that the candidate object description feature set in this application embodiment is also called the user personalized vector table. This table only stores the vector representations of high-frequency users. When querying high-frequency similar users for low-frequency comment users, the vector representation of the current low-frequency user is obtained by comparing the profile similarity with high-frequency users and weighting and summing the vectors of at least one user whose similarity meets a certain threshold. By using video content and user personalized representations simultaneously to generate comments, the automatically generated comments are more in line with the current user's personalized needs, achieving the effect of customizing automatic comments for users.

[0100] In one optional implementation, the similarity between two objects is obtained based on the similarity of their object profiles. The object profile includes at least one information tag and a tag weight corresponding to that information tag. The information tag is obtained based on the object's historical behavior analysis and can be used to represent the object's interests; therefore, it can also be called an interest tag.

[0101] For example, user A's user profile is {(Liu Moumou, 0.231), (Funny Jokes, 0.226), ..., (Baking, 0.097)}, where the values ​​are the user's preference weights for these interest tags. The user profile is constructed through iterative learning based on the user's historical playback behavior.

[0102] In this embodiment, the similarity between two users is the sum of the weighted overlap of their interest tags. Specifically, for any candidate object, the object profile of the candidate object is first obtained. Then, each interest tag in the object profile of the target object is compared with each interest tag in the object profile of the candidate object to determine the overlap between each pair of interest tags. Finally, the sum of the products of the overlap of interest tags and their corresponding weights is taken as the similarity between the target object and the candidate object.

[0103] The text similarity between any two interest tags can be considered as the overlap between them. The weight of each pair of interest tags is determined based on the label weights of each individual interest tag within that pair. For example, if the overlap between interest tag A in the target object and interest tag B in the candidate object is 0.8 (the overlap can range from 0 to 1), and the label weight for interest tag A is 'a', while the label weight for interest tag B is 'b', then the weight corresponding to the overlap between these two interest tags is (a+b) / 2. Finally, the product of the overlap between these two interest tags and their corresponding weights, x1 = 1 * (a+b) / 2 = (a+b) / 2.

[0104] Let xi represent the product of the overlap between two interest tags and their corresponding weights (i = 1, 2, ..., 6). The similarity between the target object and the candidate object is the sum of the products of the overlap between each pair of interest tags and their corresponding weights, i.e., x1 + x2 + ... + x6.

[0105] Of course, to improve the calculation speed, we can also directly compare whether two interest tags are completely identical. If they are completely identical, the corresponding overlap is 1; otherwise, it is 0. In fact, there are many ways to calculate this, and we will not make any specific restrictions here.

[0106] In the above implementation, the similarity between users is calculated by user profile similarity calculation, where user profile is a set of user interest tag weights.

[0107] It should be noted that the comment generation method in this application embodiment can also be implemented based on machine learning. Specifically, the first description information of the target multimedia content is input into the trained comment generation model; the first description information is segmented and encoded based on the encoding part of the trained comment generation model to obtain each segment of the first description information and the word vector of each segment; then the word vector of each segment and the target description features are input into the decoding part of the trained comment generation model, and the decoding part is used for decoding to obtain the comment information output by the trained comment generation model.

[0108] In this embodiment, the trained comment generation model is trained based on a training sample dataset. Each training sample in the training sample dataset includes an object profile of the sample object, descriptive information of the sample multimedia content, and real comments published by the sample object on the sample multimedia content.

[0109] Taking multimedia content, specifically video, as an example, please refer to... Figure 3The diagram shows a flowchart of a user-personalized automatic comment generation method according to an embodiment of this application. In this embodiment, a user-personalized comment generation model is trained based on a video library and a user comment library. The video library stores sample multimedia content and its description, while the user comment library stores real comments posted by sample objects on sample multimedia content, along with the description information of the sample multimedia content. Furthermore, user vectors can be derived through model training to construct a user vector library, which is the set of candidate object descriptive features listed in this embodiment.

[0110] When a user intends to comment on a video, such as when a user clicks the comment button on the client, a user vector can be obtained based on the user profile. This user vector is then used to generate a personalized comment based on the video content and the user vector. The user profile is stored in a user profile database, and the corresponding profile information can be retrieved from the database based on the user ID.

[0111] In this embodiment, comments that match the user's personality are automatically generated based on the current video content and user information, improving the user's ability to interact directly with the automatic comments and enhancing the comment interaction experience.

[0112] See Figure 4 The diagram illustrates a flowchart of the construction process for a user-personalized comment generation model in this embodiment of the application. Based on a large amount of user video comment data from the platform, an automatic user-personalized comment generation model is constructed. The trained user deep representation is exported to a user vector representation library for later use in generating video comments based on user-personalized features. Specifically, the model inputs video descriptions from the video library and user descriptions from the user comment library into the comment generation model. The model outputs the user's comment on the video, compares the obtained comment with the corresponding real comments from the user in the user comment library, adjusts the model parameters, and iteratively trains to obtain a well-trained comment generation model.

[0113] The comment generation model in the embodiments of this application will be described in detail below:

[0114] See Figure 5 The diagram shown is a structural schematic of a comment generation model according to an embodiment of this application. This model is an automatic comment generation model built upon the Transformer Encoder-Decoder model, specifically comprising two parts: an encoding part and a decoding part.

[0115] In this embodiment, the training data for the user-personalized comment generation model is shown below. To ensure the automatic generation model for user-personalized comments learns the user vector representation more thoroughly, users with more comments than the UC (i.e., the third preset threshold) and videos with more comments than the VC (i.e., the fourth preset threshold) are selected. Only comments from these users and videos are retained in the training data. The generated user vector table also only contains users who post high-frequency comments. For example, users who have posted more than 1000 comments and videos that have received more than 1500 comments are selected. The third preset threshold and the fourth preset threshold can be the same or different, and no specific limitation is made here.

[0116] Within this framework, different users can post different comments on the same video text, and the same user can post different comments on different video texts. Therefore, the training sample data structure could be: Video 1 text content, User ID1, User comment 1; Video 1 text content, User ID2, User comment 2; Video 2 text content, User ID1, User comment 3; Video 2 text content, User ID3, User comment 4; Video 2 text content, User ID4, User comment 5; ...; Video v text content, User IDu, User comment c.

[0117] Each training sample includes a video description, a user description, and a real comment posted by that user on the video. The video text content identifies the video description, the user ID represents the user description, and the user description may further include user profile information, or can be directly retrieved from a user profile database based on the user ID.

[0118] In this embodiment, the input features of the Encoder part of the user-personalized automatic comment generation model are video text content, including title text, dialogue text based on ASR recognition, and subtitle text based on OCR recognition. Considering that the dialogue text and subtitle text may be long, key information can be retained by extracting keywords. The video text content is converted into word vector representation through word segmentation and ID-based query vectors, for example... Figure 5 The words shown are 1, 2, ..., n. Then, the Transformer Encoder constructs a deep representation of the video text, i.e. Figure 5 The word 1 represents, word 2 represents, ..., word n ​​represents.

[0119] In one alternative implementation, if the final model outputs comment information that includes multiple comment words, then based on the word vectors of each segmented word (i.e., ...) in the decoding part... Figure 5 The terms shown (word 1 represents, word 2 represents...) and target description features (i.e. Figure 5When decoding the user representation (in the model) to obtain the comment information output by the trained comment generation model, an iterative loop is needed to generate each comment word in the comment information sequentially. During each iteration, the following operations are performed:

[0120] First, the previously output comment words need to be re-input into the decoding part. Then, based on the previously output comment words, the word vectors of each word segment, and the target description features, the decoding process is performed to generate the comment words for the current output.

[0121] The first input for decoding is a pre-set start marker word, for example... Figure 5 As shown, when generating comment word 1, the corresponding input starting marker word is <s>Then, when generating comment word 2, the corresponding input is comment word 1, ..., when generating comment word n, the corresponding input is comment word n-1.

[0122] In other words, when the Decoder part of the model generates each comment word, it takes the comment word generated in the previous step as input, and uses the current user vector and the word vectors of each segment as input features of the model. It queries the user vector representation through the current user ID, and calculates whether to copy words from the original video text or select words from the word list when generating specific words based on the copy mechanism.

[0123] During the training phase, the loss is calculated by comparing the generated comment words with the real comments in the training samples, and the model parameters and user vector representation are updated through error backpropagation. This approach ensures that the model-generated comments not only match the current video content but also meet the user's personalized needs.

[0124] See Figure 6 The diagram shown illustrates a training method for a comment generation model according to an embodiment of this application. This method can be executed by the server or the terminal device alone, or by both the server and the terminal device. Here, we illustrate the method as being executed by the server. The specific implementation flow of this method is as follows:

[0125] Step S601: The server selects training samples from the training sample dataset;

[0126] Step S602: The server inputs the object profiles of the sample objects in the training samples and the descriptive information of the sample multimedia content into the untrained comment generation model;

[0127] Step S603: The server obtains the predicted comments output by the untrained comment generation model;

[0128] Step S604: The server adjusts the parameters of the untrained comment generation model based on the error between the predicted comments and the real comments in the corresponding training samples;

[0129] Step S605: The server determines whether the comment generation model after parameter adjustment has converged. If it has, proceed to step S606; otherwise, return to step S601.

[0130] Step S606: The server uses the comment generation model obtained after this parameter adjustment as the trained comment generation model.

[0131] In step S601, when selecting training samples from the training sample dataset, specifically, training samples whose included sample objects have been commented on a number of times reaching a third preset threshold are selected, or training samples whose included multimedia content has been commented on a number of times reaching a fourth preset threshold are selected. Alternatively, when constructing the training sample dataset, only training samples that meet the above conditions can be selected. In this way, selection from the training sample dataset can be random and does not need to refer to the above conditions.

[0132] When using the user- and video-based personalized comment generation model constructed above to generate personalized comments for users, the video text content is first obtained, and then input into the model to construct a deep representation of the video text according to the model's Encoder format requirements. The Decoder part of the model then generates personalized automatic comments step by step based on the current user vector and the vocabulary generated in the previous step.

[0133] See Figure 7 As shown in the illustration, this is one method for obtaining user vectors for users with a low comment volume, as listed in the embodiments of this application. When querying high-frequency similar users for low-frequency comment users, the similarity between the user's profile and the high-frequency users is compared. The vectors of the top k users whose similarity meets a certain threshold are weighted and summed to obtain the vector representation of the current low-frequency user. By using video content and user personalized representations simultaneously to generate comments, the automatically generated comments are more in line with the user's personalized needs, achieving the effect of customized automatic comments for users.

[0134] In summary, this application proposes a method for generating personalized automatic comments. When generating automatic comments for users, a deep vector description of the user is introduced to learn the user's personalized description. When providing the automatic comment function to the user, video content and user personalized information are used simultaneously to improve the personalized needs of the current user in the automatic comments, further improve the user adoption rate of the automatic comments, enhance the auxiliary effect of the automatic comment function on user interaction, and improve the efficiency of user video interaction.

[0135] Based on the same inventive concept, embodiments of this application also provide a comment generation device. For example... Figure 8 As shown, it is a structural schematic diagram of a comment generation device 800 proposed in an embodiment of this application, which may include:

[0136] The word segmentation processing unit 801 is configured to, in response to a comment request triggered by a target object for target multimedia content, perform word segmentation processing on the first description information of the target multimedia content to obtain each word in the first description information; and

[0137] The feature extraction unit 802 is used to obtain target description features for the target object based on the second description information of the target object and a preset set of candidate object description features;

[0138] The prediction unit 803 is used to predict the comment information that the target object will publish for the target multimedia content based on the target description features and each word segmentation.

[0139] Optionally, the feature extraction unit 802 is specifically used for:

[0140] If the target object is one of the candidate objects, then based on the object identifier in the second description information, the corresponding target description feature is queried from the candidate object description feature set;

[0141] If the target object is not one of the candidate objects, then based on the object profile in the second description information, determine the similarity between each candidate object and the target object; obtain the description features of at least one candidate object whose similarity reaches the first preset threshold from the candidate object description feature set; and determine the target description features of the target object based on the obtained description features of at least one candidate object.

[0142] Among them, the candidate objects are those whose number of comments reaches the second preset threshold.

[0143] Optionally, the feature extraction unit 802 is specifically used for:

[0144] If the number of at least one candidate object found is less than the set number, the corresponding similarity is used as a weight to perform a weighted summation of the descriptive features of the at least one candidate object found, and the target descriptive features are obtained.

[0145] If at least one candidate object is found to reach a set number, then based on the corresponding similarity, candidate objects that meet the set number are selected, and the corresponding similarity is used as a weight to perform a weighted summation of the descriptive features of each selected candidate object to obtain the target descriptive features.

[0146] Optionally, the feature extraction unit 802 is specifically used for:

[0147] For any candidate object, obtain an object profile of the candidate object, wherein the object profile includes at least one information tag and at least one tag weight corresponding to the information tag, and the information tag is obtained based on the analysis of the object's historical behavior;

[0148] Each information tag in the object profile of the target object is compared with at least one information tag in the object profile of any candidate object to determine the degree of overlap between each pair of information tags.

[0149] The sum of the products of the overlap between any two information tags and their corresponding weights is taken as the similarity between the target object and any candidate object. The weights corresponding to any two information tags are determined based on the tag weights of each information tag in each pair of information tags.

[0150] Optionally, the word segmentation processing unit 801 is specifically used for:

[0151] Input the first descriptive information of the target multimedia content into the trained comment generation model;

[0152] The first description information is segmented and encoded based on the encoding part of the trained comment generation model to obtain each segment of the first description information and the word vector of each segment.

[0153] Optionally, the prediction unit 803 is specifically used for:

[0154] The word vectors of each segment and the target description features are input into the decoding part of the trained comment generation model. Based on the decoding part, the decoding process is performed to obtain the comment information output by the trained comment generation model.

[0155] The trained comment generation model is trained on a training sample dataset. Each training sample in the training sample dataset includes an object profile of the sample object, descriptive information of the sample multimedia content, and real comments published by the sample object on the sample multimedia content.

[0156] Optionally, if the comment information includes multiple comment terms, the prediction unit 803 is specifically used for:

[0157] An iterative loop is used to generate each comment term in the comment information sequentially; during one iteration, the following operations are performed:

[0158] Input the previously output comment word into the decoding part, where the first input into the decoding part is a pre-set starting flag word;

[0159] Based on the comment words output in the previous output, the word vectors of each word segment, and the target description features, the comment words output in the current output are generated.

[0160] Optionally, the device also includes:

[0161] Model training unit 804 is used to train the pre-trained comment generation model in the following ways:

[0162] Select training samples from the training sample dataset;

[0163] The untrained comment generation model is iteratively trained using training samples to obtain a trained comment generation model. Each training iteration includes the following operations:

[0164] Input the object profiles of the sample objects in the training samples and the descriptive information of the sample multimedia content into the untrained comment generation model to obtain the predicted comments output by the untrained comment generation model.

[0165] Based on the error between the predicted comments and the real comments in the corresponding training samples, the parameters of the untrained comment generation model are adjusted.

[0166] Optionally, model training unit 804 is specifically used for:

[0167] Select training samples from the training sample dataset whose number of comments on the sample objects reaches a third preset threshold, or select training samples whose number of comments on the multimedia content they contain reaches a fourth preset threshold.

[0168] For ease of description, the above sections are divided into modules (or units) according to their functions and described separately. Of course, in implementing this application, the functions of each module (or unit) can be implemented in one or more software or hardware components.

[0169] Having described the comment generation method and apparatus according to exemplary embodiments of this application, we will now describe an electronic device according to another exemplary embodiment of this application.

[0170] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented as: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0171] Based on the same inventive concept as the above-described method embodiments, this application also provides an electronic device. This electronic device can be used to generate user reviews. In one embodiment, the electronic device can be a server, such as... Figure 1 The server 120 is shown. In this embodiment, the structure of the electronic device can be as follows: Figure 9 As shown, it includes a memory 901, a communication module 903, and one or more processors 902.

[0172] The memory 901 is used to store computer programs executed by the processor 902. The memory 901 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0173] Memory 901 may be volatile memory, such as random-access memory (RAM); memory 901 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 901 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 901 may be a combination of the above-described memories.

[0174] Processor 902 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 902 is used to implement the above-described comment generation method when it calls a computer program stored in memory 901.

[0175] The communication module 903 is used to communicate with terminal devices and other servers.

[0176] This application embodiment does not limit the specific connection medium between the memory 901, communication module 903, and processor 902 described above. This application embodiment... Figure 9 The memory 901 and the processor 902 are connected via a bus 904, which is in... Figure 9 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 904 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 9 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.

[0177] The memory 901 stores a computer storage medium, which stores computer-executable instructions for implementing the comment generation method of this application embodiment. The processor 902 is used to execute the above-described comment generation method, such as... Figure 2 As shown.

[0178] In another embodiment, the electronic device may also be other electronic devices, such as... Figure 1 The terminal device 110 is shown. In this embodiment, the electronic device can be structured as follows: Figure 10 As shown, it includes components such as: communication component 1010, memory 1020, display unit 1030, camera 1040, sensor 1050, audio circuit 1060, Bluetooth module 1070, processor 1080, etc.

[0179] The communication component 1010 is used to communicate with the server. In some embodiments, it may include a Circuit-Based Wireless Fidelity (WiFi) module. WiFi is a short-range wireless transmission technology, and electronic devices can use WiFi modules to help users send and receive information.

[0180] The memory 1020 can be used to store software programs and data. The processor 1080 executes various functions of the terminal device 110 and performs data processing by running the software programs or data stored in the memory 1020. The memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1020 stores an operating system that enables the terminal device 110 to run. In this application, the memory 1020 may store the operating system and various applications, and may also store code that executes the comment generation method of the embodiments of this application.

[0181] The display unit 1030 can also be used to display information input by the user or information provided to the user, as well as various menus of the terminal device 110, forming a graphical user interface (GUI). Specifically, the display unit 1030 may include a display screen 1032 disposed on the front of the terminal device 110. The display screen 1032 may be configured as a liquid crystal display, a light-emitting diode, or the like. The display unit 1030 can be used to display multimedia content playback screens in the embodiments of this application.

[0182] The display unit 1030 can also be used to receive input digital or character information and generate signal inputs related to user settings and function control of the terminal device 110. Specifically, the display unit 1030 may include a touch screen 1031 disposed on the front of the terminal device 110, which can collect touch operations of the user on or near it, such as clicking a button, dragging a scroll box, etc.

[0183] The touchscreen 1031 can be placed over the display screen 1032, or the touchscreen 1031 and the display screen 1032 can be integrated to realize the input and output functions of the terminal device 110. After integration, it can be referred to as a touch display screen. In this application, the display unit 1030 can display the application and the corresponding operation steps.

[0184] Camera 1040 can be used to capture still images, which users can then post comments on via an application. There can be one or multiple cameras 1040. An object is projected onto a photosensitive element through a lens, generating an optical image. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the processor 1080 to be converted into a digital image signal.

[0185] The terminal device may also include at least one sensor 1050, such as an accelerometer 1051, a proximity sensor 1052, a fingerprint sensor 1053, and a temperature sensor 1054. The terminal device may also be equipped with other sensors such as a gyroscope, barometer, hygrometer, thermometer, infrared sensor, light sensor, and motion sensor.

[0186] Audio circuitry 1060, speaker 1061, and microphone 1062 provide an audio interface between the user and terminal device 110. Audio circuitry 1060 converts received audio data into electrical signals, which are then transmitted to speaker 1061, where they are converted into sound signals for output. Terminal device 110 may also be equipped with volume buttons for adjusting the volume of the sound signal. Conversely, microphone 1062 converts collected sound signals into electrical signals, which are then received by audio circuitry 1060, converted back into audio data, and output to communication component 1010 for transmission to, for example, another terminal device 110, or to memory 1020 for further processing.

[0187] The Bluetooth module 1070 is used to interact with other Bluetooth devices that also have a Bluetooth module via the Bluetooth protocol. For example, a terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smartwatch) that also has a Bluetooth module through the Bluetooth module 1070, thereby exchanging data.

[0188] The processor 1080 is the control center of the terminal device, connecting various parts of the terminal through various interfaces and lines. It executes various functions and processes data by running or executing software programs stored in the memory 1020 and calling data stored in the memory 1020. In some embodiments, the processor 1080 may include one or more processing units; the processor 1080 may also integrate an application processor and a baseband processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the baseband processor mainly handles wireless communication. It is understood that the baseband processor may not be integrated into the processor 1080. In this application, the processor 1080 can run the operating system, applications, user interface display and touch response, and the comment generation method of this application embodiment. Furthermore, the processor 1080 is coupled to the display unit 1030.

[0189] In some possible implementations, various aspects of the comment generation method provided in this application can also be implemented as a program product comprising program code. When the program product is run on a computer device, the program code causes the computer device to perform the steps of the comment generation method according to the various exemplary embodiments of this application described above. For example, the computer device can perform actions such as... Figure 2 The steps are shown in the figure.

[0190] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0191] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a computing device. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with a command execution system, apparatus, or device.

[0192] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.

[0193] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0194] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0195] Alternatively, if the integrated units described in the embodiments of this application are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0196] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0197] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.< / s>

Claims

1. A method for generating video comments, characterized in that, The method comprises: In the process of the target object watching a target video, in response to a comment request triggered by the target object for the target video, performing word segmentation processing on first description information of the target video to obtain each word segmentation in the first description information; and Based on second description information of the target object and an object type, screening at least one description feature from a pre-constructed candidate object description feature set, and determining a target description feature of the target object based on the at least one description feature; the object type includes a high-frequency comment object and a low-frequency comment object, wherein: if the target object is one of the candidate objects, the target object is a high-frequency comment object, and the target description feature is a description feature obtained by querying the candidate object description feature set based on an object identifier in the second description information; if the target object is not one of the candidate objects, the target object is a low-frequency comment object, and the target description feature is determined based on a description feature of at least one candidate object whose similarity to the target object reaches a first preset threshold, the similarity between each candidate object and the target object being determined based on an object portrait in the second description information; and a comment frequency of the candidate object reaches a second preset threshold; Based on the target description feature and the each word segmentation, predicting comment information to be published by the target object for the target video, the comment information conforming to current video content of the target video.

2. The method of claim 1, wherein, The description feature of each candidate object in the candidate object description feature set is obtained based on machine learning.

3. The method of claim 1, wherein, Based on the description feature of at least one candidate object whose similarity to the target object reaches a first preset threshold obtained from the candidate object description feature set, determining a target description feature of the target object, specifically comprising: If the at least one candidate object queried does not reach a set number, the corresponding similarity is used as a weight to perform weighted summation on the description feature of the at least one candidate object queried to obtain the target description feature; If the at least one candidate object queried reaches a set number, the corresponding similarity is used as a weight to perform weighted summation on the description feature of each candidate object filtered out to obtain the target description feature.

4. The method of claim 1, wherein, The similarity between each candidate object and the target object is determined by the following method: For any one candidate object, an object portrait of the any one candidate object is obtained, wherein the object portrait includes at least one information label and a label weight corresponding to the at least one information label, and the information label is obtained based on object historical behavior analysis; Each information label in the object portrait of the target object is compared with at least one information label in the object portrait of the any one candidate object to determine a coincidence degree between each two information labels; and The sum of the product of the coincidence degree between each two information labels and the corresponding weight is taken as the similarity between the target object and the arbitrary candidate object, wherein the weight corresponding to each two information labels is determined according to the label weight of each information label in the two information labels.

5. The method of claim 1, wherein, The first description information of the target video is input into a trained comment generation model. The first description information is tokenized and encoded based on an encoding part in the trained comment generation model to obtain each token in the first description information and a word vector of each token. The word vector of each token and the target description feature are input into a decoding part in the trained comment generation model, and decoding processing is performed based on the decoding part to obtain the comment information output by the trained comment generation model.

6. The method of claim 5, wherein, The word vector of each token and the target description feature are input into a decoding part in the trained comment generation model, and decoding processing is performed based on the decoding part to obtain the comment information output by the trained comment generation model. The trained comment generation model is obtained by training based on a training sample data set, wherein each training sample in the training sample data set includes an object portrait of a sample object, description information of a sample video, and a real comment published by the sample object for the sample video. If the comment information includes multiple comment words, the word vector of each token and the target description feature are input into a decoding part in the trained comment generation model, and decoding processing is performed based on the decoding part to obtain the comment information output by the trained comment generation model, which specifically includes:

7. The method of claim 6, wherein, Each comment word in the comment information is generated in turn in a loop iteration manner; wherein in one loop iteration process, the following operations are performed: The last output comment word is input into the decoding part, wherein the first input into the decoding part is a pre-set start marker word; Based on the last output comment word, the word vector of each token, and the target description feature, decoding processing is performed to generate the current output comment word. The trained comment generation model is obtained by training in the following manner:

8. The method of any one of claims 5 to 7, wherein, Training samples are selected from the training sample data set; The untrained comment generation model is trained in a loop iteration manner according to the training samples to obtain the trained comment generation model, wherein each training iteration includes the following operations: The object portrait of the sample object and the description information of the sample video in the training sample are input into the untrained comment generation model to obtain the predicted comment output by the untrained comment generation model; Based on the error between the predicted comment and the real comment in the corresponding training sample, the parameters of the untrained comment generation model are adjusted. The training samples are selected from the training sample data set, which specifically includes:

9. The method of claim 8, wherein, ​ The training samples containing the sample objects whose comment times reach a third preset threshold are selected from the training sample data set, or the training samples containing the sample videos whose comment times reach a fourth preset threshold are selected.

10. A video comment generation apparatus characterized by comprising: The method comprises the following steps: The word segmentation processing unit is configured to, in response to a comment request triggered by the target object for the target video, perform word segmentation processing on first description information of the target video to obtain each word segmentation in the first description information during the process in which the target object watches the target video. The feature extraction unit is configured to select at least one description feature from a pre-constructed candidate object description feature set based on second description information of the target object and an object type, and determine a target description feature of the target object based on the at least one description feature. The object type comprises a high-frequency comment object and a low-frequency comment object.

11. The apparatus of claim 10, wherein, If the target object is one of the candidate objects, the target object is a high-frequency comment object, and the target description feature is a description feature obtained by querying the candidate object description feature set based on an object identifier in the second description information.

12. The apparatus of claim 10, wherein, If the target object is not one of the candidate objects, the target object is a low-frequency comment object, and the target description feature is determined based on a description feature of at least one candidate object whose similarity to the target object reaches a first preset threshold and obtained from the candidate object description feature set. The similarity between each candidate object and the target object is determined based on an object portrait in the second description information. The comment times of the candidate object reach a second preset threshold.

13. The apparatus of claim 10, wherein, The prediction unit is configured to predict comment information to be published by the target object for the target video based on the target description feature and the word segmentation. The description feature of each candidate object in the candidate object description feature set is obtained based on machine learning. The feature extraction unit is specifically configured to: If the number of the at least one candidate object queried does not reach a set number, the description features of the at least one candidate object queried are weighted and summed by taking the corresponding similarity as a weight to obtain the target description feature. If the number of the at least one candidate object queried reaches the set number, the candidate objects that meet the set number are selected according to the corresponding similarity, and the description features of the selected candidate objects are weighted and summed by taking the corresponding similarity as a weight to obtain the target description feature. The feature extraction unit is specifically configured to: For any one candidate object, an object portrait of the any one candidate object is obtained. Each information label in the object portrait of the target object is compared with at least one information label in the object portrait of the any one candidate object to determine the coincidence degree between each two information labels. The sum of the products of the coincidence degree between each two information tags and the corresponding weight is taken as the similarity between the target object and the arbitrary candidate object, wherein the weight corresponding to each two information tags is determined according to the tag weight of each information tag in the two information tags.

14. An electronic device, comprising: A computer program product, comprising a computer readable storage medium storing program code thereon, the program code, when executed by a processor, causing the processor to perform the steps of any of the methods of claims 1-9.

15. A computer readable storage medium, characterized in that, A computer program product, comprising program code, the program code, when executed on an electronic device, causing the electronic device to perform the steps of any of the methods of claims 1-9.

Citation Information

Patent Citations

  • Text information generation method and text information generation device

    CN110929021A

  • Comment information publishing method, device, client, server and system

    CN110968682A